From nobody Tue Sep 29 06:58:50 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B4128430CF4; Tue, 11 Aug 2026 09:31:39 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440701; cv=none; b=NiNydqV3QR7D7FPzA169jGf/+oaUJzatt+qlYV0ri83Qzk5bW45MQWTPhBeBvvUPceTw+mY8HivFoAHB9bsGKrPzU/GceUwBxB/Fuly6YM+fdY4CbYynEoYMnQJ3Rk/F92IUYJMmf4YyH6OaxncDhKu6x95/V/k1czztTIFrbyg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440701; c=relaxed/simple; bh=8QU1iUG0ifUlrmzPk7uU396y3BChOzVOZzCpIp3VUz0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=Loa/mGkdBUmVry7R5If20/qJOHcIMefnnEMrtaXeWq9Mw5MoUi8zyQeqribUdFuHTZ2NtCN1DsxpDn0uRu8a7Gj+u0oXA5g0CxxDS+MxdY9DVmW7iVBocD/weVAUCiFWcn+wmnd4BQDNe6gtQNjTb3+fhEU3ZMozWQdtiPEiYO4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=evGaNcXC; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="evGaNcXC" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 12EA51F000E9; Tue, 11 Aug 2026 09:31:38 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786440699; bh=5l7GlFPrn2SHLxQ09KpxmwggPPwZjLWS1qYbOKh84Fs=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=evGaNcXCZyMHjQ21iq4yxq/PEGO3lOnbtrF7MkGEOl0HCOOYhpa1LcxQ7uMIJrYJ7 smHycGuJ7ykSpKAOxjNJNeT0u6i8Q4DNF5VDVDnDZpqsXR7hhEqWxOjUtbVgUldoas G3jzULvd6GZgNJujDAkJ5eSQVL+8qpedceUkJotTP7EXT3Y+UWXmbSsxWRiQO0qSWm KZJaEGhtCd1fVGzd2kJM6Vc+p5HDMombQV/Wbw7a5FAcm/AUCu2YrpOq7xQxTsRciO oEBX2RpIwohh3qnrGHTIGjsMEgC6j1nz4XfZ1++RzGIQB25SWBPwfiq2+JuNSdS0Kx oWFiao4/WkH3Q== From: Leon Romanovsky To: Bjorn Helgaas , Logan Gunthorpe , Chaitanya Kulkarni , Greg Kroah-Hartman , Jens Axboe , Alex Williamson , Leon Romanovsky , Ankit Agrawal , Jason Gunthorpe , Jonathan Corbet , Shuah Khan , "Joerg Roedel (AMD)" , Will Deacon , Robin Murphy Cc: linux-pci@vger.kernel.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, iommu@lists.linux.dev Subject: [PATCH v3 01/17] PCI/P2PDMA: Do not tear down the allocate attribute on registration failure Date: Tue, 11 Aug 2026 12:30:43 +0300 Message-ID: <20260811-fix-p2p-acs-v3-1-efc488ee7c03@nvidia.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> References: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" X-Mailer: b4 0.15-dev-18f8f Content-Transfer-Encoding: quoted-printable From: Leon Romanovsky pci_p2pdma_add_resource() installs pci_p2pdma_unmap_mappings() as a devres action with the devres allocated p2p_pgmap as its data, and only then adds the range to the pool: error =3D devm_add_action_or_reset(&pdev->dev, pci_p2pdma_unmap_mappings, p2p_pgmap); if (error) goto pages_free; p2pdma =3D rcu_dereference_protected(pdev->p2pdma, 1); error =3D gen_pool_add_owner(p2pdma->pool, ...); if (error) goto pages_free; The action removes the allocate attribute for the whole device, which tears down existing userspace mappings of every BAR already registered on it. Both failures here get that wrong, in opposite ways. devm_add_action_or_reset() runs the action when it cannot allocate its devres node, so an -ENOMEM while registering a second BAR unmaps the first one. Use devm_add_action() and let the error path unwind only what this call created. gen_pool_add_owner() allocates a chunk and can also fail with -ENOMEM. There the action is registered, and the error path frees p2p_pgmap with devm_kfree() while leaving the action pointing at it. On unbind devres runs the action and pci_p2pdma_unmap_mappings() dereferences p2p_pgmap->mem->owner->kobj, which is freed memory. Give that failure its own label and drop the action with devm_remove_action(), which removes it without running it. Fixes: 7e9c7ef83d78 ("PCI/P2PDMA: Allow userspace VMA allocations through s= ysfs") Fixes: f58ef9d1d135 ("PCI/P2PDMA: Separate the mmap() support from the core= logic") Signed-off-by: Leon Romanovsky --- drivers/pci/p2pdma.c | 8 +++++--- 1 file changed, 5 insertions(+), 3 deletions(-) diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c index b2d5266f8653..dc7aaa990fed 100644 --- a/drivers/pci/p2pdma.c +++ b/drivers/pci/p2pdma.c @@ -440,8 +440,8 @@ int pci_p2pdma_add_resource(struct pci_dev *pdev, int b= ar, size_t size, goto pgmap_free; } =20 - error =3D devm_add_action_or_reset(&pdev->dev, pci_p2pdma_unmap_mappings, - p2p_pgmap); + error =3D devm_add_action(&pdev->dev, pci_p2pdma_unmap_mappings, + p2p_pgmap); if (error) goto pages_free; =20 @@ -451,13 +451,15 @@ int pci_p2pdma_add_resource(struct pci_dev *pdev, int= bar, size_t size, range_len(&pgmap->range), dev_to_node(&pdev->dev), &pgmap->ref); if (error) - goto pages_free; + goto mappings_remove; =20 pci_info(pdev, "added peer-to-peer DMA memory %#llx-%#llx\n", pgmap->range.start, pgmap->range.end); =20 return 0; =20 +mappings_remove: + devm_remove_action(&pdev->dev, pci_p2pdma_unmap_mappings, p2p_pgmap); pages_free: devm_memunmap_pages(&pdev->dev, pgmap); pgmap_free: --=20 2.55.0 From nobody Tue Sep 29 06:58:50 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B860E42FCBB; Tue, 11 Aug 2026 09:31:30 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440692; cv=none; b=azYMCy2OLkgFmTn0G/E/wLaBp9aY673yaHSliarV1pDgRGFrImcCDhfmag/eFR5ub3Ls2Oo81OE1ZmDqKZeU2GASq/FG1tKEecHJ8Aj8FBiQ8lHmNXcOH215FFQgtyR0RzalTYskvqLwwxy/GQ6rICr9x+UpiUYkspSMnRn+y5c= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440692; c=relaxed/simple; bh=3JQp9d6LSVeGbvHLkXDapuC3a1BmRVsswp3eW5KKo+k=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=gIyD1GhMO/Yzr7Q/GbdC2IQw/1Gpw0HfoW7tH2GrQx6BXnHk7eomKdZHCpoCIHzSW2hxrOa9wNmZo2VuQgCWTyYJICDxahL7JZ4PxSu1AeElMdGWJjqMkUwTLKodOHrKX8fvF8H6W6x6wNmwA8ev356r+1WAC+948A8iYV34PoM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=RGT+3j7W; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="RGT+3j7W" Received: by smtp.kernel.org (Postfix) with ESMTPSA id C17401F000E9; Tue, 11 Aug 2026 09:31:29 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786440690; bh=0MT/zjbdZ3nnGlvcYzzHFvINU8e49DgygVrlpjq5ySs=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=RGT+3j7WliLXyW7vwG6jfNBlIkfqPczEWoi6EEUiNhvBG8Y/Q8kNYqNI/fqBLT8rI qp17uP1B0Hz9OjD3+U5YLTiFXfAr8ZXu0inLh9v7NeOeeUnpYU/RLB/mhsSjbiZ4/z EtPOF4UzyPq3sxskZ3C18Q/Sq8VjLQ9Dk6LGx479Iz4F9jGZbZnPvcwm++K8DnuEst 7BAM4j3J06lfGNOYXx6Vl4lkeyw/8uLMx/byqxg+/VLqfkK6Ew3dtbhTtPRBgsFj+e mPyEWRaqE6D2vAnBgkaz8SQWz12J7PDES3buK+/5YEoXA6bUJSv42bhWJZ8XK5eYcW be5ghtMXjME2w== From: Leon Romanovsky To: Bjorn Helgaas , Logan Gunthorpe , Chaitanya Kulkarni , Greg Kroah-Hartman , Jens Axboe , Alex Williamson , Leon Romanovsky , Ankit Agrawal , Jason Gunthorpe , Jonathan Corbet , Shuah Khan , "Joerg Roedel (AMD)" , Will Deacon , Robin Murphy Cc: linux-pci@vger.kernel.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, iommu@lists.linux.dev, Matt Evans Subject: [PATCH v3 02/17] PCI/P2PDMA: Wait for RCU readers before freeing state Date: Tue, 11 Aug 2026 12:30:44 +0300 Message-ID: <20260811-fix-p2p-acs-v3-2-efc488ee7c03@nvidia.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> References: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" X-Mailer: b4 0.15-dev-18f8f Content-Transfer-Encoding: quoted-printable From: Leon Romanovsky pci_p2pmem_find_many() scans all PCI devices without locking or protection against driver unbind, including devices with poolless P2PDMA state. pci_has_p2pmem() may observe pdev->p2pdma just before driver unbind clears it, while pci_p2pdma_release() skips the grace period when no pool is present. This allows devres to free the object while it is still in use. Clear the pointer with RCU_INIT_POINTER() and always wait for pre-existing RCU readers before returning. The same grace period continues to protect gen_pool users for pool-backed providers. Cc: Alex Williamson Cc: Matt Evans Fixes: 372d6d1b8ae3 ("PCI/P2PDMA: Refactor to separate core P2P functionali= ty from memory allocation") Signed-off-by: Leon Romanovsky --- drivers/pci/p2pdma.c | 5 ++--- 1 file changed, 2 insertions(+), 3 deletions(-) diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c index dc7aaa990fed..e8e8c7d81d22 100644 --- a/drivers/pci/p2pdma.c +++ b/drivers/pci/p2pdma.c @@ -236,9 +236,8 @@ static void pci_p2pdma_release(void *data) return; =20 /* Flush and disable pci_alloc_p2p_mem() */ - pdev->p2pdma =3D NULL; - if (p2pdma->pool) - synchronize_rcu(); + RCU_INIT_POINTER(pdev->p2pdma, NULL); + synchronize_rcu(); xa_destroy(&p2pdma->map_types); =20 if (!p2pdma->pool) --=20 2.55.0 From nobody Tue Sep 29 06:58:50 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C0D3B430306; Tue, 11 Aug 2026 09:31:33 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440694; cv=none; b=giHZ8mDGiu5m/W8Cf9PeuM1bWIzPPV70tXvzdBsQjuTMkRADNchTBFdaJDalyqoXU6SHKNFJmEMu3FuCLUdUpqWuX01n22NAkSY5FsvW1YiAr+rapl9iTp7Ax330LjIUvtG31L7IMhlrTQLozsS3PBe24a3Ml/ij/BRMBNm77Jc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440694; c=relaxed/simple; bh=pvt/6Y5wtRCbuXKjA/n0yr3jC6JNDEbDzK0m617hMjU=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=MQJnxGX0BSe8leJvcUOJEdt/YDwJikt2+A6oLoqvZs5HEOA2oOPvwx6sVdGKssDAJ3PY3ocKZ5mmBxJhxurseT9DAdKvfh5mcY0ACWUanS77mAhBwNaZivBUiBg4AHvWdlAKrVW/hJSkN5PLUlWMsuHEeMFZb+2Mz90wxgKzn1I= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=a1s5ej56; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="a1s5ej56" Received: by smtp.kernel.org (Postfix) with ESMTPSA id D57A21F000E9; Tue, 11 Aug 2026 09:31:32 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786440693; bh=UrX2Q3nfzyJlkJf2wLcoexGjPgTP4y3dmCJEYcVPMvM=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=a1s5ej56W2Uv6gRAxsgGpkBTWggp3NWeadN6i2t0QiN4pRFhugakRn4cz0bmUwrtD xm3gCDMKvCiubXsvLphyJd6i6O5U7DkTYDizJGiKmJ6qdtSBDAYL4a4jDNbvcXNGHI 9uIa9bex7BDscFYRtIPYHHlvMn+KkVKxTpX9Y9QnTQ12HYZubkbb1YNFoMiVZieiuL cNP7MC9yw8SiZzDaBibW1rGg+fWo2oTADtKDxtbGMFIK5D7LkBqniIiW/ML4IysqP1 S1zAZS4gGE2YoO2dKBFgBTjc9yMzoXLy9w49MK879Cn+44o+VGXhdpQWRGN4xtfCar b4dBDvRw0y0YQ== From: Leon Romanovsky To: Bjorn Helgaas , Logan Gunthorpe , Chaitanya Kulkarni , Greg Kroah-Hartman , Jens Axboe , Alex Williamson , Leon Romanovsky , Ankit Agrawal , Jason Gunthorpe , Jonathan Corbet , Shuah Khan , "Joerg Roedel (AMD)" , Will Deacon , Robin Murphy Cc: linux-pci@vger.kernel.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, iommu@lists.linux.dev Subject: [PATCH v3 03/17] PCI/P2PDMA: Restrict the p2pmem search to pool backed providers Date: Tue, 11 Aug 2026 12:30:45 +0300 Message-ID: <20260811-fix-p2p-acs-v3-3-efc488ee7c03@nvidia.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> References: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" X-Mailer: b4 0.15-dev-18f8f Content-Transfer-Encoding: quoted-printable From: Leon Romanovsky pci_p2pmem_find_many() exists to pick a provider that the caller will then allocate from with pci_alloc_p2pmem(), which goes straight to the gen_pool: ret =3D (void *)gen_pool_alloc_owner(p2pdma->pool, size, (void **) &ref); pci_has_p2pmem() does not ask for that pool, only for the published flag. The two used to be equivalent, because a provider could only exist by way of pci_p2pdma_add_resource(), which always creates the pool. pcim_p2pdma_init() broke that. It registers a provider for the DMABUF path and never creates a pool, so pdev->p2pdma is set while p2pdma->pool stays NULL. Nothing publishes such a provider today, so the search cannot return one yet, but the flag alone no longer says what the caller needs. Ask for the pool as well, so the search covers the providers its result is used for. A later patch documents the pdev->p2pdma lifetime and RCU rules. Signed-off-by: Leon Romanovsky --- drivers/pci/p2pdma.c | 7 ++++++- 1 file changed, 6 insertions(+), 1 deletion(-) diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c index e8e8c7d81d22..6618ef170ce1 100644 --- a/drivers/pci/p2pdma.c +++ b/drivers/pci/p2pdma.c @@ -863,7 +863,12 @@ static bool pci_has_p2pmem(struct pci_dev *pdev) =20 rcu_read_lock(); p2pdma =3D rcu_dereference(pdev->p2pdma); - res =3D p2pdma && p2pdma->p2pmem_published; + /* + * The callers hand the result to pci_alloc_p2pmem(), so only a + * provider backed by a pool is of any use here. pcim_p2pdma_init() + * creates providers without one. + */ + res =3D p2pdma && p2pdma->pool && p2pdma->p2pmem_published; rcu_read_unlock(); =20 return res; --=20 2.55.0 From nobody Tue Sep 29 06:58:50 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E6381430CE1; Tue, 11 Aug 2026 09:31:36 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440698; cv=none; b=FydxmHvWphVDDcJuCZ/l9EA8Zqj4zIpIrBqwZ9p+LNU4WeTeItikXWeWdHedw6aS7DyE7wVX9X8mSSI9rCopgT3c+YCxp9wPx4p4Ixcw4eydcPZjC8X51A4cwvYc36m2FavBHU6DDPpWs6B/3IRLf7vxgwcFStX+GH2urGVqIJk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440698; c=relaxed/simple; bh=Jvc8O9aB4fr5CPxCyMYzfrAWdkXrm+ubP7j0sCQsulE=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=e9vbsl/tq0/QZFfzzDEy7qMzWCpteKzABHiqYGxy+pY5C5Q45nKIBM30wFFTXL+BtrQYJ9VgRz70lRlZ4e+iUczH33K3VStI5Bz7YYH+N3KmujGpSEwW9ZzPB83IDGTxPdeDgUZjoL0R/N1Iaxin6XIXjDYeTnA/nChqR/yOiG0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=BAb95SzX; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="BAb95SzX" Received: by smtp.kernel.org (Postfix) with ESMTPSA id E9BE31F000E9; Tue, 11 Aug 2026 09:31:35 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786440696; bh=atU9QUcTmgMbDE3QFRiM5Hhj+AAChNXCTaJk5Hfy/L4=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=BAb95SzXBW2y2NX3+VCObdFzlsE2VUWRMJYljCDqk/vzYrpKXfm7iG7kSnqZ7SaWa qUf+LaaM4NyxyTD2jiSXHYR9Dnu26BpOW7HDlHZTxi6KmYcVyrTmrCbgxTSrtp+2X9 6bIWaXIAzO3MZqqPmYd732dLO/RUet3pq36WeSAy9zPVzaadikKsmvHNWxixu4l5Fd 7vO/XqCxHXYe+nrLu1Y20+lp6ukVRS6DztmlSkv92ThXWDKXDjPPPF7cEaqo0FGNJu zQFs5XIS5sm5zO8DeaANeocGC+E7o2NjAePVu6Kak8t1NXuno+Pg6+icQmhh8n32HH UyCDG4zLIXdcg== From: Leon Romanovsky To: Bjorn Helgaas , Logan Gunthorpe , Chaitanya Kulkarni , Greg Kroah-Hartman , Jens Axboe , Alex Williamson , Leon Romanovsky , Ankit Agrawal , Jason Gunthorpe , Jonathan Corbet , Shuah Khan , "Joerg Roedel (AMD)" , Will Deacon , Robin Murphy Cc: linux-pci@vger.kernel.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, iommu@lists.linux.dev Subject: [PATCH v3 04/17] PCI/P2PDMA: Safely terminate ACS redirect lists Date: Tue, 11 Aug 2026 12:30:46 +0300 Message-ID: <20260811-fix-p2p-acs-v3-4-efc488ee7c03@nvidia.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> References: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" X-Mailer: b4 0.15-dev-18f8f Content-Transfer-Encoding: quoted-printable From: Leon Romanovsky seq_buf marks an overflow by setting len to size + 1. The ACS diagnostic path unconditionally writes a terminator to buffer[len - 1], so a path with enough ACS ports to fill the 128-byte buffer writes one byte beyond the buffer when verbose diagnostics are requested. Use seq_buf_str() to terminate truncated output safely and remove the final semicolon only when the buffer did not overflow. Fixes: 52916982af48 ("PCI/P2PDMA: Support peer-to-peer memory") Reviewed-by: Logan Gunthorpe Signed-off-by: Leon Romanovsky --- drivers/pci/p2pdma.c | 6 ++++-- 1 file changed, 4 insertions(+), 2 deletions(-) diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c index 6618ef170ce1..a77ef9deb3c6 100644 --- a/drivers/pci/p2pdma.c +++ b/drivers/pci/p2pdma.c @@ -767,11 +767,13 @@ calc_map_type_and_dist(struct pci_dev *provider, stru= ct pci_dev *client, } =20 if (verbose) { - acs_list.buffer[acs_list.len-1] =3D 0; /* drop final semicolon */ + /* Drop the final semicolon; the list is not empty here. */ + if (!seq_buf_has_overflowed(&acs_list)) + acs_list.buffer[acs_list.len - 1] =3D '\0'; pci_warn(client, "ACS redirect is set between the client and provider (%= s)\n", pci_name(provider)); pci_warn(client, "to disable ACS redirect for this path, add the kernel = parameter: pci=3Ddisable_acs_redir=3D%s\n", - acs_list.buffer); + seq_buf_str(&acs_list)); } acs_redirects =3D true; =20 --=20 2.55.0 From nobody Tue Sep 29 06:58:50 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 03428437847; Tue, 11 Aug 2026 09:31:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440720; cv=none; b=PXMreIEmAwbv9McU+JT7ALS3Ft6gXOOEcs5BYC1BPGc1FbnP6lxFLIUQfzvhJUjy0o6yLhiNqGDAE55aaDwyCOZwTUg0EQ4xiEEgv3dCB4v+ExiQlWPH0KIOSZCx3K6Jfc+RpFiMCclsJrQIsxUkUePRaNqlyCb+KVOpmeTpEy4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440720; c=relaxed/simple; bh=eTInsLR85gfWDKjsvkqG7KAzrkPcyAsKb8FqdKSMXPA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=K6oTpBnBuOSHfsGRrxoWz05REkmV/eFGa+RaQM28D91l6nM41+w+WN7Bpc3IH7ZZY5BibT204jiUBY35XEAI7FoMMeifG+DoiSDm3GY523IICk7+ACR5MrS5s2MWoT6INu4QbjesuGFVxx3yvAMrg62V5256LNzoqw36fDyOCf0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=hBjyWbfm; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="hBjyWbfm" Received: by smtp.kernel.org (Postfix) with ESMTPSA id EB4021F000E9; Tue, 11 Aug 2026 09:31:57 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786440718; bh=1lqvi2W8LPNtWb/wvNjdUeGAFyD0nbgTZu4Qg0okxTY=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=hBjyWbfmrP7U8ajJNRSTArgqifEy21v0yjeVilRNJOaKDBqGQrp7erwxfyNiV6z60 Q5Wx9nr6HBXmDfI3EHAu4iM9l2POicFusC1RHnDDDpFmC341DoFjtCT6vJp6+pTctU gpMIiF+drFvyZisTTorUSLfAwuHM2E+sjlWjD8deeL1yF+NgnhbkXhrsEkOeexzVzn Vw/MqdWPtlfpMEJaFuxS6Q7lOpXX+Q4unMXe8CIqMlJmB7lzBJWzPO45AKbdOHTeR0 Q915Kg9ri/EbuCSZcPFvEGqM73oqfVPmwXdS1W7N6R0I+WoY7FGePEYjGjJT6K2MAw uw9O1VHIxTZGg== From: Leon Romanovsky To: Bjorn Helgaas , Logan Gunthorpe , Chaitanya Kulkarni , Greg Kroah-Hartman , Jens Axboe , Alex Williamson , Leon Romanovsky , Ankit Agrawal , Jason Gunthorpe , Jonathan Corbet , Shuah Khan , "Joerg Roedel (AMD)" , Will Deacon , Robin Murphy Cc: linux-pci@vger.kernel.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, iommu@lists.linux.dev, Matt Evans Subject: [PATCH v3 05/17] PCI/P2PDMA: Document the pdev->p2pdma lifetime and RCU rules Date: Tue, 11 Aug 2026 12:30:47 +0300 Message-ID: <20260811-fix-p2p-acs-v3-5-efc488ee7c03@nvidia.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> References: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" X-Mailer: b4 0.15-dev-18f8f Content-Transfer-Encoding: quoted-printable From: Leon Romanovsky pdev->p2pdma has two lifetime models. Provider-based entry points are quiesced by their driver before remove completes. pci_p2pmem_find_many() and the p2pmem sysfs attributes can race with unbind and therefore rely on the teardown grace period. Document publication, teardown, and how the grace period protects both the struct pci_p2pdma object and its optional gen_pool. Cc: Alex Williamson Cc: Matt Evans Signed-off-by: Leon Romanovsky --- drivers/pci/p2pdma.c | 54 ++++++++++++++++++++++++++++++++++++++++++++++++= +++- 1 file changed, 53 insertions(+), 1 deletion(-) diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c index a77ef9deb3c6..49bc8cf06240 100644 --- a/drivers/pci/p2pdma.c +++ b/drivers/pci/p2pdma.c @@ -21,6 +21,39 @@ #include #include =20 +/* + * Lifetime and RCU usage + * + * Within one driver bind, pdev->p2pdma is published exactly once, by + * pcim_p2pdma_init(), and cleared exactly once, by the pci_p2pdma_release= () + * devres action that the same function installs. It is never re-pointed a= t a + * second struct pci_p2pdma, so a reader that observes a non-NULL pointer + * always observes the same, fully initialised object. That object is devr= es + * memory allocated before the action is installed, so devres frees it only + * after pci_p2pdma_release() has returned. + * + * Most exported entry points reach pdev->p2pdma through a struct pci_dev = or a + * struct p2pdma_provider owned by the provider driver, and + * pcim_p2pdma_provider() requires callers to drop those references before= the + * driver's remove() completes. Those cannot run concurrently with + * pci_p2pdma_release(), and their rcu_dereference() calls are simply how = an + * __rcu pointer is read. + * + * pci_p2pmem_find_many() and the p2pmem sysfs attributes are the exceptio= ns. + * The first walks every PCI device, so it can reach a provider whose driv= er is + * unbinding: pci_get_device() pins the struct pci_dev, not the driver. The + * second is reachable from userspace until sysfs_remove_group() runs at t= he end + * of the release. pci_has_p2pmem() must dereference the object to determi= ne + * whether it owns a gen_pool, so even a poolless object must remain alive= until + * that RCU reader exits. The sysfs group is created with the pool. + * + * The grace period in pci_p2pdma_release() first protects the struct + * pci_p2pdma itself from being freed while pci_has_p2pmem() is using it. = For a + * pool-backed provider it also fences the gen_pool: gen_pool_alloc_owner() + * walks pool->chunks under RCU and gen_pool_destroy() frees those chunks + * without waiting for a grace period of its own, so pci_alloc_p2pmem() and + * p2pmem_alloc_mmap() hold rcu_read_lock() across the allocation. + */ struct pci_p2pdma { struct gen_pool *pool; bool p2pmem_published; @@ -235,9 +268,19 @@ static void pci_p2pdma_release(void *data) if (!p2pdma) return; =20 - /* Flush and disable pci_alloc_p2p_mem() */ + /* + * Stop new RCU readers and wait for readers that observed p2pdma before + * allowing devres to free it. This is required even without a pool, + * because pci_has_p2pmem() dereferences every non-NULL p2pdma it finds. + * For a pool-backed provider this also fences gen_pool_destroy(). + */ RCU_INIT_POINTER(pdev->p2pdma, NULL); synchronize_rcu(); + + /* + * The grace period also ensures no RCU reader can still be accessing + * map_types here. + */ xa_destroy(&p2pdma->map_types); =20 if (!p2pdma->pool) @@ -255,6 +298,9 @@ static void pci_p2pdma_release(void *data) * for a PCI device. It allocates and sets up the necessary data * structures to support P2PDMA operations, including mapping type * tracking. + * + * The state is published once per driver bind and torn down by a devres + * action on unbind. Repeated calls for the same device are a no-op. */ int pcim_p2pdma_init(struct pci_dev *pdev) { @@ -786,6 +832,12 @@ calc_map_type_and_dist(struct pci_dev *provider, struc= t pci_dev *client, map_type =3D PCI_P2PDMA_MAP_NOT_SUPPORTED; } done: + /* + * pci_p2pmem_find_many() reaches this with a provider whose driver may + * be unbinding, so the store runs under RCU: pci_p2pdma_release() + * clears the pointer and waits for readers before destroying + * map_types. See "Lifetime and RCU usage" above. + */ rcu_read_lock(); p2pdma =3D rcu_dereference(provider->p2pdma); if (p2pdma) --=20 2.55.0 From nobody Tue Sep 29 06:58:50 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0EC36432BD8; Tue, 11 Aug 2026 09:31:42 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440704; cv=none; b=O7a6DeGj982OTMrUrXjW7Ht/KNBQnJP3lekLCEl0xAPPODuZxmQnLJ0r22BvqAfmH1R09Iy4KVQPjnaCymR+8jkc18rllLj8oPulNBmB6bz+UgVA2mCfOaZDkd9wBGJdixb4jotqLlIHMqbYzyvkePpnEcp0dUIjMdbORKyHe7Q= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440704; c=relaxed/simple; bh=sjQq/3/8/CO1p0c3gG4LtOOqzshhDDbyqBdeCh4NksU=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=DYX0IylKRGPq7DOUPE+S3wS2IgMWabeUDmqG3c4i8FWkO0quj8FJJkUBA7fhh3eZjwxy1YTtfPULeRVLECI9Co5UFBRDbqHFTnng2cWtkkHE9Kk+Q91e0QHzk4z61jaCLYwLbSx9xbGQz/5QyCPB0J5cSxbtqW5ptG++A6uh8Ww= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=K1mxOb7x; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="K1mxOb7x" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 332C21F000E9; Tue, 11 Aug 2026 09:31:42 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786440702; bh=wFlI5ofQqed5SZyALyQHqTpJ4HaBv8zoy6KujpstP9Y=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=K1mxOb7xHoTEC80miMPtvsgHnex4Ymp0ItQ25yRiXpiUeYnorUEtc8dSYxKFu8g0Q OF5gqlsp0fCsQVdioJz2L0Im8ZFVkh1q1xDvh0WYGVc8RCchcii2+jh1LmKsED6UnO hvSfuuBLbZohWtidtQKs5HB5fucQTKaaxBVSfmhqh+d0ox3vZbqWBX9vJ+tHzAYE4T Q570Yz1QNZ6TVKKBVU0UKf169uIPakVUuO1vY1WSF8R+8zsi4G6RQUJ2CmcqfWETt2 SwEsnY5AgiVSnvJdpX6wdW/sQD5uEGMkj7olA/q9NtBcCydj015fYTn9beB7wuJ6MN 2yIPtK09R9b6Q== From: Leon Romanovsky To: Bjorn Helgaas , Logan Gunthorpe , Chaitanya Kulkarni , Greg Kroah-Hartman , Jens Axboe , Alex Williamson , Leon Romanovsky , Ankit Agrawal , Jason Gunthorpe , Jonathan Corbet , Shuah Khan , "Joerg Roedel (AMD)" , Will Deacon , Robin Murphy Cc: linux-pci@vger.kernel.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, iommu@lists.linux.dev Subject: [PATCH v3 06/17] PCI/P2PDMA: Gate the host bridge whitelist warning on verbose Date: Tue, 11 Aug 2026 12:30:48 +0300 Message-ID: <20260811-fix-p2p-acs-v3-6-efc488ee7c03@nvidia.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> References: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" X-Mailer: b4 0.15-dev-18f8f Content-Transfer-Encoding: quoted-printable From: Leon Romanovsky calc_map_type_and_dist() prints every other diagnostic under its verbose argument, but reaches the "Host bridge not in P2PDMA whitelist" warning through host_bridge_whitelist(), which it hands acs_redirects instead. A caller that asked for a silent answer still gets the warning whenever any port on the path has an ACS redirect bit set, the CPU is not whitelisted by cpu_supports_p2pdma(), and the host bridge is not in pci_p2pdma_whitelist[]. pci_p2pmem_find_many() is such a caller. It sweeps every device with published p2pmem and asks for the distance to each client with verbose=3Dfalse, and pci_p2pdma_distance_many() recomputes rather than consulting the map_types cache, so the warning repeats on every sweep. The argument was never meant to say "ACS redirects were found". When commit cf201bfe8cdc ("PCI/P2PDMA: Warn if host bridge not in whitelist") added it, acs_redirects was a bool pointer that the quiet entry point passed as NULL: if (verbose) map =3D calc_map_type_and_dist_warn(provider, pci_client, &distance); else map =3D calc_map_type_and_dist(provider, pci_client, &distance, NULL, NULL); so the argument was true on exactly the path that commit describes. Folding the two entry points into one verbose flag turned the pointer into a value and left the call site alone, silently narrowing the warning to paths that carry an ACS redirect. Pass verbose. This also restores the warning for a verbose caller that takes the host bridge route with no ACS redirect on the path, which until now was told it could not use peer-to-peer DMA without being told which vendor and device would have to be added to the whitelist. Fixes: d1b8dc09dd71 ("PCI/P2PDMA: Simplify distance calculation") Signed-off-by: Leon Romanovsky --- drivers/pci/p2pdma.c | 4 +--- 1 file changed, 1 insertion(+), 3 deletions(-) diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c index 49bc8cf06240..a364008bbf50 100644 --- a/drivers/pci/p2pdma.c +++ b/drivers/pci/p2pdma.c @@ -750,7 +750,6 @@ calc_map_type_and_dist(struct pci_dev *provider, struct= pci_dev *client, { enum pci_p2pdma_map_type map_type =3D PCI_P2PDMA_MAP_THRU_HOST_BRIDGE; struct pci_dev *a =3D provider, *b =3D client, *bb; - bool acs_redirects =3D false; struct pci_p2pdma *p2pdma; struct seq_buf acs_list; int acs_cnt =3D 0; @@ -821,11 +820,10 @@ calc_map_type_and_dist(struct pci_dev *provider, stru= ct pci_dev *client, pci_warn(client, "to disable ACS redirect for this path, add the kernel = parameter: pci=3Ddisable_acs_redir=3D%s\n", seq_buf_str(&acs_list)); } - acs_redirects =3D true; =20 map_through_host_bridge: if (!cpu_supports_p2pdma() && - !host_bridge_whitelist(provider, client, acs_redirects)) { + !host_bridge_whitelist(provider, client, verbose)) { if (verbose) pci_warn(client, "cannot be used for peer-to-peer DMA as the client and= provider (%s) do not share an upstream bridge or whitelisted host bridge\n= ", pci_name(provider)); --=20 2.55.0 From nobody Tue Sep 29 06:58:50 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3A37D42AF9B; Tue, 11 Aug 2026 09:31:47 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440711; cv=none; b=NK1gXA9T77x78ioZ/4Otgbtd8bExBfXq0KP6e9ZjgOLuL0ksiVKl9lbmv0mLT3/grxH1jO6FVVJARw51k7fcBY7480YIcAKT78y5fU9NsRz55zpHoBL/COQpOu8NRbZNjBq9GtJQfeI5e6oJ2yZmNrfSiJ8K80uxwZvfQHdUpEU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440711; c=relaxed/simple; bh=nxWNJX3YUnQ3wS+Ge+YXFuztqyh0Jb5li7SRLsAC6nw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=VL4SH8KlSsNTGW/BAXdJc/5gxCB8P7agxKQTkT3m7jSDfOzrE04WsRm4E7FPPQGUr/SilfBH/mlWd7yMAH6a249R6MplPhHPMR/EAKx1xWIA/GzbIYKKtiePUR/yfPhSq89mxSAqaoQXFFYHyYUmS8btrR2MnlU+EX8VD/6sjPU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=jpJ1NCVD; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="jpJ1NCVD" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 5C78B1F000E9; Tue, 11 Aug 2026 09:31:45 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786440705; bh=mF5n3YOfW5QrsqIP73nDBBVRmhqCEMlU86R6zm1GxNY=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=jpJ1NCVD2jjpQURmMMkbm0pOkveTCWptwy52n+JyUOt8zZpyBX4im0Hz/vQaO7pJj /0s7h4KpqsQc7uV/joZIM+t0bhPfxfkQd8OJKIHGsO2PGREhO404xXo+iyrEYzv+7s snvAqElSUZFBIJLyaHWeIPjC77U0i5WAiRBaN+j24xcM/YV/xVbMHhQ1oTsXf7eh7S sKn4k0WApcOEGMgp24Ms7pKexZho4In5zfhjWstnJ1+K/wThBxGEjJw1E5s2aHuKHY FOTlXmhXbkabcjn5HI0SP2AxzoVIamlb7ucwbSzdtubhdQpMgf85V5s4+25C1KKbTa qqYkvZhJo2KoQ== From: Leon Romanovsky To: Bjorn Helgaas , Logan Gunthorpe , Chaitanya Kulkarni , Greg Kroah-Hartman , Jens Axboe , Alex Williamson , Leon Romanovsky , Ankit Agrawal , Jason Gunthorpe , Jonathan Corbet , Shuah Khan , "Joerg Roedel (AMD)" , Will Deacon , Robin Murphy Cc: linux-pci@vger.kernel.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, iommu@lists.linux.dev Subject: [PATCH v3 07/17] PCI/P2PDMA: Document the Address Type assumption Date: Tue, 11 Aug 2026 12:30:49 +0300 Message-ID: <20260811-fix-p2p-acs-v3-7-efc488ee7c03@nvidia.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> References: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" X-Mailer: b4 0.15-dev-18f8f Content-Transfer-Encoding: quoted-printable From: Leon Romanovsky P2PDMA selects a mapping from the ACS controls that govern Requests carrying an Untranslated address. PCIe r7.0, sec 6.12.3 routes a Translated Request directly to the peer when ACS Direct Translated P2P is enabled, regardless of P2P Request Redirect and P2P Egress Control. Translation Blocking takes precedence and prevents that direct route. Document this assumption because an ATS capable client can otherwise reach the peer directly whichever mapping P2PDMA selects. Reviewed-by: Logan Gunthorpe Signed-off-by: Leon Romanovsky --- Documentation/driver-api/pci/p2pdma.rst | 7 +++++++ 1 file changed, 7 insertions(+) diff --git a/Documentation/driver-api/pci/p2pdma.rst b/Documentation/driver= -api/pci/p2pdma.rst index d3f406cca694..a7fd426c3685 100644 --- a/Documentation/driver-api/pci/p2pdma.rst +++ b/Documentation/driver-api/pci/p2pdma.rst @@ -15,6 +15,13 @@ then based on the ACS settings the transaction can route= entirely within the PCIe hierarchy and never reach the root port. The kernel will evaluate the PCIe topology and always permit P2P in these well-defined cases. =20 +This evaluation covers the ACS controls that govern Requests carrying an +Untranslated address. Unless ACS Translation Blocking is enabled, a Port +with ACS Direct Translated P2P enabled routes a Request carrying a Transla= ted +address directly to the peer regardless of those controls. An ATS capable +client may therefore reach the peer on the direct path whichever mapping t= he +kernel selects. + However, if the P2P transaction reaches the host bridge then it might have= to hairpin back out the same root port, be routed inside the CPU SOC to anoth= er PCIe root port, or routed internally to the SOC. --=20 2.55.0 From nobody Tue Sep 29 06:58:50 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9CFFC432E81; Tue, 11 Aug 2026 09:31:50 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440715; cv=none; b=B13BjLjhOPohzOMk+oR6TMNPoR/5cOy52E1srPrwduMQbDX6ZqpcBwowIpmmRzONAEuSGpdmDRcbh3GwjYYaIKh3lTnu03B/ZarJPMwp8XJ4mNG7ZPFqBo8uGcAtqJjZj9xnCinjOQZKLhY0sIocpDBsfaN0Youc3a+HD+p88Sw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440715; c=relaxed/simple; bh=LdDpnv3b8uEfo6MzTz/l8sGRCNsIkdazjpnPKSY5Z80=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=p6vQC+yjYoVIiUMT5vRREtC/kmuiQU0jySTFqW1s07SAzT2N9g536lslmFld+wn3hiMpXBkfNFCG64X0DopwUuxNDiyC0E3dYtC/TZwYmiYKcjKo3Dw4ik3oSswQA+mcSY5iiw42k5MHZQYvbPktZskANQHqakBV/W9rekaZB88= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=OTVhFjx2; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="OTVhFjx2" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 92AC91F00A3A; Tue, 11 Aug 2026 09:31:48 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786440709; bh=MUUhR6qcUEIri/XXh4neJYaArNo8BNREmI9KueUYZ10=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=OTVhFjx2SwZeZiDiTmEUXh/T0bcYPq1vT9KiF63NMv9i6ccfh5Thgk2hJdmJ9Ur/4 iSNHrgYe/etkdOXzcrzkZuA6/XHsNO32PuVB4GFf4BnsoFBHBalDdBE56ff+ZILO+7 WEiPD3XsikU0W9LXEVy7WwhQU6evC6eR19GJ9pTJlAYkkLtEYYlyKZtt4tS53NkLYc CD5tmvpfrgem7k4nXgqp0S2Txbu1n5JlgRs0N7I3dPaGCJ2SgkxPDXJHROvk8aivBS y1tLaurp/FmHVfvIg8IVP4UO53PaOXyPO9srcJIQlFKKiOGSEBVHgFDv91AysOmr2a DqKA/TEfaWtOA== From: Leon Romanovsky To: Bjorn Helgaas , Logan Gunthorpe , Chaitanya Kulkarni , Greg Kroah-Hartman , Jens Axboe , Alex Williamson , Leon Romanovsky , Ankit Agrawal , Jason Gunthorpe , Jonathan Corbet , Shuah Khan , "Joerg Roedel (AMD)" , Will Deacon , Robin Murphy Cc: linux-pci@vger.kernel.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, iommu@lists.linux.dev Subject: [PATCH v3 08/17] PCI: Account for Direct Translated P2P in ACS isolation checks Date: Tue, 11 Aug 2026 12:30:50 +0300 Message-ID: <20260811-fix-p2p-acs-v3-8-efc488ee7c03@nvidia.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> References: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" X-Mailer: b4 0.15-dev-18f8f Content-Transfer-Encoding: quoted-printable From: Leon Romanovsky PCIe r7.0, sec 6.12.3: peer-to-peer Memory Requests whose Address Type (AT) field indicates a Translated address must be routed to the peer Port/Function without redirection, regardless of ACS P2P Request Redirect and ACS P2P Egress Control settings. Request Redirect therefore does not isolate devices below a Port with ACS Direct Translated P2P enabled. Sec 6.12.1.1 makes such a Request an ACS Violation once Translation Blocking is enabled, and that error "must take precedence over ... ACS P2P control mechanisms". Report isolation only in that case. Without Translation Blocking, devices below such a Port now share an IOMMU group. This only holds for a caller that needs Request Redirect to isolate peers. pci_enable_pasid() asks for Request Redirect for a different reason: a Request carrying a PASID is routed by address alone (sec 2.2.10.4), so it has to be redirected Upstream to reach the translation agent. Direct Translated P2P says nothing about that, because a Translated Request already carries an address the agent produced for that PASID (sec 10.1.3). Give pci_acs_enabled() and pci_acs_path_enabled() a scope so each caller states which Requests its answer has to cover, and apply the rule above only for PCI_ACS_SCOPE_ALL. pci_acs_flags_enabled() and the Intel SPT PCH quirk both need the rule, so it lives in pci_acs_rr_ineffective(). Fixes: ad805758c0eb ("PCI: add ACS validation utility") Signed-off-by: Leon Romanovsky --- drivers/iommu/iommu.c | 8 +++++--- drivers/pci/ats.c | 11 +++++++++- drivers/pci/pci.c | 26 +++++++++++++++-------- drivers/pci/pci.h | 30 +++++++++++++++++++++++++-- drivers/pci/quirks.c | 57 ++++++++++++++++++++++++++++++++++-------------= ---- include/linux/pci.h | 30 +++++++++++++++++++++++---- 6 files changed, 124 insertions(+), 38 deletions(-) diff --git a/drivers/iommu/iommu.c b/drivers/iommu/iommu.c index e8f13dcebbde..6ab32d714ce5 100644 --- a/drivers/iommu/iommu.c +++ b/drivers/iommu/iommu.c @@ -1502,13 +1502,14 @@ static struct iommu_group *get_pci_function_alias_g= roup(struct pci_dev *pdev, struct pci_dev *tmp =3D NULL; struct iommu_group *group; =20 - if (!pdev->multifunction || pci_acs_enabled(pdev, REQ_ACS_FLAGS)) + if (!pdev->multifunction || + pci_acs_enabled(pdev, REQ_ACS_FLAGS, PCI_ACS_SCOPE_ALL)) return NULL; =20 for_each_pci_dev(tmp) { if (tmp =3D=3D pdev || tmp->bus !=3D pdev->bus || PCI_SLOT(tmp->devfn) !=3D PCI_SLOT(pdev->devfn) || - pci_acs_enabled(tmp, REQ_ACS_FLAGS)) + pci_acs_enabled(tmp, REQ_ACS_FLAGS, PCI_ACS_SCOPE_ALL)) continue; =20 group =3D get_pci_alias_group(tmp, devfns); @@ -1652,7 +1653,8 @@ struct iommu_group *pci_device_group(struct device *d= ev) if (!bus->self) continue; =20 - if (pci_acs_path_enabled(bus->self, NULL, REQ_ACS_FLAGS)) + if (pci_acs_path_enabled(bus->self, NULL, REQ_ACS_FLAGS, + PCI_ACS_SCOPE_ALL)) break; =20 pdev =3D bus->self; diff --git a/drivers/pci/ats.c b/drivers/pci/ats.c index 96efa00d9743..35c3949f39e0 100644 --- a/drivers/pci/ats.c +++ b/drivers/pci/ats.c @@ -463,7 +463,16 @@ int pci_enable_pasid(struct pci_dev *pdev, int feature= s) if (!pasid) return -EINVAL; =20 - if (!pci_acs_path_enabled(pdev, NULL, PCI_ACS_RR | PCI_ACS_UF)) + /* + * A Request carrying a PASID is routed by address alone (PCIe r7.0, + * sec 2.2.10.4), so it has to be redirected Upstream to reach the + * translation agent. Only Untranslated Requests are at stake here: + * a Translated Request already carries an address the agent produced + * for this PASID (sec 10.1.3), so ACS Direct Translated P2P routing it + * to a peer is not a way around the agent. + */ + if (!pci_acs_path_enabled(pdev, NULL, PCI_ACS_RR | PCI_ACS_UF, + PCI_ACS_SCOPE_UNTRANSLATED)) return -EINVAL; =20 pci_read_config_word(pdev, pasid + PCI_PASID_CAP, &supported); diff --git a/drivers/pci/pci.c b/drivers/pci/pci.c index 77b17b13ee61..492bb26a99de 100644 --- a/drivers/pci/pci.c +++ b/drivers/pci/pci.c @@ -3545,7 +3545,8 @@ void pci_configure_ari(struct pci_dev *dev) } } =20 -static bool pci_acs_flags_enabled(struct pci_dev *pdev, u16 acs_flags) +static bool pci_acs_flags_enabled(struct pci_dev *pdev, u16 acs_flags, + enum pci_acs_scope scope) { int pos; u16 ctrl; @@ -3554,6 +3555,11 @@ static bool pci_acs_flags_enabled(struct pci_dev *pd= ev, u16 acs_flags) if (!pos) return false; =20 + pci_read_config_word(pdev, pos + PCI_ACS_CTRL, &ctrl); + + if (pci_acs_rr_ineffective(ctrl, acs_flags, scope)) + return false; + /* * Except for egress control, capabilities are either required * or only required if controllable. Features missing from the @@ -3561,7 +3567,6 @@ static bool pci_acs_flags_enabled(struct pci_dev *pde= v, u16 acs_flags) */ acs_flags &=3D (pdev->acs_capabilities | PCI_ACS_EC); =20 - pci_read_config_word(pdev, pos + PCI_ACS_CTRL, &ctrl); return (ctrl & acs_flags) =3D=3D acs_flags; } =20 @@ -3569,6 +3574,7 @@ static bool pci_acs_flags_enabled(struct pci_dev *pde= v, u16 acs_flags) * pci_acs_enabled - test ACS against required flags for a given device * @pdev: device to test * @acs_flags: required PCI ACS flags + * @scope: which peer-to-peer Requests the answer has to cover * * Return true if the device supports the provided flags. Automatically * filters out flags that are not implemented on multifunction devices. @@ -3581,11 +3587,12 @@ static bool pci_acs_flags_enabled(struct pci_dev *p= dev, u16 acs_flags) * it much easier for callers of this function to ignore the actual type * or topology of the device when testing ACS support. */ -bool pci_acs_enabled(struct pci_dev *pdev, u16 acs_flags) +bool pci_acs_enabled(struct pci_dev *pdev, u16 acs_flags, + enum pci_acs_scope scope) { int ret; =20 - ret =3D pci_dev_specific_acs_enabled(pdev, acs_flags); + ret =3D pci_dev_specific_acs_enabled(pdev, acs_flags, scope); if (ret >=3D 0) return ret > 0; =20 @@ -3620,7 +3627,7 @@ bool pci_acs_enabled(struct pci_dev *pdev, u16 acs_fl= ags) */ case PCI_EXP_TYPE_DOWNSTREAM: case PCI_EXP_TYPE_ROOT_PORT: - return pci_acs_flags_enabled(pdev, acs_flags); + return pci_acs_flags_enabled(pdev, acs_flags, scope); /* * PCIe 3.0, 6.12.1.2 specifies ACS capabilities that should be * implemented by the remaining PCIe types to indicate peer-to-peer @@ -3635,7 +3642,7 @@ bool pci_acs_enabled(struct pci_dev *pdev, u16 acs_fl= ags) if (!pdev->multifunction) break; =20 - return pci_acs_flags_enabled(pdev, acs_flags); + return pci_acs_flags_enabled(pdev, acs_flags, scope); } =20 /* @@ -3650,19 +3657,20 @@ bool pci_acs_enabled(struct pci_dev *pdev, u16 acs_= flags) * @start: starting downstream device * @end: ending upstream device or NULL to search to the root bus * @acs_flags: required flags + * @scope: which peer-to-peer Requests the answer has to cover * * Walk up a device tree from start to end testing PCI ACS support. If * any step along the way does not support the required flags, return fals= e. */ -bool pci_acs_path_enabled(struct pci_dev *start, - struct pci_dev *end, u16 acs_flags) +bool pci_acs_path_enabled(struct pci_dev *start, struct pci_dev *end, + u16 acs_flags, enum pci_acs_scope scope) { struct pci_dev *pdev, *parent =3D start; =20 do { pdev =3D parent; =20 - if (!pci_acs_enabled(pdev, acs_flags)) + if (!pci_acs_enabled(pdev, acs_flags, scope)) return false; =20 if (pci_is_root_bus(pdev->bus)) diff --git a/drivers/pci/pci.h b/drivers/pci/pci.h index 4469e1a77f3c..6230adb39166 100644 --- a/drivers/pci/pci.h +++ b/drivers/pci/pci.h @@ -1045,15 +1045,41 @@ resource_size_t pci_min_window_alignment(struct pci= _bus *bus, =20 void pci_acs_init(struct pci_dev *dev); void pci_enable_acs(struct pci_dev *dev); + +/* + * PCIe r7.0, sec 6.12.3: ACS P2P Request Redirect does not by itself keep= a + * peer-to-peer Request off the direct path to its target. + * + * Direct Translated P2P routes a Request carrying a Translated address to= the + * peer regardless of Request Redirect, so Request Redirect does not isola= te + * unless Translation Blocking rejects the Request first (sec 6.12.1.1). = It + * says nothing about an Untranslated Request, so a caller asking only abo= ut + * those is unaffected. + * + * @ctrl is the ACS Control register, @acs_flags the controls the caller a= sked + * for, and @scope the Requests its answer has to cover. + */ +static inline bool pci_acs_rr_ineffective(u32 ctrl, u16 acs_flags, + enum pci_acs_scope scope) +{ + if (!(acs_flags & PCI_ACS_RR)) + return false; + + return scope =3D=3D PCI_ACS_SCOPE_ALL && + (ctrl & PCI_ACS_DT) && !(ctrl & PCI_ACS_TB); +} + #ifdef CONFIG_PCI_QUIRKS -int pci_dev_specific_acs_enabled(struct pci_dev *dev, u16 acs_flags); +int pci_dev_specific_acs_enabled(struct pci_dev *dev, u16 acs_flags, + enum pci_acs_scope scope); int pci_dev_specific_enable_acs(struct pci_dev *dev); int pci_dev_specific_disable_acs_redir(struct pci_dev *dev); void pci_disable_broken_acs_cap(struct pci_dev *pdev); int pcie_failed_link_retrain(struct pci_dev *dev); #else static inline int pci_dev_specific_acs_enabled(struct pci_dev *dev, - u16 acs_flags) + u16 acs_flags, + enum pci_acs_scope scope) { return -ENOTTY; } diff --git a/drivers/pci/quirks.c b/drivers/pci/quirks.c index b09f27f7846f..8b50cd0e5114 100644 --- a/drivers/pci/quirks.c +++ b/drivers/pci/quirks.c @@ -4707,7 +4707,8 @@ static int pci_acs_ctrl_enabled(u16 acs_ctrl_req, u16= acs_ctrl_ena) * 1022:780f [AMD] FCH PCI Bridge * 1022:7809 [AMD] FCH USB OHCI Controller */ -static int pci_quirk_amd_sb_acs(struct pci_dev *dev, u16 acs_flags) +static int pci_quirk_amd_sb_acs(struct pci_dev *dev, u16 acs_flags, + enum pci_acs_scope scope) { #ifdef CONFIG_ACPI struct acpi_table_header *header =3D NULL; @@ -4752,7 +4753,8 @@ static bool pci_quirk_cavium_acs_match(struct pci_dev= *dev) } } =20 -static int pci_quirk_cavium_acs(struct pci_dev *dev, u16 acs_flags) +static int pci_quirk_cavium_acs(struct pci_dev *dev, u16 acs_flags, + enum pci_acs_scope scope) { if (!pci_quirk_cavium_acs_match(dev)) return -ENOTTY; @@ -4769,7 +4771,8 @@ static int pci_quirk_cavium_acs(struct pci_dev *dev, = u16 acs_flags) PCI_ACS_SV | PCI_ACS_RR | PCI_ACS_CR | PCI_ACS_UF); } =20 -static int pci_quirk_xgene_acs(struct pci_dev *dev, u16 acs_flags) +static int pci_quirk_xgene_acs(struct pci_dev *dev, u16 acs_flags, + enum pci_acs_scope scope) { /* * X-Gene Root Ports matching this quirk do not allow peer-to-peer @@ -4785,7 +4788,8 @@ static int pci_quirk_xgene_acs(struct pci_dev *dev, u= 16 acs_flags) * But the implementation could block peer-to-peer transactions between th= em * and provide ACS-like functionality. */ -static int pci_quirk_zhaoxin_pcie_ports_acs(struct pci_dev *dev, u16 acs_f= lags) +static int pci_quirk_zhaoxin_pcie_ports_acs(struct pci_dev *dev, u16 acs_f= lags, + enum pci_acs_scope scope) { if (!pci_is_pcie(dev) || ((pci_pcie_type(dev) !=3D PCI_EXP_TYPE_ROOT_PORT) && @@ -4856,7 +4860,8 @@ static bool pci_quirk_intel_pch_acs_match(struct pci_= dev *dev) return false; } =20 -static int pci_quirk_intel_pch_acs(struct pci_dev *dev, u16 acs_flags) +static int pci_quirk_intel_pch_acs(struct pci_dev *dev, u16 acs_flags, + enum pci_acs_scope scope) { if (!pci_quirk_intel_pch_acs_match(dev)) return -ENOTTY; @@ -4878,7 +4883,8 @@ static int pci_quirk_intel_pch_acs(struct pci_dev *de= v, u16 acs_flags) * Port to pass traffic to another Root Port. All PCIe transactions are * terminated inside the Root Port. */ -static int pci_quirk_qcom_rp_acs(struct pci_dev *dev, u16 acs_flags) +static int pci_quirk_qcom_rp_acs(struct pci_dev *dev, u16 acs_flags, + enum pci_acs_scope scope) { return pci_acs_ctrl_enabled(acs_flags, PCI_ACS_SV | PCI_ACS_RR | PCI_ACS_CR | PCI_ACS_UF); @@ -4890,13 +4896,15 @@ static int pci_quirk_qcom_rp_acs(struct pci_dev *de= v, u16 acs_flags) * and validate bus numbers in requests, but does not provide an ACS * capability. */ -static int pci_quirk_nxp_rp_acs(struct pci_dev *dev, u16 acs_flags) +static int pci_quirk_nxp_rp_acs(struct pci_dev *dev, u16 acs_flags, + enum pci_acs_scope scope) { return pci_acs_ctrl_enabled(acs_flags, PCI_ACS_SV | PCI_ACS_RR | PCI_ACS_CR | PCI_ACS_UF); } =20 -static int pci_quirk_al_acs(struct pci_dev *dev, u16 acs_flags) +static int pci_quirk_al_acs(struct pci_dev *dev, u16 acs_flags, + enum pci_acs_scope scope) { if (pci_pcie_type(dev) !=3D PCI_EXP_TYPE_ROOT_PORT) return -ENOTTY; @@ -4976,7 +4984,8 @@ static bool pci_quirk_intel_spt_pch_acs_match(struct = pci_dev *dev) =20 #define INTEL_SPT_ACS_CTRL (PCI_ACS_CAP + 4) =20 -static int pci_quirk_intel_spt_pch_acs(struct pci_dev *dev, u16 acs_flags) +static int pci_quirk_intel_spt_pch_acs(struct pci_dev *dev, u16 acs_flags, + enum pci_acs_scope scope) { int pos; u32 cap, ctrl; @@ -4990,14 +4999,18 @@ static int pci_quirk_intel_spt_pch_acs(struct pci_d= ev *dev, u16 acs_flags) =20 /* see pci_acs_flags_enabled() */ pci_read_config_dword(dev, pos + PCI_ACS_CAP, &cap); - acs_flags &=3D (cap | PCI_ACS_EC); - pci_read_config_dword(dev, pos + INTEL_SPT_ACS_CTRL, &ctrl); =20 + if (pci_acs_rr_ineffective(ctrl, acs_flags, scope)) + return 0; + + acs_flags &=3D (cap | PCI_ACS_EC); + return pci_acs_ctrl_enabled(acs_flags, ctrl); } =20 -static int pci_quirk_mf_endpoint_acs(struct pci_dev *dev, u16 acs_flags) +static int pci_quirk_mf_endpoint_acs(struct pci_dev *dev, u16 acs_flags, + enum pci_acs_scope scope) { /* * SV, TB, and UF are not relevant to multifunction endpoints. @@ -5013,7 +5026,8 @@ static int pci_quirk_mf_endpoint_acs(struct pci_dev *= dev, u16 acs_flags) PCI_ACS_CR | PCI_ACS_UF | PCI_ACS_DT); } =20 -static int pci_quirk_rciep_acs(struct pci_dev *dev, u16 acs_flags) +static int pci_quirk_rciep_acs(struct pci_dev *dev, u16 acs_flags, + enum pci_acs_scope scope) { /* * Intel RCiEP's are required to allow p2p only on translated @@ -5027,7 +5041,8 @@ static int pci_quirk_rciep_acs(struct pci_dev *dev, u= 16 acs_flags) PCI_ACS_SV | PCI_ACS_RR | PCI_ACS_CR | PCI_ACS_UF); } =20 -static int pci_quirk_brcm_acs(struct pci_dev *dev, u16 acs_flags) +static int pci_quirk_brcm_acs(struct pci_dev *dev, u16 acs_flags, + enum pci_acs_scope scope) { /* * iProc PAXB Root Ports don't advertise an ACS capability, but @@ -5039,7 +5054,8 @@ static int pci_quirk_brcm_acs(struct pci_dev *dev, u1= 6 acs_flags) PCI_ACS_SV | PCI_ACS_RR | PCI_ACS_CR | PCI_ACS_UF); } =20 -static int pci_quirk_loongson_acs(struct pci_dev *dev, u16 acs_flags) +static int pci_quirk_loongson_acs(struct pci_dev *dev, u16 acs_flags, + enum pci_acs_scope scope) { /* * Loongson PCIe Root Ports don't advertise an ACS capability, but @@ -5060,7 +5076,8 @@ static int pci_quirk_loongson_acs(struct pci_dev *dev= , u16 acs_flags) * RP1000/RP2000 10G NICs(sp). * FF5xxx 40G/25G/10G NICs(aml). */ -static int pci_quirk_wangxun_nic_acs(struct pci_dev *dev, u16 acs_flags) +static int pci_quirk_wangxun_nic_acs(struct pci_dev *dev, u16 acs_flags, + enum pci_acs_scope scope) { switch (dev->device) { case 0x0100 ... 0x010F: /* EM */ @@ -5077,7 +5094,8 @@ static int pci_quirk_wangxun_nic_acs(struct pci_dev = *dev, u16 acs_flags) static const struct pci_dev_acs_enabled { u16 vendor; u16 device; - int (*acs_enabled)(struct pci_dev *dev, u16 acs_flags); + int (*acs_enabled)(struct pci_dev *dev, u16 acs_flags, + enum pci_acs_scope scope); } pci_dev_acs_enabled[] =3D { { PCI_VENDOR_ID_ATI, 0x4385, pci_quirk_amd_sb_acs }, { PCI_VENDOR_ID_ATI, 0x439c, pci_quirk_amd_sb_acs }, @@ -5256,7 +5274,8 @@ static const struct pci_dev_acs_enabled { * 0: Device does not provide all the desired controls * >0: Device provides all the controls in @acs_flags */ -int pci_dev_specific_acs_enabled(struct pci_dev *dev, u16 acs_flags) +int pci_dev_specific_acs_enabled(struct pci_dev *dev, u16 acs_flags, + enum pci_acs_scope scope) { const struct pci_dev_acs_enabled *i; int ret; @@ -5272,7 +5291,7 @@ int pci_dev_specific_acs_enabled(struct pci_dev *dev,= u16 acs_flags) i->vendor =3D=3D (u16)PCI_ANY_ID) && (i->device =3D=3D dev->device || i->device =3D=3D (u16)PCI_ANY_ID)) { - ret =3D i->acs_enabled(dev, acs_flags); + ret =3D i->acs_enabled(dev, acs_flags, scope); if (ret >=3D 0) return ret; } diff --git a/include/linux/pci.h b/include/linux/pci.h index 64b308b6e61c..867c0f0970bd 100644 --- a/include/linux/pci.h +++ b/include/linux/pci.h @@ -277,6 +277,26 @@ enum pci_bus_flags { PCI_BUS_FLAGS_NO_EXTCFG =3D (__force pci_bus_flags_t) 8, }; =20 +/** + * enum pci_acs_scope - which peer-to-peer Requests an ACS check must cover + * @PCI_ACS_SCOPE_ALL: every peer-to-peer Request, including one carrying a + * Translated address. ACS Direct Translated P2P routes those to the peer + * regardless of P2P Request Redirect (PCIe r7.0, sec 6.12.3), so it + * defeats isolation unless ACS Translation Blocking rejects them first + * (sec 6.12.1.1). + * @PCI_ACS_SCOPE_UNTRANSLATED: only Requests carrying an Untranslated add= ress. + * ACS Direct Translated P2P does not apply to those, so it says nothing + * about whether they reach the Root Complex. + * + * A caller proving that peers cannot reach each other wants + * %PCI_ACS_SCOPE_ALL. A caller that only needs Untranslated Requests rou= ted + * Upstream, such as pci_enable_pasid(), wants %PCI_ACS_SCOPE_UNTRANSLATED. + */ +enum pci_acs_scope { + PCI_ACS_SCOPE_ALL, + PCI_ACS_SCOPE_UNTRANSLATED, +}; + /* Values from Link Status register, PCIe r3.1, sec 7.8.8 */ enum pcie_link_width { PCIE_LNK_WIDTH_RESRV =3D 0x00, @@ -2213,7 +2233,8 @@ static inline struct pci_dev *pci_dev_get(struct pci_= dev *dev) { return NULL; } =20 #define dev_is_pci(d) (false) #define dev_is_pf(d) (false) -static inline bool pci_acs_enabled(struct pci_dev *pdev, u16 acs_flags) +static inline bool pci_acs_enabled(struct pci_dev *pdev, u16 acs_flags, + enum pci_acs_scope scope) { return false; } static inline int pci_irqd_intx_xlate(struct irq_domain *d, struct device_node *node, @@ -2710,9 +2731,10 @@ static inline bool pci_dev_is_disconnected(const str= uct pci_dev *dev) } =20 void pci_request_acs(void); -bool pci_acs_enabled(struct pci_dev *pdev, u16 acs_flags); -bool pci_acs_path_enabled(struct pci_dev *start, - struct pci_dev *end, u16 acs_flags); +bool pci_acs_enabled(struct pci_dev *pdev, u16 acs_flags, + enum pci_acs_scope scope); +bool pci_acs_path_enabled(struct pci_dev *start, struct pci_dev *end, + u16 acs_flags, enum pci_acs_scope scope); int pci_enable_atomic_ops_to_root(struct pci_dev *dev, u32 cap_mask); =20 #define PCI_VPD_LRDT 0x80 /* Large Resource Data Type */ --=20 2.55.0 From nobody Tue Sep 29 06:58:50 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 47F0A43441C; Tue, 11 Aug 2026 09:31:53 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440716; cv=none; b=RwfV2M9dEpH9n4AuW0Ktgsrmr8sHoJgl66TYWqWaChayW8mv06H5P2oeV+kVQ9JLlSA1gYBeqQ2ZvyAgEWxRIf1IcfT1Hi3Qnfe5si5XCfiUVo8Qvxkzn9lcmI4kWG5OT/TLGGpM4N/u2xIM0T89SVeEpaGO0HMdG7D/FRPd71A= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440716; c=relaxed/simple; bh=lV4jEK4svh+6UshbsA84C+hjExtusy0e+WQBHxPePYE=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=r9IL7VjbE8nxt8LfO6dXdGAHve30M5IAVRULAFXWI1QnTHbsvOrDQtD9zvkLAa21JYoFRaKfJ3OiQnFY5VnhXHLfzMnLl0ouZyb+nbftWOk/bxFMKdG9WVv1Yf1JsPPQuG710QUatIIHHsJq4kFSRQ7NjF7LFWpJyct7p5plp7w= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=ULF6cKOt; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="ULF6cKOt" Received: by smtp.kernel.org (Postfix) with ESMTPSA id B8CEC1F000E9; Tue, 11 Aug 2026 09:31:51 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786440712; bh=048mN//99k0AQtFFFqqA1UJMkxuCpS/5Uf5l8gCWCE4=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=ULF6cKOty2zW/8CWeek7MWqZ/kw0rxjSkjR/t6Y2BiN25Zx7Th62pZB9QPugcw0zd 1rHVU81ITwu2VIfURWR9g8LGbpBrQaK0UUrxGMzf6DmpkTxmQpKAcrKRl30uDoai1f NcUlPUiyCv87EQOX0RCV3LBjrPz+JnJGhLvCPV0aVu63EvqBw+FZCL4va/O72hgurF 2IkYdhcyyDZGxwSAZqK4k/8G5nkVRZVhrNr42ECD9p6QMCJgHI1gA2B10KFpEfD1d6 N8O8s8BiNN1i1hb8b8yfTdW4CR3IXsKAMgITKxd89tfC0LruopkidkHKe4s1NF12qr u19FRa5Nn8dOw== From: Leon Romanovsky To: Bjorn Helgaas , Logan Gunthorpe , Chaitanya Kulkarni , Greg Kroah-Hartman , Jens Axboe , Alex Williamson , Leon Romanovsky , Ankit Agrawal , Jason Gunthorpe , Jonathan Corbet , Shuah Khan , "Joerg Roedel (AMD)" , Will Deacon , Robin Murphy Cc: linux-pci@vger.kernel.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, iommu@lists.linux.dev Subject: [PATCH v3 09/17] PCI: Add ACS egress control vector accessor Date: Tue, 11 Aug 2026 12:30:51 +0300 Message-ID: <20260811-fix-p2p-acs-v3-9-efc488ee7c03@nvidia.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> References: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" X-Mailer: b4 0.15-dev-18f8f Content-Transfer-Encoding: quoted-printable From: Leon Romanovsky Whether ACS P2P Egress Control routes, redirects, or blocks a peer-to-peer request depends on the Egress Control Vector bit for the target port, not on the enable bit alone (PCIe r7.0, sec 6.12.3). Provide a helper to read that bit for a peer Root or Switch Downstream Port. Report an unreadable or uncovered vector as an error rather than as a clear bit, so callers do not mistake it for permission to route directly. Each bit corresponds to a Port Number within one Switch or Root Complex (sec 7.7.12.4), so both ports have to number their ports in the same place. Downstream Ports of one Switch share its internal bus, and Root Ports of one Root Complex share a root bus, so require a shared bus and reject anything else. A target numbered elsewhere has no bit in this vector and would select an unrelated one. Reviewed-by: Logan Gunthorpe Signed-off-by: Leon Romanovsky --- drivers/pci/pci.c | 68 +++++++++++++++++++++++++++++++++++++++++++++++++++= ++++ drivers/pci/pci.h | 1 + 2 files changed, 69 insertions(+) diff --git a/drivers/pci/pci.c b/drivers/pci/pci.c index 492bb26a99de..8d165c9534ff 100644 --- a/drivers/pci/pci.c +++ b/drivers/pci/pci.c @@ -3545,6 +3545,74 @@ void pci_configure_ari(struct pci_dev *dev) } } =20 +/* + * PCIe r7.0, sec 7.7.12: only for Root Ports and Switch Downstream Ports = does + * each Egress Control Vector bit correspond to a Port Number. Elsewhere = the + * vector is indexed by Function or Function Group Number, so a Link + * Capabilities Port Number must not be used to select a bit. + * + * pcie_downstream_port() is too permissive here because it also accepts a + * PCI/PCI-X to PCIe Bridge. + */ +static bool pci_acs_egress_vector_port(const struct pci_dev *dev) +{ + int type =3D pci_pcie_type(dev); + + return type =3D=3D PCI_EXP_TYPE_ROOT_PORT || + type =3D=3D PCI_EXP_TYPE_DOWNSTREAM; +} + +/** + * pci_acs_egress_ctrl_is_set - Read an ACS Egress Control Vector bit + * @pdev: ingress Root or Switch Downstream Port + * @target: target Root or Switch Downstream Port + * + * Return: 1 if @pdev's Egress Control Vector bit for @target is set, 0 if + * it is clear, or a negative errno if the bit cannot be determined. + */ +int pci_acs_egress_ctrl_is_set(struct pci_dev *pdev, struct pci_dev *targe= t) +{ + unsigned int vector_size; + u32 lnkcap, vector; + u8 target_port; + int ret; + + if (!(pdev->acs_capabilities & PCI_ACS_EC) || + !pci_acs_egress_vector_port(pdev) || + !pci_acs_egress_vector_port(target)) + return -EOPNOTSUPP; + + /* + * Each vector bit corresponds to a Port Number within one Switch or + * Root Complex (PCIe r7.0, sec 7.7.12.4). Downstream Ports of one + * Switch share its internal bus and Root Ports of one Root Complex + * share a root bus, so anything else numbers its ports elsewhere and + * would index an unrelated bit here. + */ + if (pdev->bus !=3D target->bus) + return -EOPNOTSUPP; + + ret =3D pcie_capability_read_dword(target, PCI_EXP_LNKCAP, &lnkcap); + if (ret) + return pcibios_err_to_errno(ret); + + target_port =3D FIELD_GET(PCI_EXP_LNKCAP_PN, lnkcap); + vector_size =3D pdev->acs_capabilities >> 8; + + /* An Egress Control Vector Size of 0 encodes 256 bits. */ + if (vector_size && target_port >=3D vector_size) + return -ERANGE; + + ret =3D pci_read_config_dword(pdev, + pdev->acs_cap + PCI_ACS_EGRESS_CTL_V + + (target_port / 32) * sizeof(vector), + &vector); + if (ret) + return pcibios_err_to_errno(ret); + + return !!(vector & BIT(target_port % 32)); +} + static bool pci_acs_flags_enabled(struct pci_dev *pdev, u16 acs_flags, enum pci_acs_scope scope) { diff --git a/drivers/pci/pci.h b/drivers/pci/pci.h index 6230adb39166..d3ea9b2bb7fc 100644 --- a/drivers/pci/pci.h +++ b/drivers/pci/pci.h @@ -1069,6 +1069,7 @@ static inline bool pci_acs_rr_ineffective(u32 ctrl, u= 16 acs_flags, (ctrl & PCI_ACS_DT) && !(ctrl & PCI_ACS_TB); } =20 +int pci_acs_egress_ctrl_is_set(struct pci_dev *pdev, struct pci_dev *targe= t); #ifdef CONFIG_PCI_QUIRKS int pci_dev_specific_acs_enabled(struct pci_dev *dev, u16 acs_flags, enum pci_acs_scope scope); --=20 2.55.0 From nobody Tue Sep 29 06:58:50 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2256F434E50; Tue, 11 Aug 2026 09:31:55 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440718; cv=none; b=qffg8S/GLNyVYRN75bKjUNQX5XuGV7qthyD3CKVGUM7WvuTgWiDjF4LT3wZ5dUqjG9bY90EFzgCVWy0easy9M/ie7zWMVfUrNQeX9geHu2g1jltqSjdYInUTOy3eH6Wua9nO5U88DrdzaRPC4KVg7jmbQIeFsFzCm1w6DBKQ7pY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440718; c=relaxed/simple; bh=wAGPi+FQZKFYHbcujproushB6aoFKHI2znb3lKf16pA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=Hqk7Ez86wphwalKXuBes0ahQ+8iGPAY3qXewRciWDulHh8s3J6USxPx/OJVFKZIsO50VBgDRcFan8BzqeN+yMhzh0ta6wYKsL79QqEr5UXa4IqJ3eCRcJfd0viE0eb+LLpgaYFZ7+fbpWiqVU4cBypDmAmI/byGwtLSZCyPn6y0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=HYRz8Q4t; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="HYRz8Q4t" Received: by smtp.kernel.org (Postfix) with ESMTPSA id CB2CB1F00A3D; Tue, 11 Aug 2026 09:31:54 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786440715; bh=0uxyPV/6i8QHnH03dgNAW09i5KijziMNkEBo8lvxSrI=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=HYRz8Q4tZ6qpPRETvt37gVxbB/jlB1QHzM8KOHSo4gr2LIsjFWjASBLM8BPRcvHT8 vhxqog7eQYnP3lPpbEujx7Ty9+BMAx1tEpho7eSsRe35tl8A5Q+zVt7ChQACBOBbt1 3+00Re6aI2OA/PvaR2VptQO2bpRcMUzwX+udxh5M2qQazL6LAq6wKIDvpg6S9dZk6a 1uzw+68mon4W0ioGAElHs+CJo+5aR+LywsaGLwbqsokjzPVEp/1RGs+Bz0scp5dtb8 dNJDIrsmJzglHIqV/oFYewJRQHor2IkZogUBGcB18sbFrPzJyY/JUo8fqlL6K9YlSL N/tqkbD/6fcjg== From: Leon Romanovsky To: Bjorn Helgaas , Logan Gunthorpe , Chaitanya Kulkarni , Greg Kroah-Hartman , Jens Axboe , Alex Williamson , Leon Romanovsky , Ankit Agrawal , Jason Gunthorpe , Jonathan Corbet , Shuah Khan , "Joerg Roedel (AMD)" , Will Deacon , Robin Murphy Cc: linux-pci@vger.kernel.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, iommu@lists.linux.dev Subject: [PATCH v3 10/17] PCI: Account for ACS egress control in isolation checks Date: Tue, 11 Aug 2026 12:30:52 +0300 Message-ID: <20260811-fix-p2p-acs-v3-10-efc488ee7c03@nvidia.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> References: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" X-Mailer: b4 0.15-dev-18f8f Content-Transfer-Encoding: quoted-printable From: Leon Romanovsky pci_acs_enabled() treats P2P Request Redirect as effective whenever its control bit is set. PCIe r7.0, sec 6.12.3, table 6-11 lets an enabled Egress Control Vector override it: a clear vector bit routes the request directly. IOMMU grouping uses this check to prove peer requests cannot bypass the IOMMU, but cannot know every applicable vector bit, so Request Redirect gives no such guarantee while Egress Control is enabled. Report Request Redirect as ineffective there, merging the devices into one IOMMU group, and report no isolation when the register cannot be read. Apply the same rule to the Intel SPT PCH quirk. Unlike Direct Translated P2P this holds for an Untranslated Request too, so it applies in both scopes. pci_enable_pasid() therefore fails on a path where a port has Egress Control enabled, because Request Redirect no longer shows that a Request carrying a PASID reaches the translation agent. Fixes: ad805758c0eb ("PCI: add ACS validation utility") Reviewed-by: Logan Gunthorpe Signed-off-by: Leon Romanovsky --- drivers/pci/pci.c | 3 ++- drivers/pci/pci.h | 13 +++++++++++-- drivers/pci/quirks.c | 7 +++++-- 3 files changed, 18 insertions(+), 5 deletions(-) diff --git a/drivers/pci/pci.c b/drivers/pci/pci.c index 8d165c9534ff..a633f473590f 100644 --- a/drivers/pci/pci.c +++ b/drivers/pci/pci.c @@ -3623,7 +3623,8 @@ static bool pci_acs_flags_enabled(struct pci_dev *pde= v, u16 acs_flags, if (!pos) return false; =20 - pci_read_config_word(pdev, pos + PCI_ACS_CTRL, &ctrl); + if (pci_read_config_word(pdev, pos + PCI_ACS_CTRL, &ctrl)) + return false; =20 if (pci_acs_rr_ineffective(ctrl, acs_flags, scope)) return false; diff --git a/drivers/pci/pci.h b/drivers/pci/pci.h index d3ea9b2bb7fc..32394e349766 100644 --- a/drivers/pci/pci.h +++ b/drivers/pci/pci.h @@ -1056,6 +1056,12 @@ void pci_enable_acs(struct pci_dev *dev); * says nothing about an Untranslated Request, so a caller asking only abo= ut * those is unaffected. * + * Egress Control can override Request Redirect for any peer Request, + * Untranslated ones included, so it applies in either scope. This + * target-independent test cannot prove that every applicable Egress Contr= ol + * Vector bit is set, so Request Redirect does not guarantee that the Requ= est + * leaves the direct path while Egress Control is enabled. + * * @ctrl is the ACS Control register, @acs_flags the controls the caller a= sked * for, and @scope the Requests its answer has to cover. */ @@ -1065,8 +1071,11 @@ static inline bool pci_acs_rr_ineffective(u32 ctrl, = u16 acs_flags, if (!(acs_flags & PCI_ACS_RR)) return false; =20 - return scope =3D=3D PCI_ACS_SCOPE_ALL && - (ctrl & PCI_ACS_DT) && !(ctrl & PCI_ACS_TB); + if (scope =3D=3D PCI_ACS_SCOPE_ALL && + (ctrl & PCI_ACS_DT) && !(ctrl & PCI_ACS_TB)) + return true; + + return ctrl & PCI_ACS_EC; } =20 int pci_acs_egress_ctrl_is_set(struct pci_dev *pdev, struct pci_dev *targe= t); diff --git a/drivers/pci/quirks.c b/drivers/pci/quirks.c index 8b50cd0e5114..cee6be63cadd 100644 --- a/drivers/pci/quirks.c +++ b/drivers/pci/quirks.c @@ -4998,8 +4998,11 @@ static int pci_quirk_intel_spt_pch_acs(struct pci_de= v *dev, u16 acs_flags, return -ENOTTY; =20 /* see pci_acs_flags_enabled() */ - pci_read_config_dword(dev, pos + PCI_ACS_CAP, &cap); - pci_read_config_dword(dev, pos + INTEL_SPT_ACS_CTRL, &ctrl); + if (pci_read_config_dword(dev, pos + PCI_ACS_CAP, &cap)) + return 0; + + if (pci_read_config_dword(dev, pos + INTEL_SPT_ACS_CTRL, &ctrl)) + return 0; =20 if (pci_acs_rr_ineffective(ctrl, acs_flags, scope)) return 0; --=20 2.55.0 From nobody Tue Sep 29 06:58:50 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A091142EEA4; Tue, 11 Aug 2026 09:32:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440741; cv=none; b=rSNDKbWy/r+zEOaA5BZkRuqV4EDMFl5LKLPeiisq11AmEluMTxiUB6TwlsoPboQjwvPDCE7Oq1bg3HC3sbYWX5q7nqEZI/c2SFVpeSAngYZVZ77SO/s2c8h7K3RODs7C8rSv6CBeUiCst8Q7/O/JroCrbew1s4aGXuI+MM68xcg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440741; c=relaxed/simple; bh=Pt1nevMTgJ0hyu2UqYgqPOmzI5fgDrwSDPKbuwPAmCw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=tfy4VyIy1Zg9CoegB2AcreQdaKv93LtMBpA3gD+EVOknE/1o2Gfgs6oQ8CA0ElVds57MytuswdUAD+QWihSUzBRpgFMCS/ypdICRdMPSA0y/DpzbUR9ejMcg1mPAFmWlgOosH/U6JF6nZoSyZOw+7Fl6vW9eq8D3SktnRhSqu5k= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=hlu1PxGu; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="hlu1PxGu" Received: by smtp.kernel.org (Postfix) with ESMTPSA id F29A51F000E9; Tue, 11 Aug 2026 09:32:19 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786440740; bh=ffLJzyWnTCRIbRrdEkbpXBdeTZ14TQNfRzt8dE7+yJg=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=hlu1PxGu9/0mtIhIrSV+P7xR0aFEX7dwbF6vhlhkIB1O8Ir/xgO7B1fAitns/Jnna shEhZE/fhKwBdk2m7dMDrGT4Z9eE0rBA9CQjNmPIwVZ584tYt9FqcfrOdFkntKME0c x7aSIduzqio2g76hiJAFrpyiqgJLTB4WL00IliYlOO0S3prZPjY4Fhvnt+53BDsdNN QswvCXJZjssnMG7tOAt49/BufIk4djnqLYV8LkgyfbJTBATv50F07XnMtZDadHKTsq exHydF58F5IbFnbPhQNgVz20k4+jed5On+QAszjrzpe2v6iuRW+/YabYHWFzFUlaEx w9wimDwGrio5Q== From: Leon Romanovsky To: Bjorn Helgaas , Logan Gunthorpe , Chaitanya Kulkarni , Greg Kroah-Hartman , Jens Axboe , Alex Williamson , Leon Romanovsky , Ankit Agrawal , Jason Gunthorpe , Jonathan Corbet , Shuah Khan , "Joerg Roedel (AMD)" , Will Deacon , Robin Murphy Cc: linux-pci@vger.kernel.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, iommu@lists.linux.dev Subject: [PATCH v3 11/17] PCI/P2PDMA: Derive peer-to-peer routing from ACS control bits Date: Tue, 11 Aug 2026 12:30:53 +0300 Message-ID: <20260811-fix-p2p-acs-v3-11-efc488ee7c03@nvidia.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> References: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" X-Mailer: b4 0.15-dev-18f8f Content-Transfer-Encoding: quoted-printable From: Leon Romanovsky The redirect test reports an upstream redirect for any ACS redirect or egress control bit. PCIe r7.0, sec 6.12.3, table 6-11 ties the outcome to the control bits and the target vector, and an enabled egress control bit alone does not redirect. Compute the result from the control bits so the vector can be honored next, using the host-bridge route while the peer target is unknown. Reviewed-by: Logan Gunthorpe Signed-off-by: Leon Romanovsky --- drivers/pci/p2pdma.c | 49 ++++++++++++++++++++++++++++++++++++------------- 1 file changed, 36 insertions(+), 13 deletions(-) diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c index a364008bbf50..69cef8ca9557 100644 --- a/drivers/pci/p2pdma.c +++ b/drivers/pci/p2pdma.c @@ -21,6 +21,8 @@ #include #include =20 +#include "pci.h" + /* * Lifetime and RCU usage * @@ -536,26 +538,45 @@ static struct pci_dev *find_parent_pci_dev(struct dev= ice *dev) return NULL; } =20 -/* - * Check if a PCI bridge has its ACS redirection bits set to redirect P2P - * TLPs upstream via ACS. Returns 1 if the packets will be redirected - * upstream, 0 otherwise. - */ -static int pci_bridge_has_acs_redir(struct pci_dev *pdev) +enum pci_acs_p2pdma_state { + PCI_ACS_P2PDMA_DIRECT, + PCI_ACS_P2PDMA_REDIRECT, +}; + +static enum pci_acs_p2pdma_state +pci_acs_p2pdma_state(struct pci_dev *pdev, struct pci_dev *target) { int pos; u16 ctrl; =20 pos =3D pdev->acs_cap; if (!pos) - return 0; + return PCI_ACS_P2PDMA_DIRECT; =20 - pci_read_config_word(pdev, pos + PCI_ACS_CTRL, &ctrl); + if (pci_read_config_word(pdev, pos + PCI_ACS_CTRL, &ctrl)) + return PCI_ACS_P2PDMA_REDIRECT; =20 - if (ctrl & (PCI_ACS_RR | PCI_ACS_CR | PCI_ACS_EC)) - return 1; + if (!(ctrl & PCI_ACS_EC)) + return ctrl & (PCI_ACS_RR | PCI_ACS_CR) ? + PCI_ACS_P2PDMA_REDIRECT : PCI_ACS_P2PDMA_DIRECT; =20 - return 0; + /* + * The vector cannot be read without the peer target, so redirect + * upstream until the paths diverge. + */ + if (!target) + return PCI_ACS_P2PDMA_REDIRECT; + + /* + * PCIe r7.0, sec 6.12.3, table 6-11: a set or indeterminate egress + * control vector bit keeps the request off the direct path; a clear + * bit permits it, subject only to completion redirect. + */ + if (pci_acs_egress_ctrl_is_set(pdev, target)) + return PCI_ACS_P2PDMA_REDIRECT; + + return ctrl & PCI_ACS_CR ? PCI_ACS_P2PDMA_REDIRECT : + PCI_ACS_P2PDMA_DIRECT; } =20 static void seq_buf_print_bus_devfn(struct seq_buf *buf, struct pci_dev *p= dev) @@ -767,7 +788,8 @@ calc_map_type_and_dist(struct pci_dev *provider, struct= pci_dev *client, while (a) { dist_b =3D 0; =20 - if (pci_bridge_has_acs_redir(a)) { + if (pci_acs_p2pdma_state(a, NULL) =3D=3D + PCI_ACS_P2PDMA_REDIRECT) { seq_buf_print_bus_devfn(&acs_list, a); acs_cnt++; } @@ -796,7 +818,8 @@ calc_map_type_and_dist(struct pci_dev *provider, struct= pci_dev *client, if (a =3D=3D bb) break; =20 - if (pci_bridge_has_acs_redir(bb)) { + if (pci_acs_p2pdma_state(bb, NULL) =3D=3D + PCI_ACS_P2PDMA_REDIRECT) { seq_buf_print_bus_devfn(&acs_list, bb); acs_cnt++; } --=20 2.55.0 From nobody Tue Sep 29 06:58:50 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A6E3D4399E3; Tue, 11 Aug 2026 09:32:01 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440723; cv=none; b=DgJ50gl0BuOAfh6HYYEhrg024dyhbVuFPUlanwvtvDhEWIGLvOp6sERib7ZbnryKdPs1WNrWGgEb6Qk/6yWfY3GO2TrO/bkPf7kQ0Zj1DpbSOlPW9dpa9bYwIi9v0LELEvNdWyK8+4Zzn6nhB8DwPY32J1rPNBCuhWsV21+4Qqc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440723; c=relaxed/simple; bh=sDeEMTnZhYHGCJpPaFb0hBaDV1Kogv9eqZ8Om7GiYrM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=ZjQj8flbxf4tfbBfi2TE0uyIaS9sC5ZH86reSQCjGXS48ZL0vRe63yfXtqkCuhVoF/drD3i+m7mylKdij9LAUlDNxMbvvpRy6Be6q4idIUDVIuo5Tg71pzYicudHdqTMAds4zaRET2ZuwSgj9j53gkWPeMK37LJe0fRzKijh4y8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=byi4owM9; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="byi4owM9" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 0AE191F000E9; Tue, 11 Aug 2026 09:32:00 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786440721; bh=D2EpG6HPdiSnlBbrmNMridF7es1dC7ICUwry8aozP2I=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=byi4owM9wyN1ThEoTwczNy5HmqXDW8QjvsruNitcmmzYzt97GxX5l3BA1mVc7hd9w Mkv0QcOtrXLK0LZBgUjG5Sf0ZMur1EwN/63BCNxZZwXSbblGLy2Ud5mRPeiAhOQzd3 OQCL7zP3D2Y3CglPfeH6+3qYZ/gmJW1jElXYUWXDWsgIr/1+Lc3xlReQlWUMDoL3Fk 6aYX0mPbYDNm5SdsgrkFu9+fJDmsCLPIAtTdXR0oz3IQZnjQINmZ/AqCTWy9DxmuQA SkgA7p9mlLJZD4ae0TPMxvxW4yHFQAf9lKT0favd55XGp/L9G3ZM/mTni4/j64iChv xiUYUZ3exoVsg== From: Leon Romanovsky To: Bjorn Helgaas , Logan Gunthorpe , Chaitanya Kulkarni , Greg Kroah-Hartman , Jens Axboe , Alex Williamson , Leon Romanovsky , Ankit Agrawal , Jason Gunthorpe , Jonathan Corbet , Shuah Khan , "Joerg Roedel (AMD)" , Will Deacon , Robin Murphy Cc: linux-pci@vger.kernel.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, iommu@lists.linux.dev Subject: [PATCH v3 12/17] PCI/P2PDMA: Honor ACS egress control vectors Date: Tue, 11 Aug 2026 12:30:54 +0300 Message-ID: <20260811-fix-p2p-acs-v3-12-efc488ee7c03@nvidia.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> References: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" X-Mailer: b4 0.15-dev-18f8f Content-Transfer-Encoding: quoted-printable From: Leon Romanovsky An enabled Egress Control bit does not itself send a peer request upstream. PCIe r7.0, sec 6.12.3, table 6-11 makes the outcome depend on the Egress Control Vector bit for the target port: a clear bit routes the request directly regardless of P2P Request Redirect. Read the vector where the paths diverge below their common upstream port. Keep a clear vector bit on the direct path, subject to P2P Completion Redirect. A set bit with Request Redirect clear is an ACS Violation. ACS acts only on peer-to-peer Requests, so route it, and an indeterminate vector, through the host bridge. Fixes: 52916982af48 ("PCI/P2PDMA: Support peer-to-peer memory") Reviewed-by: Logan Gunthorpe Signed-off-by: Leon Romanovsky --- drivers/pci/p2pdma.c | 114 +++++++++++++++++++++++++++++++++--------------= ---- 1 file changed, 75 insertions(+), 39 deletions(-) diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c index 69cef8ca9557..879c92d66f5b 100644 --- a/drivers/pci/p2pdma.c +++ b/drivers/pci/p2pdma.c @@ -541,12 +541,13 @@ static struct pci_dev *find_parent_pci_dev(struct dev= ice *dev) enum pci_acs_p2pdma_state { PCI_ACS_P2PDMA_DIRECT, PCI_ACS_P2PDMA_REDIRECT, + PCI_ACS_P2PDMA_NOT_SUPPORTED, }; =20 static enum pci_acs_p2pdma_state pci_acs_p2pdma_state(struct pci_dev *pdev, struct pci_dev *target) { - int pos; + int pos, ret; u16 ctrl; =20 pos =3D pdev->acs_cap; @@ -554,26 +555,26 @@ pci_acs_p2pdma_state(struct pci_dev *pdev, struct pci= _dev *target) return PCI_ACS_P2PDMA_DIRECT; =20 if (pci_read_config_word(pdev, pos + PCI_ACS_CTRL, &ctrl)) - return PCI_ACS_P2PDMA_REDIRECT; + return PCI_ACS_P2PDMA_NOT_SUPPORTED; =20 - if (!(ctrl & PCI_ACS_EC)) + /* EC applies only at the path divergence where the target is known. */ + if (!target || !(ctrl & PCI_ACS_EC)) return ctrl & (PCI_ACS_RR | PCI_ACS_CR) ? PCI_ACS_P2PDMA_REDIRECT : PCI_ACS_P2PDMA_DIRECT; =20 /* - * The vector cannot be read without the peer target, so redirect - * upstream until the paths diverge. + * PCIe r7.0, sec 6.12.3, table 6-11: a set Egress Control Vector + * bit redirects the request only when Request Redirect is set. With + * Request Redirect clear, the request is handled as an ACS Violation. + * A clear vector bit permits direct routing, subject to Completion + * Redirect. */ - if (!target) - return PCI_ACS_P2PDMA_REDIRECT; - - /* - * PCIe r7.0, sec 6.12.3, table 6-11: a set or indeterminate egress - * control vector bit keeps the request off the direct path; a clear - * bit permits it, subject only to completion redirect. - */ - if (pci_acs_egress_ctrl_is_set(pdev, target)) - return PCI_ACS_P2PDMA_REDIRECT; + ret =3D pci_acs_egress_ctrl_is_set(pdev, target); + if (ret < 0) + return PCI_ACS_P2PDMA_NOT_SUPPORTED; + if (ret) + return ctrl & PCI_ACS_RR ? PCI_ACS_P2PDMA_REDIRECT : + PCI_ACS_P2PDMA_NOT_SUPPORTED; =20 return ctrl & PCI_ACS_CR ? PCI_ACS_P2PDMA_REDIRECT : PCI_ACS_P2PDMA_DIRECT; @@ -754,9 +755,9 @@ static unsigned long map_types_idx(struct pci_dev *clie= nt) * then to Device B. The mapping type returned depends on the ACS * redirection setting of the ports along the path. * - * If ACS redirect is set on any port in the path, traffic between the - * devices will go through the host bridge, so return - * PCI_P2PDMA_MAP_THRU_HOST_BRIDGE; otherwise return + * If ACS redirects traffic on any port in the path, or blocks the direct + * path or leaves its routing indeterminate, return + * PCI_P2PDMA_MAP_THRU_HOST_BRIDGE. Otherwise, return * PCI_P2PDMA_MAP_BUS_ADDR. * * Any two devices that have a data path that goes through the host bridge @@ -770,10 +771,13 @@ calc_map_type_and_dist(struct pci_dev *provider, stru= ct pci_dev *client, int *dist, bool verbose) { enum pci_p2pdma_map_type map_type =3D PCI_P2PDMA_MAP_THRU_HOST_BRIDGE; - struct pci_dev *a =3D provider, *b =3D client, *bb; + struct pci_dev *a =3D provider, *b =3D client, *bb, *target; + struct pci_dev *a_child =3D NULL, *b_child =3D NULL; + struct pci_dev *acs_unsupported =3D NULL; + enum pci_acs_p2pdma_state state; struct pci_p2pdma *p2pdma; struct seq_buf acs_list; - int acs_cnt =3D 0; + int acs_redirect_cnt =3D 0; int dist_a =3D 0; int dist_b =3D 0; char buf[128]; @@ -787,60 +791,92 @@ calc_map_type_and_dist(struct pci_dev *provider, stru= ct pci_dev *client, */ while (a) { dist_b =3D 0; - - if (pci_acs_p2pdma_state(a, NULL) =3D=3D - PCI_ACS_P2PDMA_REDIRECT) { - seq_buf_print_bus_devfn(&acs_list, a); - acs_cnt++; - } - + b_child =3D NULL; bb =3D b; =20 while (bb) { if (a =3D=3D bb) - goto check_b_path_acs; + goto check_paths_acs; =20 + b_child =3D bb; bb =3D pci_upstream_bridge(bb); dist_b++; } =20 + a_child =3D a; a =3D pci_upstream_bridge(a); dist_a++; } =20 + /* + * The paths share no upstream bridge, so there is no direct path for + * ACS to gate: PCI_P2PDMA_MAP_BUS_ADDR is not reachable here and the + * request can only get to the peer through the host bridge. + */ *dist =3D dist_a + dist_b; goto map_through_host_bridge; =20 -check_b_path_acs: - bb =3D b; +check_paths_acs: + *dist =3D dist_a + dist_b; + bb =3D provider; =20 while (bb) { + target =3D bb =3D=3D a_child ? b_child : NULL; + state =3D pci_acs_p2pdma_state(bb, target); + if (state !=3D PCI_ACS_P2PDMA_DIRECT) { + seq_buf_print_bus_devfn(&acs_list, bb); + if (state =3D=3D PCI_ACS_P2PDMA_REDIRECT) + acs_redirect_cnt++; + else if (!acs_unsupported) + acs_unsupported =3D bb; + } + if (a =3D=3D bb) break; =20 - if (pci_acs_p2pdma_state(bb, NULL) =3D=3D - PCI_ACS_P2PDMA_REDIRECT) { + bb =3D pci_upstream_bridge(bb); + } + + bb =3D client; + + while (bb && a !=3D bb) { + target =3D bb =3D=3D b_child ? a_child : NULL; + state =3D pci_acs_p2pdma_state(bb, target); + if (state !=3D PCI_ACS_P2PDMA_DIRECT) { seq_buf_print_bus_devfn(&acs_list, bb); - acs_cnt++; + if (state =3D=3D PCI_ACS_P2PDMA_REDIRECT) + acs_redirect_cnt++; + else if (!acs_unsupported) + acs_unsupported =3D bb; } =20 bb =3D pci_upstream_bridge(bb); } =20 - *dist =3D dist_a + dist_b; - - if (!acs_cnt) { + /* + * Below a shared upstream bridge, a path that no port redirects or + * blocks routes the request directly. + */ + if (!acs_unsupported && !acs_redirect_cnt) { map_type =3D PCI_P2PDMA_MAP_BUS_ADDR; goto done; } =20 + /* + * ACS controls only act on Requests routed peer-to-peer, so a blocked + * or indeterminate direct path still leaves the host-bridge route. + */ if (verbose) { /* Drop the final semicolon; the list is not empty here. */ if (!seq_buf_has_overflowed(&acs_list)) acs_list.buffer[acs_list.len - 1] =3D '\0'; - pci_warn(client, "ACS redirect is set between the client and provider (%= s)\n", - pci_name(provider)); - pci_warn(client, "to disable ACS redirect for this path, add the kernel = parameter: pci=3Ddisable_acs_redir=3D%s\n", + if (acs_unsupported) + pci_warn(client, "ACS leaves no usable direct P2P path to provider %s a= t %s\n", + pci_name(provider), pci_name(acs_unsupported)); + else + pci_warn(client, "ACS redirect is set between the client and provider (= %s)\n", + pci_name(provider)); + pci_warn(client, "to disable ACS controls for this path, add the kernel = parameter: pci=3Ddisable_acs_redir=3D%s\n", seq_buf_str(&acs_list)); } =20 --=20 2.55.0 From nobody Tue Sep 29 06:58:50 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2690A43A80C; Tue, 11 Aug 2026 09:32:04 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440726; cv=none; b=ceAcpto2S3/N4mrdOGLRdVRFzBbbIFpU+QD3yuqF5vUnRr3+m2UMJ+YbDAvtdC3n/kwO5Pl2kbAuiHS0qm1IzcU1ENEX+korGMegFSVARl3O2rGp9sGXCQwjSm3PGCu6D8JgQBaff8eUHiuoErNR/qYczd1dRhNO6nUTYR4e7gM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440726; c=relaxed/simple; bh=TLQDvKOAPD4WOMF0kj2k/X8n13r0aMes4uuRLi3koa0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=Wwyq4B13yD+XQOtoB22rKhFG1OGxLEiAc3vdWxt9H0pdgh2mNgcnoPsYRYbX+FdBl59O8p9QyPNJEhC/BnKl3iE14mzIlgPeVzh7WBOz/z4ocOECGtHIQElmOBEIeNWYjeEJlYU1MVuFWteJbOFS32TdrJq/OYyxcVhXXCH3rpk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=Z/OnwWEr; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="Z/OnwWEr" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 2EB741F00A3A; Tue, 11 Aug 2026 09:32:03 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786440724; bh=09VxA+XpyBE+sTpo60zn6eYSOkH91ZUsB7fB5VsT4XQ=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=Z/OnwWErtDUuZVKKcmxlNIq5eiLWannOsjjzs8nu5OotwpyNFKnKH7pg3gncEsL48 qy22p8mlq1ayXEltVZXCY0Oid8jqZPWkBm7M7mDT/9VnkB3/Pou8+Ssa7MvRB13BeT +/BGlJRetFiXxcwex9kAJUHwo5Nf8edJMYKLB8OewTVdy0cKEzLFTg2sujBMw1VWnQ hsMzAcUyu7z+lPmfUuDN0075AagKR94h43f6O+p0kCdcm9QmTD7N+bbG1BhQZ1+IZy 1A3s3KmwvTXiObejfh3AArGuZSyTH3sD5U4C6v/nE17aqrjCrbxIM/bhpqN5WNfnZ1 BppWHg73GxvqA== From: Leon Romanovsky To: Bjorn Helgaas , Logan Gunthorpe , Chaitanya Kulkarni , Greg Kroah-Hartman , Jens Axboe , Alex Williamson , Leon Romanovsky , Ankit Agrawal , Jason Gunthorpe , Jonathan Corbet , Shuah Khan , "Joerg Roedel (AMD)" , Will Deacon , Robin Murphy Cc: linux-pci@vger.kernel.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, iommu@lists.linux.dev Subject: [PATCH v3 13/17] PCI/P2PDMA: Document ACS egress control handling Date: Tue, 11 Aug 2026 12:30:55 +0300 Message-ID: <20260811-fix-p2p-acs-v3-13-efc488ee7c03@nvidia.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> References: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" X-Mailer: b4 0.15-dev-18f8f Content-Transfer-Encoding: quoted-printable From: Leon Romanovsky Document the ACS P2P Egress Control outcomes used by P2PDMA: a clear target vector bit permits direct routing, a set bit with Request Redirect enabled sends the request upstream, and a set bit with Request Redirect disabled causes an ACS Violation that P2PDMA rejects. Also record that pci=3Ddisable_acs_redir=3D clears P2P Request Redirect, Completion Redirect, and Egress Control. Reviewed-by: Logan Gunthorpe Signed-off-by: Leon Romanovsky --- Documentation/admin-guide/kernel-parameters.txt | 9 +++++---- Documentation/driver-api/pci/p2pdma.rst | 8 ++++++++ 2 files changed, 13 insertions(+), 4 deletions(-) diff --git a/Documentation/admin-guide/kernel-parameters.txt b/Documentatio= n/admin-guide/kernel-parameters.txt index b5493a7f8f22..5c3ed4fd439c 100644 --- a/Documentation/admin-guide/kernel-parameters.txt +++ b/Documentation/admin-guide/kernel-parameters.txt @@ -5226,10 +5226,11 @@ Kernel parameters disable_acs_redir=3D[; ...] Specify one or more PCI devices (in the format specified above) separated by semicolons. - Each device specified will have the PCI ACS - redirect capabilities forced off which will - allow P2P traffic between devices through - bridges without forcing it upstream. Note: + Each device specified will have the PCI ACS P2P + Request Redirect, Completion Redirect, and Egress + Control features forced off. This may allow P2P + traffic through bridges that would otherwise be + redirected upstream or blocked. Note: this removes isolation between devices and may put more devices in an IOMMU group. config_acs=3D diff --git a/Documentation/driver-api/pci/p2pdma.rst b/Documentation/driver= -api/pci/p2pdma.rst index a7fd426c3685..b759a1b828e0 100644 --- a/Documentation/driver-api/pci/p2pdma.rst +++ b/Documentation/driver-api/pci/p2pdma.rst @@ -15,6 +15,14 @@ then based on the ACS settings the transaction can route= entirely within the PCIe hierarchy and never reach the root port. The kernel will evaluate the PCIe topology and always permit P2P in these well-defined cases. =20 +ACS P2P Egress Control does not, by itself, force a transaction upstream. A +clear Egress Control Vector bit for the peer port permits direct routing; a +set bit redirects the request upstream when P2P Request Redirect is enable= d. +When Request Redirect is disabled, a set vector bit causes an ACS Violation +instead. The kernel evaluates these controls together and routes P2P DMA +through the host bridge when the direct path is blocked or cannot be +determined. + This evaluation covers the ACS controls that govern Requests carrying an Untranslated address. Unless ACS Translation Blocking is enabled, a Port with ACS Direct Translated P2P enabled routes a Request carrying a Transla= ted --=20 2.55.0 From nobody Tue Sep 29 06:58:50 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4076843C04D; Tue, 11 Aug 2026 09:32:08 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440729; cv=none; b=s9O6LkRHl/L+RkgAg2JjfOIISCl+ylXQ9MPOPVvdx6jwDmaLXxelCe4DNJqjWFfH6YzW2MK1kFJGMR46liiiz6A8PkeoOvUWprMbMxAeWnEO5bBiSGss4J3m7lIj69thEesIpIsL/elgfHzpuTtHnpplaCJ79IE79t9tevNk8sc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440729; c=relaxed/simple; bh=lwKvxUkh7+QlZngGOshWcJ+saAspfq0PUjvd2c6Y2/8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=Qff8dmatIW6DnfNDWbiIRaEY+1GokcsQpLB8dCTDXaAc0Tb63Hdr1lSggYBu5lbUqdVKJt4BLjYkhN1wAOrTZRuRORwF/rk8pkUDF8pMr2RHAhhmacjhSV+7tVVshrjlJWDtTCxzGwhKyViyuipGXpuzbIIBrenHgEjbizlwr1A= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=AjVgCymt; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="AjVgCymt" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 5B5371F000E9; Tue, 11 Aug 2026 09:32:07 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786440727; bh=a2ZO+QxTCsqdrI1xx7qw4wLfYHv6DBQMqIE+FlOHc50=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=AjVgCymtSHPbYOpippUg0nz5lYVEp8Y1djTsNI6oYJ7k284Rq/kHZn+n3sh9HK8Lk yMm7TaZNfG1FVt6FQ1oCoEZLjSMy/omI3fXr74y8o2P8q3aYnPhiaQZGbKxFSxV47/ yav91bQB+V5qcuSLjjA14XOwU7kqEnbHGxY0fCSi/cP4ehrZopwwsEX1C0uQjUT/Y/ 4UCSYrK2esMHeY9BJkXKSGJ78DohTjzRPKRmQoK+UN9v01y9GKeDbuC0skUWOReqHh G0G9LQ4Me7ecrfhaTiQ0QYmcu3/jC5di//hhHtsbCOpANT+1QKEKvIYcglgH/IzXoP o3fSXKEyrGYlQ== From: Leon Romanovsky To: Bjorn Helgaas , Logan Gunthorpe , Chaitanya Kulkarni , Greg Kroah-Hartman , Jens Axboe , Alex Williamson , Leon Romanovsky , Ankit Agrawal , Jason Gunthorpe , Jonathan Corbet , Shuah Khan , "Joerg Roedel (AMD)" , Will Deacon , Robin Murphy Cc: linux-pci@vger.kernel.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, iommu@lists.linux.dev Subject: [PATCH v3 14/17] PCI/P2PDMA: Extract pure ACS routing decision helpers Date: Tue, 11 Aug 2026 12:30:56 +0300 Message-ID: <20260811-fix-p2p-acs-v3-14-efc488ee7c03@nvidia.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> References: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" X-Mailer: b4 0.15-dev-18f8f Content-Transfer-Encoding: quoted-printable From: Leon Romanovsky The ACS Egress Control routing decision (PCIe r7.0, sec 6.12.3, table 6-11) and the Egress Control Vector Size rule were embedded in functions that also perform config-space I/O and walk the PCIe hierarchy. That made the branch-heavy logic -- in particular the paths that require an Egress Control Vector, which are unreachable on most hardware -- difficult to exercise in isolation. Factor the logic into two pure helpers: - pci_acs_p2pdma_decision() maps the ACS control word, whether the target port is known, and the target's Egress Control Vector bit to a routing state. - pci_acs_egress_port_valid() applies the "a vector size of 0 encodes 256 bits" rule to decide whether a target port is within the vector. pci_acs_p2pdma_state() and pci_acs_egress_ctrl_is_set() now call these. No functional change intended: pci_acs_egress_ctrl_is_set() still checks the port range before reading the vector DWORD. The helpers are exposed under CONFIG_KUNIT via VISIBLE_IF_KUNIT so the following patch can unit-test them. Reviewed-by: Logan Gunthorpe Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Leon Romanovsky --- drivers/pci/p2pdma.c | 60 +++++++++++++++++++++++++++++-------------------= ---- drivers/pci/pci.c | 26 +++++++++++++++++++---- drivers/pci/pci.h | 17 +++++++++++++++ 3 files changed, 73 insertions(+), 30 deletions(-) diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c index 879c92d66f5b..2c38ed56a57b 100644 --- a/drivers/pci/p2pdma.c +++ b/drivers/pci/p2pdma.c @@ -538,16 +538,40 @@ static struct pci_dev *find_parent_pci_dev(struct dev= ice *dev) return NULL; } =20 -enum pci_acs_p2pdma_state { - PCI_ACS_P2PDMA_DIRECT, - PCI_ACS_P2PDMA_REDIRECT, - PCI_ACS_P2PDMA_NOT_SUPPORTED, -}; +/* + * PCIe r7.0, sec 6.12.3, table 6-11: decide how a peer-to-peer request at= an + * ACS-capable ingress port routes, given its Egress Control register @ctr= l, + * whether the target port is known (@has_target), and that target's Egress + * Control Vector bit (@egress: 1 set, 0 clear, negative if it could not be + * read). + * + * Egress Control applies only where the target is known (the path diverge= nce). + * There, a set vector bit redirects the request only when Request Redirec= t is + * set; with Request Redirect clear it is an ACS Violation. A clear vecto= r bit + * permits direct routing, subject to Completion Redirect. + */ +VISIBLE_IF_KUNIT enum pci_acs_p2pdma_state +pci_acs_p2pdma_decision(u16 ctrl, bool has_target, int egress) +{ + if (!has_target || !(ctrl & PCI_ACS_EC)) + return ctrl & (PCI_ACS_RR | PCI_ACS_CR) ? + PCI_ACS_P2PDMA_REDIRECT : PCI_ACS_P2PDMA_DIRECT; + + if (egress < 0) + return PCI_ACS_P2PDMA_NOT_SUPPORTED; + if (egress) + return ctrl & PCI_ACS_RR ? PCI_ACS_P2PDMA_REDIRECT : + PCI_ACS_P2PDMA_NOT_SUPPORTED; + + return ctrl & PCI_ACS_CR ? PCI_ACS_P2PDMA_REDIRECT : + PCI_ACS_P2PDMA_DIRECT; +} +EXPORT_SYMBOL_IF_KUNIT(pci_acs_p2pdma_decision); =20 static enum pci_acs_p2pdma_state pci_acs_p2pdma_state(struct pci_dev *pdev, struct pci_dev *target) { - int pos, ret; + int pos, egress =3D 0; u16 ctrl; =20 pos =3D pdev->acs_cap; @@ -557,27 +581,11 @@ pci_acs_p2pdma_state(struct pci_dev *pdev, struct pci= _dev *target) if (pci_read_config_word(pdev, pos + PCI_ACS_CTRL, &ctrl)) return PCI_ACS_P2PDMA_NOT_SUPPORTED; =20 - /* EC applies only at the path divergence where the target is known. */ - if (!target || !(ctrl & PCI_ACS_EC)) - return ctrl & (PCI_ACS_RR | PCI_ACS_CR) ? - PCI_ACS_P2PDMA_REDIRECT : PCI_ACS_P2PDMA_DIRECT; + /* Egress Control is evaluated only where the target is known. */ + if (target && (ctrl & PCI_ACS_EC)) + egress =3D pci_acs_egress_ctrl_is_set(pdev, target); =20 - /* - * PCIe r7.0, sec 6.12.3, table 6-11: a set Egress Control Vector - * bit redirects the request only when Request Redirect is set. With - * Request Redirect clear, the request is handled as an ACS Violation. - * A clear vector bit permits direct routing, subject to Completion - * Redirect. - */ - ret =3D pci_acs_egress_ctrl_is_set(pdev, target); - if (ret < 0) - return PCI_ACS_P2PDMA_NOT_SUPPORTED; - if (ret) - return ctrl & PCI_ACS_RR ? PCI_ACS_P2PDMA_REDIRECT : - PCI_ACS_P2PDMA_NOT_SUPPORTED; - - return ctrl & PCI_ACS_CR ? PCI_ACS_P2PDMA_REDIRECT : - PCI_ACS_P2PDMA_DIRECT; + return pci_acs_p2pdma_decision(ctrl, !!target, egress); } =20 static void seq_buf_print_bus_devfn(struct seq_buf *buf, struct pci_dev *p= dev) diff --git a/drivers/pci/pci.c b/drivers/pci/pci.c index a633f473590f..d900fdb6f37d 100644 --- a/drivers/pci/pci.c +++ b/drivers/pci/pci.c @@ -3562,6 +3562,26 @@ static bool pci_acs_egress_vector_port(const struct = pci_dev *dev) type =3D=3D PCI_EXP_TYPE_DOWNSTREAM; } =20 +/** + * pci_acs_egress_port_valid - Is a target port within the Egress Control = Vector + * @acs_caps: the ingress port's ACS Capability register + * @target_port: the target Downstream Port number + * + * The Egress Control Vector Size occupies bits 15:8 of the ACS Capability + * register (PCIe r7.0, sec 7.7.12). A size of 0 encodes 256 bits, so + * every port number is addressable. + * + * Return: %true if @target_port has a bit in the Egress Control Vector. + */ +VISIBLE_IF_KUNIT +bool pci_acs_egress_port_valid(u16 acs_caps, u8 target_port) +{ + unsigned int vector_size =3D acs_caps >> 8; + + return !vector_size || target_port < vector_size; +} +EXPORT_SYMBOL_IF_KUNIT(pci_acs_egress_port_valid); + /** * pci_acs_egress_ctrl_is_set - Read an ACS Egress Control Vector bit * @pdev: ingress Root or Switch Downstream Port @@ -3572,7 +3592,6 @@ static bool pci_acs_egress_vector_port(const struct p= ci_dev *dev) */ int pci_acs_egress_ctrl_is_set(struct pci_dev *pdev, struct pci_dev *targe= t) { - unsigned int vector_size; u32 lnkcap, vector; u8 target_port; int ret; @@ -3597,10 +3616,8 @@ int pci_acs_egress_ctrl_is_set(struct pci_dev *pdev,= struct pci_dev *target) return pcibios_err_to_errno(ret); =20 target_port =3D FIELD_GET(PCI_EXP_LNKCAP_PN, lnkcap); - vector_size =3D pdev->acs_capabilities >> 8; =20 - /* An Egress Control Vector Size of 0 encodes 256 bits. */ - if (vector_size && target_port >=3D vector_size) + if (!pci_acs_egress_port_valid(pdev->acs_capabilities, target_port)) return -ERANGE; =20 ret =3D pci_read_config_dword(pdev, @@ -3612,6 +3629,7 @@ int pci_acs_egress_ctrl_is_set(struct pci_dev *pdev, = struct pci_dev *target) =20 return !!(vector & BIT(target_port % 32)); } +EXPORT_SYMBOL_IF_KUNIT(pci_acs_egress_ctrl_is_set); =20 static bool pci_acs_flags_enabled(struct pci_dev *pdev, u16 acs_flags, enum pci_acs_scope scope) diff --git a/drivers/pci/pci.h b/drivers/pci/pci.h index 32394e349766..06a18aa663bc 100644 --- a/drivers/pci/pci.h +++ b/drivers/pci/pci.h @@ -2,6 +2,7 @@ #ifndef DRIVERS_PCI_H #define DRIVERS_PCI_H =20 +#include #include #include #include @@ -1079,6 +1080,22 @@ static inline bool pci_acs_rr_ineffective(u32 ctrl, = u16 acs_flags, } =20 int pci_acs_egress_ctrl_is_set(struct pci_dev *pdev, struct pci_dev *targe= t); + +/* + * Peer-to-peer routing decision for an ACS-capable ingress port, per + * PCIe r7.0, sec 6.12.3, table 6-11. + */ +enum pci_acs_p2pdma_state { + PCI_ACS_P2PDMA_DIRECT, /* peer-to-peer permitted directly */ + PCI_ACS_P2PDMA_REDIRECT, /* redirected upstream to host bridge */ + PCI_ACS_P2PDMA_NOT_SUPPORTED, /* no usable peer-to-peer route */ +}; + +#if IS_ENABLED(CONFIG_KUNIT) +bool pci_acs_egress_port_valid(u16 acs_caps, u8 target_port); +enum pci_acs_p2pdma_state pci_acs_p2pdma_decision(u16 ctrl, bool has_targe= t, + int egress); +#endif #ifdef CONFIG_PCI_QUIRKS int pci_dev_specific_acs_enabled(struct pci_dev *dev, u16 acs_flags, enum pci_acs_scope scope); --=20 2.55.0 From nobody Tue Sep 29 06:58:50 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8C7F343C7A9; Tue, 11 Aug 2026 09:32:11 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440733; cv=none; b=GVz3H9XCSPvsJiIvKR1RyAYNwnhgDMLyagUGANLeFCbDK/yUaGrAive3LCBk53u0hIjXkGVnFImX4RZtKg6PbS0un0o7R5bmLTL4iD3AALR1Qvskbz/PUmABsZpuS/I8BYxyT4PZhTtH5KRaBCyI8gZsT6I7pV+jo1gUvep8axs= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440733; c=relaxed/simple; bh=k1VKo034gATB3naQi33zmZ76nnuEjG89B1MqTal7ZGg=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=J1qqV72OJZONtwM1d9QILYxRVQOGVfJibew2Hz2bRONOCEZaACxz3CuqZiTe4/6tOD9V1rw7HmhmfzbZvwhpK2z2D/Qvahf/c3a//1jn9rBPCJ3uQzi5SSVKVQcosuM30SUTa5eUh76TPMwuFWOVUXw8teOh18CoB20skZlhH1o= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=LuGLpCwm; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="LuGLpCwm" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 8C7361F000E9; Tue, 11 Aug 2026 09:32:10 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786440731; bh=TFWZLl7GIDQamJP4g/1bkCB0urk85cTRZoDQeKa9AYA=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=LuGLpCwmp1DfAaiFCBSZ/R43j3iuP3aQgwZwRVK01h6hFqqs8djRWQcOu6jXkN/d5 Bab7MOvJgK4ZYZnt/458l8iM04z+Ss5KOO1axO9Bj0jaVbc1+4alX+lQ2CnDRkvQ0m WptS7MXLGXre7cGcuxInnbF7CAhBAGnLHBetlm15uIE3I3FhfsJ29N58ImQsb5bV9a ks+nwJMoe1l2sDJ46Y73AfkG84kTVTk3wAwB/If5WrUBDMCAzqpC4JJQ4gJljgRASy 9uSxefOIDgPH+ZJeCMB8JbSK22MqeKc0aXZcuhk31i9T98Z7zBpajtFjuEHnQS8TOS 06X1/Q8RKdpbw== From: Leon Romanovsky To: Bjorn Helgaas , Logan Gunthorpe , Chaitanya Kulkarni , Greg Kroah-Hartman , Jens Axboe , Alex Williamson , Leon Romanovsky , Ankit Agrawal , Jason Gunthorpe , Jonathan Corbet , Shuah Khan , "Joerg Roedel (AMD)" , Will Deacon , Robin Murphy Cc: linux-pci@vger.kernel.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, iommu@lists.linux.dev Subject: [PATCH v3 15/17] PCI/P2PDMA: Add KUnit tests for ACS routing decisions Date: Tue, 11 Aug 2026 12:30:57 +0300 Message-ID: <20260811-fix-p2p-acs-v3-15-efc488ee7c03@nvidia.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> References: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" X-Mailer: b4 0.15-dev-18f8f Content-Transfer-Encoding: quoted-printable From: Leon Romanovsky Add a KUnit suite exercising the ACS peer-to-peer routing logic: - pci_acs_p2pdma_decision(): the full PCIe table 6-11 truth table, including the Egress Control Vector branches (bit set/clear, with and without Request Redirect and Completion Redirect) that require a switch implementing the Egress Control Vector and so cannot be reached on commonly available hardware. - pci_acs_egress_port_valid(): the vector-size boundary, including the "size 0 encodes 256 bits" case. - pci_acs_egress_ctrl_is_set(): driven through a fake pci_ops returning canned config space, covering target Port Number extraction from LNKCAP, the vector DWORD offset (target_port / 32), the bit position (target_port % 32), the -ERANGE bound, and the unsupported-port and shared-bus guards -- all without real hardware. Run with: cat > /tmp/pci-acs.kunitconfig <<'EOF' CONFIG_KUNIT=3Dy CONFIG_PCI=3Dy CONFIG_ZONE_DEVICE=3Dy CONFIG_MEMORY_HOTPLUG=3Dy CONFIG_MEMORY_HOTREMOVE=3Dy CONFIG_SPARSEMEM_VMEMMAP=3Dy CONFIG_PCI_P2PDMA=3Dy CONFIG_PCI_ACS_KUNIT_TEST=3Dy EOF ./tools/testing/kunit/kunit.py run --arch=3Dx86_64 \ --kunitconfig=3D/tmp/pci-acs.kunitconfig --jobs=3D$(nproc) pci_acs Assisted-by: Claude:claude-opus-4-8 Reviewed-by: Logan Gunthorpe Signed-off-by: Leon Romanovsky --- drivers/pci/Kconfig | 15 ++ drivers/pci/Makefile | 1 + drivers/pci/pci_acs_test.c | 416 +++++++++++++++++++++++++++++++++++++++++= ++++ 3 files changed, 432 insertions(+) diff --git a/drivers/pci/Kconfig b/drivers/pci/Kconfig index 0c7408509ba2..30ad7f407c6f 100644 --- a/drivers/pci/Kconfig +++ b/drivers/pci/Kconfig @@ -226,6 +226,21 @@ config PCI_P2PDMA =20 If unsure, say N. =20 +config PCI_ACS_KUNIT_TEST + tristate "KUnit tests for PCI ACS P2P routing" if !KUNIT_ALL_TESTS + depends on PCI_P2PDMA && KUNIT + default KUNIT_ALL_TESTS + help + Enable KUnit tests for the PCI ACS peer-to-peer routing decision + logic (PCIe ACS Egress Control, table 6-11), including the code + paths that require an ACS Egress Control Vector and so cannot be + exercised on typical peer-to-peer hardware. + + For more information on KUnit and unit tests in general, refer to + the KUnit documentation in Documentation/dev-tools/kunit/. + + If unsure, say N. + config PCI_LABEL def_bool y if (DMI || ACPI) select NLS diff --git a/drivers/pci/Makefile b/drivers/pci/Makefile index 41ebc3b9a518..6305d128d3df 100644 --- a/drivers/pci/Makefile +++ b/drivers/pci/Makefile @@ -31,6 +31,7 @@ obj-$(CONFIG_PCI_STUB) +=3D pci-stub.o obj-$(CONFIG_PCI_PF_STUB) +=3D pci-pf-stub.o obj-$(CONFIG_PCI_ECAM) +=3D ecam.o obj-$(CONFIG_PCI_P2PDMA) +=3D p2pdma.o +obj-$(CONFIG_PCI_ACS_KUNIT_TEST) +=3D pci_acs_test.o obj-$(CONFIG_XEN_PCIDEV_FRONTEND) +=3D xen-pcifront.o obj-$(CONFIG_VGA_ARB) +=3D vgaarb.o obj-$(CONFIG_PCI_DOE) +=3D doe.o diff --git a/drivers/pci/pci_acs_test.c b/drivers/pci/pci_acs_test.c new file mode 100644 index 000000000000..08d4654b95a9 --- /dev/null +++ b/drivers/pci/pci_acs_test.c @@ -0,0 +1,416 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * KUnit tests for PCI ACS peer-to-peer routing decision logic. + * + * These exercise the pure helpers factored out of the ACS Egress Control + * handling (PCIe r7.0, sec 6.12.3, table 6-11). They cover the code paths + * that require an ACS Egress Control Vector, which cannot be reached on t= he + * peer-to-peer hardware commonly available for testing. + */ +#include + +#include +#include + +#include "pci.h" + +/* pci_acs_p2pdma_decision(): the table 6-11 truth table. */ + +struct acs_decision_case { + const char *desc; + u16 ctrl; + bool has_target; + int egress; + enum pci_acs_p2pdma_state expect; +}; + +/* Shorthands to keep the table below readable. */ +#define ACS_DIRECT PCI_ACS_P2PDMA_DIRECT +#define ACS_REDIR PCI_ACS_P2PDMA_REDIRECT +#define ACS_NO_P2P PCI_ACS_P2PDMA_NOT_SUPPORTED + +static const struct acs_decision_case acs_decision_cases[] =3D { + /* No target known: Egress Control is ignored, RR/CR decide. */ + { "no_target/none", 0, false, 0, ACS_DIRECT }, + { "no_target/rr", PCI_ACS_RR, false, 0, ACS_REDIR }, + { "no_target/cr", PCI_ACS_CR, false, 0, ACS_REDIR }, + { "no_target/ec_only", PCI_ACS_EC, false, 0, ACS_DIRECT }, + + /* Target known but EC clear: RR/CR decide, egress not consulted. */ + { "ec_clear/none", 0, true, 0, ACS_DIRECT }, + { "ec_clear/rr", PCI_ACS_RR, true, 0, ACS_REDIR }, + { "ec_clear/cr", PCI_ACS_CR, true, 0, ACS_REDIR }, + { "ec_clear/rr_cr", PCI_ACS_RR | PCI_ACS_CR, true, 0, ACS_REDIR }, + + /* EC set but vector unreadable: never a usable P2P route. */ + { "ec/eopnotsupp", PCI_ACS_EC | PCI_ACS_RR, true, -EOPNOTSUPP, ACS_NO_P2P= }, + { "ec/erange", PCI_ACS_EC | PCI_ACS_CR, true, -ERANGE, ACS_NO_P2P }, + + /* EC set, vector bit set: redirect iff RR, else ACS Violation. */ + { "ec/vec_set/none", PCI_ACS_EC, true, 1, ACS_NO_P2P }, + { "ec/vec_set/cr", PCI_ACS_EC | PCI_ACS_CR, true, 1, ACS_NO_P2P }, + { "ec/vec_set/rr", PCI_ACS_EC | PCI_ACS_RR, true, 1, ACS_REDIR }, + { "ec/vec_set/rr_cr", PCI_ACS_EC | PCI_ACS_RR | PCI_ACS_CR, true, 1, + ACS_REDIR }, + + /* EC set, vector bit clear: direct unless CR redirects. */ + { "ec/vec_clear/none", PCI_ACS_EC, true, 0, ACS_DIRECT }, + { "ec/vec_clear/rr", PCI_ACS_EC | PCI_ACS_RR, true, 0, ACS_DIRECT }, + { "ec/vec_clear/cr", PCI_ACS_EC | PCI_ACS_CR, true, 0, ACS_REDIR }, + { "ec/vec_clear/rr_cr", PCI_ACS_EC | PCI_ACS_RR | PCI_ACS_CR, true, 0, + ACS_REDIR }, +}; + +#undef ACS_DIRECT +#undef ACS_REDIR +#undef ACS_NO_P2P + +static void acs_decision_desc(const struct acs_decision_case *c, char *des= c) +{ + strscpy(desc, c->desc, KUNIT_PARAM_DESC_SIZE); +} + +KUNIT_ARRAY_PARAM(acs_decision, acs_decision_cases, acs_decision_desc); + +static void pci_acs_p2pdma_decision_test(struct kunit *test) +{ + const struct acs_decision_case *c =3D test->param_value; + + KUNIT_EXPECT_EQ(test, + pci_acs_p2pdma_decision(c->ctrl, c->has_target, c->egress), + c->expect); +} + +/* pci_acs_egress_port_valid(): the Egress Control Vector Size rule. */ + +struct egress_valid_case { + const char *desc; + u16 acs_caps; + u8 target_port; + bool expect; +}; + +static const struct egress_valid_case egress_valid_cases[] =3D { + /* A Vector Size of 0 encodes 256 bits, so every port is addressable. */ + { "size0/port0", 0x0000, 0, true }, + { "size0/port255", 0x0000, 255, true }, + /* Vector Size N (bits 15:8): ports [0, N) are addressable. */ + { "size1/port0", 0x0100, 0, true }, + { "size1/port1", 0x0100, 1, false }, + { "size8/port7", 0x0800, 7, true }, + { "size8/port8", 0x0800, 8, false }, + { "size255/port254", 0xff00, 254, true }, + { "size255/port255", 0xff00, 255, false }, +}; + +static void egress_valid_desc(const struct egress_valid_case *c, char *des= c) +{ + strscpy(desc, c->desc, KUNIT_PARAM_DESC_SIZE); +} + +KUNIT_ARRAY_PARAM(egress_valid, egress_valid_cases, egress_valid_desc); + +static void pci_acs_egress_port_valid_test(struct kunit *test) +{ + const struct egress_valid_case *c =3D test->param_value; + + KUNIT_EXPECT_EQ(test, + pci_acs_egress_port_valid(c->acs_caps, c->target_port), + c->expect); +} + +/* + * pci_acs_egress_ctrl_is_set(): drive the config-space reads with a fake = pci_ops + * so the Egress Control Vector lookup is exercised without real hardware = -- + * the target Port Number from LNKCAP, the vector DWORD at target_port/32,= and + * the bit at target_port%32. + */ + +/* PCIe Capabilities register value: device/port @type, capability version= 2. */ +#define ACS_TEST_PCIE_FLAGS(type) (((type) << 4) | 0x2) +#define ACS_DOWNSTREAM ACS_TEST_PCIE_FLAGS(PCI_EXP_TYPE_DOWNSTREAM) +#define ACS_ENDPOINT ACS_TEST_PCIE_FLAGS(PCI_EXP_TYPE_ENDPOINT) +#define ACS_ROOT_PORT ACS_TEST_PCIE_FLAGS(PCI_EXP_TYPE_ROOT_PORT) +#define ACS_PCIE_BRIDGE ACS_TEST_PCIE_FLAGS(PCI_EXP_TYPE_PCIE_BRIDGE) + +struct acs_fake_cfg { + unsigned int pdev_devfn; + unsigned int target_devfn; + u16 pdev_acs_cap; + u8 target_pcie_cap; + u8 target_port; + u32 egress_vector[8]; /* full 256-bit vector */ +}; + +static int acs_fake_cfg_read(struct pci_bus *bus, unsigned int devfn, + int where, int size, u32 *val) +{ + struct acs_fake_cfg *cfg =3D bus->sysdata; + + *val =3D 0; + if (size !=3D 4) + return PCIBIOS_SUCCESSFUL; + + if (devfn =3D=3D cfg->target_devfn && + where =3D=3D cfg->target_pcie_cap + PCI_EXP_LNKCAP) { + *val =3D FIELD_PREP(PCI_EXP_LNKCAP_PN, cfg->target_port); + } else if (devfn =3D=3D cfg->pdev_devfn) { + int base =3D cfg->pdev_acs_cap + PCI_ACS_EGRESS_CTL_V; + + if (where >=3D base && + where < base + (int)sizeof(cfg->egress_vector)) + *val =3D cfg->egress_vector[(where - base) / 4]; + } + return PCIBIOS_SUCCESSFUL; +} + +static int acs_fake_cfg_write(struct pci_bus *bus, unsigned int devfn, + int where, int size, u32 val) +{ + return PCIBIOS_SUCCESSFUL; +} + +static struct pci_ops acs_fake_ops =3D { + .read =3D acs_fake_cfg_read, + .write =3D acs_fake_cfg_write, +}; + +static struct acs_fake_cfg acs_base_cfg(void) +{ + return (struct acs_fake_cfg){ + .pdev_devfn =3D PCI_DEVFN(0, 0), + .target_devfn =3D PCI_DEVFN(1, 0), + .pdev_acs_cap =3D 0x100, + .target_pcie_cap =3D 0x40, + }; +} + +static int acs_egress_ctrl_set(struct kunit *test, struct acs_fake_cfg *cf= g, + u16 pdev_acs_caps, u16 pdev_flags, u16 target_flags) +{ + struct pci_bus *bus =3D kunit_kzalloc(test, sizeof(*bus), GFP_KERNEL); + struct pci_dev *pdev =3D kunit_kzalloc(test, sizeof(*pdev), GFP_KERNEL); + struct pci_dev *target =3D kunit_kzalloc(test, sizeof(*target), GFP_KERNE= L); + + KUNIT_ASSERT_NOT_NULL(test, bus); + KUNIT_ASSERT_NOT_NULL(test, pdev); + KUNIT_ASSERT_NOT_NULL(test, target); + + bus->ops =3D &acs_fake_ops; + bus->sysdata =3D cfg; + + pdev->bus =3D bus; + pdev->devfn =3D cfg->pdev_devfn; + pdev->acs_cap =3D cfg->pdev_acs_cap; + pdev->acs_capabilities =3D pdev_acs_caps; + pdev->pcie_cap =3D 0x40; + pdev->pcie_flags_reg =3D pdev_flags; + + target->bus =3D bus; + target->devfn =3D cfg->target_devfn; + target->pcie_cap =3D cfg->target_pcie_cap; + target->pcie_flags_reg =3D target_flags; + + return pci_acs_egress_ctrl_is_set(pdev, target); +} + +static void acs_egress_vector_bit_set_test(struct kunit *test) +{ + struct acs_fake_cfg cfg =3D acs_base_cfg(); + + cfg.target_port =3D 5; + cfg.egress_vector[0] =3D BIT(5); + + KUNIT_EXPECT_EQ(test, + acs_egress_ctrl_set(test, &cfg, PCI_ACS_EC | (32 << 8), + ACS_DOWNSTREAM, ACS_DOWNSTREAM), + 1); +} + +static void acs_egress_vector_bit_clear_test(struct kunit *test) +{ + struct acs_fake_cfg cfg =3D acs_base_cfg(); + + cfg.target_port =3D 5; /* vector left all-zero */ + + KUNIT_EXPECT_EQ(test, + acs_egress_ctrl_set(test, &cfg, PCI_ACS_EC | (32 << 8), + ACS_DOWNSTREAM, ACS_DOWNSTREAM), + 0); +} + +static void acs_egress_high_port_index_test(struct kunit *test) +{ + struct acs_fake_cfg cfg =3D acs_base_cfg(); + + /* Port 40 lives in vector DWORD 1, bit 8: exercises target_port/32. */ + cfg.target_port =3D 40; + cfg.egress_vector[1] =3D BIT(40 % 32); + + KUNIT_EXPECT_EQ(test, + acs_egress_ctrl_set(test, &cfg, PCI_ACS_EC | (64 << 8), + ACS_DOWNSTREAM, ACS_DOWNSTREAM), + 1); +} + +static void acs_egress_port_out_of_range_test(struct kunit *test) +{ + struct acs_fake_cfg cfg =3D acs_base_cfg(); + + /* Vector Size 8, port 40 is beyond it. */ + cfg.target_port =3D 40; + + KUNIT_EXPECT_EQ(test, + acs_egress_ctrl_set(test, &cfg, PCI_ACS_EC | (8 << 8), + ACS_DOWNSTREAM, ACS_DOWNSTREAM), + -ERANGE); +} + +static void acs_egress_no_ec_cap_test(struct kunit *test) +{ + struct acs_fake_cfg cfg =3D acs_base_cfg(); + + cfg.target_port =3D 5; + + /* acs_capabilities without PCI_ACS_EC: unsupported. */ + KUNIT_EXPECT_EQ(test, + acs_egress_ctrl_set(test, &cfg, 32 << 8, + ACS_DOWNSTREAM, ACS_DOWNSTREAM), + -EOPNOTSUPP); +} + +static void acs_egress_pdev_not_downstream_test(struct kunit *test) +{ + struct acs_fake_cfg cfg =3D acs_base_cfg(); + + cfg.target_port =3D 5; + + KUNIT_EXPECT_EQ(test, + acs_egress_ctrl_set(test, &cfg, PCI_ACS_EC | (32 << 8), + ACS_ENDPOINT, ACS_DOWNSTREAM), + -EOPNOTSUPP); +} + +static void acs_egress_target_not_downstream_test(struct kunit *test) +{ + struct acs_fake_cfg cfg =3D acs_base_cfg(); + + cfg.target_port =3D 5; + + KUNIT_EXPECT_EQ(test, + acs_egress_ctrl_set(test, &cfg, PCI_ACS_EC | (32 << 8), + ACS_DOWNSTREAM, ACS_ENDPOINT), + -EOPNOTSUPP); +} + +/* + * The vector is indexed by Port Number only for Root Ports and Switch + * Downstream Ports, so a PCI/PCI-X to PCIe Bridge must not be indexed by = its + * Link Capabilities Port Number. + */ +static void acs_egress_pdev_pcie_bridge_test(struct kunit *test) +{ + struct acs_fake_cfg cfg =3D acs_base_cfg(); + + cfg.target_port =3D 5; + + KUNIT_EXPECT_EQ(test, + acs_egress_ctrl_set(test, &cfg, PCI_ACS_EC | (32 << 8), + ACS_PCIE_BRIDGE, ACS_DOWNSTREAM), + -EOPNOTSUPP); +} + +static void acs_egress_target_pcie_bridge_test(struct kunit *test) +{ + struct acs_fake_cfg cfg =3D acs_base_cfg(); + + cfg.target_port =3D 5; + + KUNIT_EXPECT_EQ(test, + acs_egress_ctrl_set(test, &cfg, PCI_ACS_EC | (32 << 8), + ACS_DOWNSTREAM, ACS_PCIE_BRIDGE), + -EOPNOTSUPP); +} + +/* + * Each vector bit is a Port Number within one Switch or Root Complex (PCIe + * r7.0, sec 7.7.12.4), so a target that does not share the ingress port's= bus + * has no bit here even when the bit at its Port Number is set. + */ +static void acs_egress_target_other_bus_test(struct kunit *test) +{ + struct acs_fake_cfg cfg =3D acs_base_cfg(); + struct pci_bus *bus =3D kunit_kzalloc(test, sizeof(*bus), GFP_KERNEL); + struct pci_bus *other =3D kunit_kzalloc(test, sizeof(*other), GFP_KERNEL); + struct pci_dev *pdev =3D kunit_kzalloc(test, sizeof(*pdev), GFP_KERNEL); + struct pci_dev *target =3D kunit_kzalloc(test, sizeof(*target), GFP_KERNE= L); + + KUNIT_ASSERT_NOT_NULL(test, bus); + KUNIT_ASSERT_NOT_NULL(test, other); + KUNIT_ASSERT_NOT_NULL(test, pdev); + KUNIT_ASSERT_NOT_NULL(test, target); + + cfg.target_port =3D 5; + cfg.egress_vector[0] =3D BIT(5); + + bus->ops =3D &acs_fake_ops; + bus->sysdata =3D &cfg; + other->ops =3D &acs_fake_ops; + other->sysdata =3D &cfg; + + pdev->bus =3D bus; + pdev->devfn =3D cfg.pdev_devfn; + pdev->acs_cap =3D cfg.pdev_acs_cap; + pdev->acs_capabilities =3D PCI_ACS_EC | (32 << 8); + pdev->pcie_cap =3D 0x40; + pdev->pcie_flags_reg =3D ACS_DOWNSTREAM; + + target->bus =3D other; + target->devfn =3D cfg.target_devfn; + target->pcie_cap =3D cfg.target_pcie_cap; + target->pcie_flags_reg =3D ACS_DOWNSTREAM; + + KUNIT_EXPECT_EQ(test, pci_acs_egress_ctrl_is_set(pdev, target), + -EOPNOTSUPP); +} + +/* A Root Port is a valid ingress and egress port for the vector. */ +static void acs_egress_root_port_test(struct kunit *test) +{ + struct acs_fake_cfg cfg =3D acs_base_cfg(); + + cfg.target_port =3D 5; + cfg.egress_vector[0] =3D BIT(5); + + KUNIT_EXPECT_EQ(test, + acs_egress_ctrl_set(test, &cfg, PCI_ACS_EC | (32 << 8), + ACS_ROOT_PORT, ACS_ROOT_PORT), + 1); +} + +static struct kunit_case pci_acs_test_cases[] =3D { + KUNIT_CASE_PARAM(pci_acs_p2pdma_decision_test, acs_decision_gen_params), + KUNIT_CASE_PARAM(pci_acs_egress_port_valid_test, egress_valid_gen_params), + KUNIT_CASE(acs_egress_vector_bit_set_test), + KUNIT_CASE(acs_egress_vector_bit_clear_test), + KUNIT_CASE(acs_egress_high_port_index_test), + KUNIT_CASE(acs_egress_port_out_of_range_test), + KUNIT_CASE(acs_egress_no_ec_cap_test), + KUNIT_CASE(acs_egress_pdev_not_downstream_test), + KUNIT_CASE(acs_egress_target_not_downstream_test), + KUNIT_CASE(acs_egress_pdev_pcie_bridge_test), + KUNIT_CASE(acs_egress_target_pcie_bridge_test), + KUNIT_CASE(acs_egress_target_other_bus_test), + KUNIT_CASE(acs_egress_root_port_test), + {} +}; + +static struct kunit_suite pci_acs_test_suite =3D { + .name =3D "pci_acs", + .test_cases =3D pci_acs_test_cases, +}; +kunit_test_suite(pci_acs_test_suite); + +MODULE_IMPORT_NS("EXPORTED_FOR_KUNIT_TESTING"); +MODULE_LICENSE("GPL"); +MODULE_DESCRIPTION("KUnit tests for PCI ACS peer-to-peer routing decisions= "); --=20 2.55.0 From nobody Tue Sep 29 06:58:50 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5B35342E01F; Tue, 11 Aug 2026 09:32:14 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440735; cv=none; b=OKtTepEjqyQ+KWza9YDwszkqVvYgSLrdjaO7F4iITYjG5Wt1BsP0qfwwXfPwrZiTiXk1FBtcaE5z60tkDGw0UMcs4gjzbewcenb0026gNRROKbfmH5DNYq0hFO0dULGfeqZ5n0GFcBg2CQKxCxDK16NAOQewvdsnzs5YNahacyw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440735; c=relaxed/simple; bh=1Zupz39mijiB2eKoFgrRCIUPxUSyOm+t3/Myl03mv6k=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=EeV6sdbHyTkU+G81h3EAeDvgTKOn2+HoBdJj/zEuRVukNTbNpj6tcXCIErtd2rl5QJgnS0i/LRtFh7rUv3VaT/RGAL9GRXEyDSNXmtti1CcsCcLYwYNoxLQqgrqB0rYrFRq2RWAFbd/Mp8J2gQhKojNyk+lmLDCBhVgN4WkowRQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=iulMCgJM; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="iulMCgJM" Received: by smtp.kernel.org (Postfix) with ESMTPSA id AEF991F00A3A; Tue, 11 Aug 2026 09:32:13 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786440734; bh=QS+4DdLbcZFSX5l8hs7wBp22auL1V3b3oXdmJboJGfo=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=iulMCgJMqyYJqazy7CDsWpae/lI2hpBla1XmJxBZiOK+l/ojBmn2QJDoOpJd+/Aq/ bIGYYfrkgtgHTmFKmeoMKk65U8njgkaLQJ43g9DsgZJI4TZbixbiuQPZUIjKlfH3lD lP+TEXsOsqKSksUAz5iYADg00843LML6054GLIENOXNw3m1Itt/mwsbCCAjon7Hu4q aqTWf5Gqr6E/pXmDb5obck6m8MjMd3JNTgJusA2y1CpNDp2320O76ntBxL1gwZQ8/2 hPYAACB2pVJEcUgvDx+8JjBsxgXjRG2Xb1Nblwz3FpZOpqHsx0WUywgDQ6XrCG6FJ+ F/GXaGY/aw0Aw== From: Leon Romanovsky To: Bjorn Helgaas , Logan Gunthorpe , Chaitanya Kulkarni , Greg Kroah-Hartman , Jens Axboe , Alex Williamson , Leon Romanovsky , Ankit Agrawal , Jason Gunthorpe , Jonathan Corbet , Shuah Khan , "Joerg Roedel (AMD)" , Will Deacon , Robin Murphy Cc: linux-pci@vger.kernel.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, iommu@lists.linux.dev Subject: [PATCH v3 16/17] PCI/P2PDMA: Add KUnit coverage for the ACS P2P routing walk Date: Tue, 11 Aug 2026 12:30:58 +0300 Message-ID: <20260811-fix-p2p-acs-v3-16-efc488ee7c03@nvidia.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> References: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" X-Mailer: b4 0.15-dev-18f8f Content-Transfer-Encoding: quoted-printable From: Leon Romanovsky Extend the ACS KUnit suite with end-to-end coverage of calc_map_type_and_dist(), the provider-to-client hierarchy walk. A fabricated PCIe fabric (host bridge, Root Port, Switch Upstream Port, two Switch Downstream Ports and the provider/client endpoints) with a fake pci_ops backing the ACS Control, Egress Control Vector and LNKCAP reads lets the walk run without real hardware. The tests assert: - BUS_ADDR when no port on the path enables ACS; - THRU_HOST_BRIDGE when a Downstream Port's Egress Control Vector routes the peer with Request Redirect clear (an ACS Violation) at the path divergence, leaving only the host-bridge route; - BUS_ADDR when Egress Control is enabled but the peer's vector bit is clear; - THRU_HOST_BRIDGE when Request Redirect redirects the request and the host bridge is whitelisted. calc_map_type_and_dist() is exposed under CONFIG_KUNIT via VISIBLE_IF_KUNIT. Reviewed-by: Logan Gunthorpe Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Leon Romanovsky --- drivers/pci/p2pdma.c | 3 +- drivers/pci/pci.h | 4 + drivers/pci/pci_acs_test.c | 204 +++++++++++++++++++++++++++++++++++++++++= ++++ 3 files changed, 210 insertions(+), 1 deletion(-) diff --git a/drivers/pci/p2pdma.c b/drivers/pci/p2pdma.c index 2c38ed56a57b..34929bc6efb6 100644 --- a/drivers/pci/p2pdma.c +++ b/drivers/pci/p2pdma.c @@ -774,7 +774,7 @@ static unsigned long map_types_idx(struct pci_dev *clie= nt) * ports per above. If the device is not in the whitelist, return * PCI_P2PDMA_MAP_NOT_SUPPORTED. */ -static enum pci_p2pdma_map_type +VISIBLE_IF_KUNIT enum pci_p2pdma_map_type calc_map_type_and_dist(struct pci_dev *provider, struct pci_dev *client, int *dist, bool verbose) { @@ -911,6 +911,7 @@ calc_map_type_and_dist(struct pci_dev *provider, struct= pci_dev *client, rcu_read_unlock(); return map_type; } +EXPORT_SYMBOL_IF_KUNIT(calc_map_type_and_dist); =20 /** * pci_p2pdma_distance_many - Determine the cumulative distance between diff --git a/drivers/pci/pci.h b/drivers/pci/pci.h index 06a18aa663bc..c55e6ea7c9de 100644 --- a/drivers/pci/pci.h +++ b/drivers/pci/pci.h @@ -6,6 +6,7 @@ #include #include #include +#include #include =20 struct pcie_tlp_log; @@ -1095,6 +1096,9 @@ enum pci_acs_p2pdma_state { bool pci_acs_egress_port_valid(u16 acs_caps, u8 target_port); enum pci_acs_p2pdma_state pci_acs_p2pdma_decision(u16 ctrl, bool has_targe= t, int egress); +enum pci_p2pdma_map_type calc_map_type_and_dist(struct pci_dev *provider, + struct pci_dev *client, + int *dist, bool verbose); #endif #ifdef CONFIG_PCI_QUIRKS int pci_dev_specific_acs_enabled(struct pci_dev *dev, u16 acs_flags, diff --git a/drivers/pci/pci_acs_test.c b/drivers/pci/pci_acs_test.c index 08d4654b95a9..130605b91b47 100644 --- a/drivers/pci/pci_acs_test.c +++ b/drivers/pci/pci_acs_test.c @@ -10,6 +10,7 @@ #include =20 #include +#include #include =20 #include "pci.h" @@ -388,6 +389,205 @@ static void acs_egress_root_port_test(struct kunit *t= est) 1); } =20 +/* + * calc_map_type_and_dist(): drive the full provider->client hierarchy walk + * over a fabricated PCIe fabric matching the canonical "two devices behin= d one + * switch" tree: + * + * host bridge / root bus + * Root Port + * Switch Upstream Port + * Switch Downstream Port 0 -- provider + * Switch Downstream Port 1 -- client + * + * A fake pci_ops answers the ACS Control, Egress Control Vector and LNKCAP + * reads for the two downstream ports, so the ACS Egress Control evaluated= at + * the path divergence (Downstream Port 0 targeting Downstream Port 1) dec= ides + * the mapping without any real hardware. + */ + +struct acs_dn_cfg { + u16 acs_ctrl; /* ACS Control register value */ + u8 port; /* this port's LNKCAP Port Number */ + u32 egress[8]; /* Egress Control Vector (256 bits) */ +}; + +struct acs_fabric { + struct pci_dev *provider; + struct pci_dev *client; + struct pci_dev *dn0; /* Downstream Port 0 (provider side) */ + struct pci_dev *dn1; /* Downstream Port 1 (client side) */ + struct acs_dn_cfg dn0_cfg; + struct acs_dn_cfg dn1_cfg; +}; + +static void acs_dn_read(struct pci_dev *dn, struct acs_dn_cfg *c, + int where, int size, u32 *val) +{ + int vec =3D dn->acs_cap + PCI_ACS_EGRESS_CTL_V; + + if (size =3D=3D 4 && where =3D=3D dn->pcie_cap + PCI_EXP_LNKCAP) + *val =3D FIELD_PREP(PCI_EXP_LNKCAP_PN, c->port); + else if (dn->acs_cap && size =3D=3D 2 && where =3D=3D dn->acs_cap + PCI_A= CS_CTRL) + *val =3D c->acs_ctrl; + else if (dn->acs_cap && size =3D=3D 4 && + where >=3D vec && where < vec + (int)sizeof(c->egress)) + *val =3D c->egress[(where - vec) / 4]; +} + +static int acs_fabric_read(struct pci_bus *bus, unsigned int devfn, + int where, int size, u32 *val) +{ + struct acs_fabric *f =3D bus->sysdata; + + *val =3D 0; + if (bus =3D=3D f->dn0->bus && devfn =3D=3D f->dn0->devfn) + acs_dn_read(f->dn0, &f->dn0_cfg, where, size, val); + else if (bus =3D=3D f->dn1->bus && devfn =3D=3D f->dn1->devfn) + acs_dn_read(f->dn1, &f->dn1_cfg, where, size, val); + return PCIBIOS_SUCCESSFUL; +} + +static int acs_fabric_write(struct pci_bus *bus, unsigned int devfn, + int where, int size, u32 val) +{ + return PCIBIOS_SUCCESSFUL; +} + +static struct pci_ops acs_fabric_ops =3D { + .read =3D acs_fabric_read, + .write =3D acs_fabric_write, +}; + +static struct pci_bus *acs_add_bus(struct kunit *test, struct pci_bus *par= ent, + struct pci_dev *self, u8 nr, void *sysdata) +{ + struct pci_bus *bus =3D kunit_kzalloc(test, sizeof(*bus), GFP_KERNEL); + + KUNIT_ASSERT_NOT_NULL(test, bus); + bus->parent =3D parent; + bus->self =3D self; + bus->number =3D nr; + bus->ops =3D &acs_fabric_ops; + bus->sysdata =3D sysdata; + INIT_LIST_HEAD(&bus->devices); + return bus; +} + +static struct pci_dev *acs_add_dev(struct kunit *test, struct pci_bus *bus, + unsigned int devfn, int pcie_type) +{ + struct pci_dev *dev =3D kunit_kzalloc(test, sizeof(*dev), GFP_KERNEL); + + KUNIT_ASSERT_NOT_NULL(test, dev); + dev->bus =3D bus; + dev->devfn =3D devfn; + dev->pcie_cap =3D 0x40; + dev->pcie_flags_reg =3D ACS_TEST_PCIE_FLAGS(pcie_type); + list_add_tail(&dev->bus_list, &bus->devices); + return dev; +} + +static void acs_build_fabric(struct kunit *test, struct acs_fabric *f) +{ + struct pci_bus *bus0, *bus1, *bus2, *bus3, *bus4; + struct pci_dev *rootport, *swup; + struct pci_host_bridge *host; + + host =3D kunit_kzalloc(test, sizeof(*host), GFP_KERNEL); + KUNIT_ASSERT_NOT_NULL(test, host); + + bus0 =3D acs_add_bus(test, NULL, NULL, 0, f); /* root bus */ + /* The Root Port doubles as the whitelisted host-bridge device. */ + rootport =3D acs_add_dev(test, bus0, PCI_DEVFN(0, 0), + PCI_EXP_TYPE_ROOT_PORT); + rootport->vendor =3D PCI_VENDOR_ID_GOOGLE; + rootport->device =3D 0x1234; + host->bus =3D bus0; + bus0->bridge =3D &host->dev; + + bus1 =3D acs_add_bus(test, bus0, rootport, 1, f); + swup =3D acs_add_dev(test, bus1, PCI_DEVFN(0, 0), PCI_EXP_TYPE_UPSTREAM); + + bus2 =3D acs_add_bus(test, bus1, swup, 2, f); + f->dn0 =3D acs_add_dev(test, bus2, PCI_DEVFN(0, 0), PCI_EXP_TYPE_DOWNSTRE= AM); + f->dn1 =3D acs_add_dev(test, bus2, PCI_DEVFN(1, 0), PCI_EXP_TYPE_DOWNSTRE= AM); + + bus3 =3D acs_add_bus(test, bus2, f->dn0, 3, f); + f->provider =3D acs_add_dev(test, bus3, PCI_DEVFN(0, 0), + PCI_EXP_TYPE_ENDPOINT); + + bus4 =3D acs_add_bus(test, bus2, f->dn1, 4, f); + f->client =3D acs_add_dev(test, bus4, PCI_DEVFN(0, 0), + PCI_EXP_TYPE_ENDPOINT); +} + +static enum pci_p2pdma_map_type acs_walk_map(struct acs_fabric *f) +{ + int dist; + + return calc_map_type_and_dist(f->provider, f->client, &dist, false); +} + +static void acs_walk_bus_addr_test(struct kunit *test) +{ + struct acs_fabric f =3D {}; + + acs_build_fabric(test, &f); + /* No ACS on the path: peer-to-peer is allowed directly. */ + KUNIT_EXPECT_EQ(test, acs_walk_map(&f), PCI_P2PDMA_MAP_BUS_ADDR); +} + +static void acs_walk_ec_violation_test(struct kunit *test) +{ + struct acs_fabric f =3D {}; + + acs_build_fabric(test, &f); + /* + * Downstream Port 0 has Egress Control enabled with the vector bit for + * the client's Downstream Port 1 set and Request Redirect clear: an ACS + * Violation, so the direct path is unusable and the request has to take + * the host-bridge route. + */ + f.dn0->acs_cap =3D 0x100; + f.dn0->acs_capabilities =3D PCI_ACS_EC | (64 << 8); + f.dn0_cfg.acs_ctrl =3D PCI_ACS_EC; + f.dn1_cfg.port =3D 5; + f.dn0_cfg.egress[0] =3D BIT(5); + + KUNIT_EXPECT_EQ(test, acs_walk_map(&f), + PCI_P2PDMA_MAP_THRU_HOST_BRIDGE); +} + +static void acs_walk_ec_vector_clear_test(struct kunit *test) +{ + struct acs_fabric f =3D {}; + + acs_build_fabric(test, &f); + /* Egress Control enabled but the vector bit for the peer is clear. */ + f.dn0->acs_cap =3D 0x100; + f.dn0->acs_capabilities =3D PCI_ACS_EC | (64 << 8); + f.dn0_cfg.acs_ctrl =3D PCI_ACS_EC; + f.dn1_cfg.port =3D 5; /* egress vector left all-zero */ + + KUNIT_EXPECT_EQ(test, acs_walk_map(&f), PCI_P2PDMA_MAP_BUS_ADDR); +} + +static void acs_walk_thru_host_bridge_test(struct kunit *test) +{ + struct acs_fabric f =3D {}; + + acs_build_fabric(test, &f); + /* Request Redirect set: traffic is redirected up to the host bridge. */ + f.dn0->acs_cap =3D 0x100; + f.dn0->acs_capabilities =3D PCI_ACS_RR; + f.dn0_cfg.acs_ctrl =3D PCI_ACS_RR; + + /* The Google root port is whitelisted, so the host-bridge path is OK. */ + KUNIT_EXPECT_EQ(test, acs_walk_map(&f), + PCI_P2PDMA_MAP_THRU_HOST_BRIDGE); +} + static struct kunit_case pci_acs_test_cases[] =3D { KUNIT_CASE_PARAM(pci_acs_p2pdma_decision_test, acs_decision_gen_params), KUNIT_CASE_PARAM(pci_acs_egress_port_valid_test, egress_valid_gen_params), @@ -402,6 +602,10 @@ static struct kunit_case pci_acs_test_cases[] =3D { KUNIT_CASE(acs_egress_target_pcie_bridge_test), KUNIT_CASE(acs_egress_target_other_bus_test), KUNIT_CASE(acs_egress_root_port_test), + KUNIT_CASE(acs_walk_bus_addr_test), + KUNIT_CASE(acs_walk_ec_violation_test), + KUNIT_CASE(acs_walk_ec_vector_clear_test), + KUNIT_CASE(acs_walk_thru_host_bridge_test), {} }; =20 --=20 2.55.0 From nobody Tue Sep 29 06:58:50 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7BE9543DA4B; Tue, 11 Aug 2026 09:32:17 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440738; cv=none; b=MU5Pf0Cl4VPFOxL8VnuX5cW/31SEtB27N/lVz5v6bWK3ajf270PnF1T856vL6xCinpdN4lqI3Z0HpUFFzmaDai+Ya4bZlncbx2iOSFygixeOd3y5XpbWBINd4PHNAY+tN9Bsgt/twdgTn2Hl4r64u5PGloJNEUmQFWgam8bxnA4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786440738; c=relaxed/simple; bh=xHX0745GTSixc57wJHnKKtkxs/DmMLtP43QRNGhBymw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=cCqE1lyqpeGpavjIfQnKaHZt1sZR7nLhg0BwvOn8sRr5/Wl1337z9COAey1AhlE2ZXJ7KpS+39E0+M6qWkVUABmDLf/HeQcWSxe+L7eN74D6pH9N6FMBc5GYKsUvHAZvXOS7kR9kZDSqJilRos+ZOzoJaU/a6zekEA0PD9P+1TY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=GqC3t24+; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="GqC3t24+" Received: by smtp.kernel.org (Postfix) with ESMTPSA id CD1B11F000E9; Tue, 11 Aug 2026 09:32:16 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786440737; bh=2+bUtGYXfogM9+SmfMCRcPaKXM9gF1dgNXfMyMgg5D4=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=GqC3t24+5c+5CyD2T6PGMgExE1yf9+At05jIQh6yOVnSm+DjMjpXpp69EB0eUDOU0 zrLieoqnnLNWzi3d480D/gQBPQpvdLtwTBY7rIjG469qkPKdcLQhwaIp/0akHXe+1R g6zBUjrLFmR9Jrkpb0A5Kz5Q26CGdY8pPxoXgfpQv20nzamedHXyYgQ/qs4g1iOBQf pozcVY3CauZ70ovjiH9C/GjHVoNOtxk/0nfrfdSkAKNpn4aZICLCIKvBKCPWjpDfja bIJi8L7EEVspdMceOXsEEBEKWGL0aDXATk7NIo8CgnZbCb4iNWkznVCfpqIXg6iIW5 USlR1Qwrlkjbw== From: Leon Romanovsky To: Bjorn Helgaas , Logan Gunthorpe , Chaitanya Kulkarni , Greg Kroah-Hartman , Jens Axboe , Alex Williamson , Leon Romanovsky , Ankit Agrawal , Jason Gunthorpe , Jonathan Corbet , Shuah Khan , "Joerg Roedel (AMD)" , Will Deacon , Robin Murphy Cc: linux-pci@vger.kernel.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, iommu@lists.linux.dev Subject: [PATCH v3 17/17] PCI: Add KUnit coverage for ACS isolation checks Date: Tue, 11 Aug 2026 12:30:59 +0300 Message-ID: <20260811-fix-p2p-acs-v3-17-efc488ee7c03@nvidia.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> References: <20260811-fix-p2p-acs-v3-0-efc488ee7c03@nvidia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" X-Mailer: b4 0.15-dev-18f8f Content-Transfer-Encoding: quoted-printable From: Leon Romanovsky Direct Translated P2P and Egress Control both let a peer request reach the peer without Request Redirect, and whether Request Redirect still isolates depends further on Translation Blocking and on the flags the caller requests. Firmware owns these bits, so the combinations are not reachable on a given machine. Drive pci_acs_flags_enabled() with a fake pci_ops supplying the ACS Control register and check each combination, in both scopes: an Untranslated-only caller such as pci_enable_pasid() is unaffected by Direct Translated P2P but still loses Request Redirect to Egress Control. Cover the two ways the check gives up as well: a device with no ACS capability, and a control register that cannot be read. The function is exposed under CONFIG_KUNIT via VISIBLE_IF_KUNIT. Reviewed-by: Logan Gunthorpe Assisted-by: Claude:claude-opus-5 Signed-off-by: Leon Romanovsky --- drivers/pci/pci.c | 6 +- drivers/pci/pci.h | 2 + drivers/pci/pci_acs_test.c | 181 +++++++++++++++++++++++++++++++++++++++++= ++++ 3 files changed, 187 insertions(+), 2 deletions(-) diff --git a/drivers/pci/pci.c b/drivers/pci/pci.c index d900fdb6f37d..7ee1fca60a12 100644 --- a/drivers/pci/pci.c +++ b/drivers/pci/pci.c @@ -3631,8 +3631,9 @@ int pci_acs_egress_ctrl_is_set(struct pci_dev *pdev, = struct pci_dev *target) } EXPORT_SYMBOL_IF_KUNIT(pci_acs_egress_ctrl_is_set); =20 -static bool pci_acs_flags_enabled(struct pci_dev *pdev, u16 acs_flags, - enum pci_acs_scope scope) +VISIBLE_IF_KUNIT +bool pci_acs_flags_enabled(struct pci_dev *pdev, u16 acs_flags, + enum pci_acs_scope scope) { int pos; u16 ctrl; @@ -3656,6 +3657,7 @@ static bool pci_acs_flags_enabled(struct pci_dev *pde= v, u16 acs_flags, =20 return (ctrl & acs_flags) =3D=3D acs_flags; } +EXPORT_SYMBOL_IF_KUNIT(pci_acs_flags_enabled); =20 /** * pci_acs_enabled - test ACS against required flags for a given device diff --git a/drivers/pci/pci.h b/drivers/pci/pci.h index c55e6ea7c9de..f8f9a15e411a 100644 --- a/drivers/pci/pci.h +++ b/drivers/pci/pci.h @@ -1093,6 +1093,8 @@ enum pci_acs_p2pdma_state { }; =20 #if IS_ENABLED(CONFIG_KUNIT) +bool pci_acs_flags_enabled(struct pci_dev *pdev, u16 acs_flags, + enum pci_acs_scope scope); bool pci_acs_egress_port_valid(u16 acs_caps, u8 target_port); enum pci_acs_p2pdma_state pci_acs_p2pdma_decision(u16 ctrl, bool has_targe= t, int egress); diff --git a/drivers/pci/pci_acs_test.c b/drivers/pci/pci_acs_test.c index 130605b91b47..11f145c89001 100644 --- a/drivers/pci/pci_acs_test.c +++ b/drivers/pci/pci_acs_test.c @@ -120,6 +120,184 @@ static void pci_acs_egress_port_valid_test(struct kun= it *test) c->expect); } =20 +/* + * pci_acs_flags_enabled(): Direct Translated P2P and Egress Control both = let a + * peer request reach the peer without Request Redirect, so neither may re= port + * isolation. Translation Blocking rejects a Translated Request before it= is + * routed, which restores the Request Redirect guarantee. A fake pci_ops + * supplies the ACS Control register. + */ + +/* Flags an IOMMU asks for; see REQ_ACS_FLAGS in drivers/iommu/iommu.c. */ +#define ACS_REQ_FLAGS (PCI_ACS_SV | PCI_ACS_RR | PCI_ACS_CR | PCI_ACS_UF) +#define ACS_ALL_CAPS (PCI_ACS_SV | PCI_ACS_TB | PCI_ACS_RR | PCI_ACS_CR | \ + PCI_ACS_UF | PCI_ACS_EC | PCI_ACS_DT) +#define ACS_TEST_CAP 0x100 +/* Flags pci_enable_pasid() asks for; see drivers/pci/ats.c. */ +#define ACS_PASID_FLAGS (PCI_ACS_RR | PCI_ACS_UF) + +struct acs_ctrl_cfg { + unsigned int devfn; + u16 cap; /* offset the ACS capability answers at */ + u16 ctrl; + bool fail_read; +}; + +static int acs_ctrl_read(struct pci_bus *bus, unsigned int devfn, + int where, int size, u32 *val) +{ + struct acs_ctrl_cfg *cfg =3D bus->sysdata; + + *val =3D 0; + if (cfg->fail_read) + return PCIBIOS_DEVICE_NOT_FOUND; + + if (devfn =3D=3D cfg->devfn && size =3D=3D 2 && + where =3D=3D cfg->cap + PCI_ACS_CTRL) + *val =3D cfg->ctrl; + return PCIBIOS_SUCCESSFUL; +} + +static int acs_ctrl_write(struct pci_bus *bus, unsigned int devfn, + int where, int size, u32 val) +{ + return PCIBIOS_SUCCESSFUL; +} + +static struct pci_ops acs_ctrl_ops =3D { + .read =3D acs_ctrl_read, + .write =3D acs_ctrl_write, +}; + +struct acs_isolation_case { + const char *desc; + u16 ctrl; /* ACS Control register */ + u16 req; /* flags the caller asks for */ + enum pci_acs_scope scope; /* Requests the answer must cover */ + bool expect; /* isolation reported? */ +}; + +static const struct acs_isolation_case acs_isolation_cases[] =3D { + { "plain_rr", ACS_REQ_FLAGS, ACS_REQ_FLAGS, PCI_ACS_SCOPE_ALL, true }, + /* Direct Translated P2P bypasses Request Redirect ... */ + { "dt", ACS_REQ_FLAGS | PCI_ACS_DT, ACS_REQ_FLAGS, + PCI_ACS_SCOPE_ALL, false }, + /* ... unless Translation Blocking rejects the Translated Request. */ + { "dt_tb", ACS_REQ_FLAGS | PCI_ACS_DT | PCI_ACS_TB, ACS_REQ_FLAGS, + PCI_ACS_SCOPE_ALL, true }, + { "tb_only", ACS_REQ_FLAGS | PCI_ACS_TB, ACS_REQ_FLAGS, + PCI_ACS_SCOPE_ALL, true }, + /* Egress Control can override Request Redirect as well. */ + { "ec", ACS_REQ_FLAGS | PCI_ACS_EC, ACS_REQ_FLAGS, + PCI_ACS_SCOPE_ALL, false }, + { "ec_dt_tb", ACS_REQ_FLAGS | PCI_ACS_EC | PCI_ACS_DT | PCI_ACS_TB, + ACS_REQ_FLAGS, PCI_ACS_SCOPE_ALL, false }, + /* Without Request Redirect requested, neither bit is consulted. */ + { "no_rr_dt", PCI_ACS_SV | PCI_ACS_CR | PCI_ACS_UF | PCI_ACS_DT, + PCI_ACS_SV | PCI_ACS_CR | PCI_ACS_UF, PCI_ACS_SCOPE_ALL, true }, + /* A control bit the caller asked for is simply missing. */ + { "rr_not_enabled", PCI_ACS_SV | PCI_ACS_CR | PCI_ACS_UF, ACS_REQ_FLAGS, + PCI_ACS_SCOPE_ALL, false }, + + /* + * An Untranslated-only caller such as pci_enable_pasid() is not + * affected by Direct Translated P2P, but is still affected by Egress + * Control, which acts on Untranslated peer Requests too. + */ + { "untrans/plain_rr", ACS_PASID_FLAGS, ACS_PASID_FLAGS, + PCI_ACS_SCOPE_UNTRANSLATED, true }, + { "untrans/dt", ACS_PASID_FLAGS | PCI_ACS_DT, ACS_PASID_FLAGS, + PCI_ACS_SCOPE_UNTRANSLATED, true }, + { "untrans/dt_tb", ACS_PASID_FLAGS | PCI_ACS_DT | PCI_ACS_TB, + ACS_PASID_FLAGS, PCI_ACS_SCOPE_UNTRANSLATED, true }, + { "untrans/ec", ACS_PASID_FLAGS | PCI_ACS_EC, ACS_PASID_FLAGS, + PCI_ACS_SCOPE_UNTRANSLATED, false }, + { "untrans/rr_not_enabled", PCI_ACS_UF, ACS_PASID_FLAGS, + PCI_ACS_SCOPE_UNTRANSLATED, false }, +}; + +static void acs_isolation_desc(const struct acs_isolation_case *c, char *d= esc) +{ + strscpy(desc, c->desc, KUNIT_PARAM_DESC_SIZE); +} + +KUNIT_ARRAY_PARAM(acs_isolation, acs_isolation_cases, acs_isolation_desc); + +static void pci_acs_flags_enabled_test(struct kunit *test) +{ + const struct acs_isolation_case *c =3D test->param_value; + struct acs_ctrl_cfg cfg =3D { .devfn =3D PCI_DEVFN(0, 0), + .cap =3D ACS_TEST_CAP, .ctrl =3D c->ctrl }; + struct pci_bus *bus =3D kunit_kzalloc(test, sizeof(*bus), GFP_KERNEL); + struct pci_dev *pdev =3D kunit_kzalloc(test, sizeof(*pdev), GFP_KERNEL); + + KUNIT_ASSERT_NOT_NULL(test, bus); + KUNIT_ASSERT_NOT_NULL(test, pdev); + + bus->ops =3D &acs_ctrl_ops; + bus->sysdata =3D &cfg; + + pdev->bus =3D bus; + pdev->devfn =3D cfg.devfn; + pdev->acs_cap =3D ACS_TEST_CAP; + pdev->acs_capabilities =3D ACS_ALL_CAPS; + + KUNIT_EXPECT_EQ(test, pci_acs_flags_enabled(pdev, c->req, c->scope), + c->expect); +} + +static bool acs_isolated(struct kunit *test, struct acs_ctrl_cfg *cfg, + u16 acs_cap, u16 acs_flags) +{ + struct pci_bus *bus =3D kunit_kzalloc(test, sizeof(*bus), GFP_KERNEL); + struct pci_dev *pdev =3D kunit_kzalloc(test, sizeof(*pdev), GFP_KERNEL); + + KUNIT_ASSERT_NOT_NULL(test, bus); + KUNIT_ASSERT_NOT_NULL(test, pdev); + + bus->ops =3D &acs_ctrl_ops; + bus->sysdata =3D cfg; + + pdev->bus =3D bus; + pdev->devfn =3D cfg->devfn; + pdev->acs_cap =3D acs_cap; + pdev->acs_capabilities =3D ACS_ALL_CAPS; + + return pci_acs_flags_enabled(pdev, acs_flags, PCI_ACS_SCOPE_ALL); +} + +/* + * Without an ACS capability there is no control register to consult. The + * fake answers at offset 0 here, so dropping the acs_cap guard would read= an + * isolating control word rather than nothing. + */ +static void pci_acs_flags_no_cap_test(struct kunit *test) +{ + struct acs_ctrl_cfg cfg =3D { .devfn =3D PCI_DEVFN(0, 0), .cap =3D 0, + .ctrl =3D ACS_REQ_FLAGS }; + + KUNIT_EXPECT_FALSE(test, acs_isolated(test, &cfg, 0, ACS_REQ_FLAGS)); +} + +/* + * An unreadable ACS Control register reads back as all ones, which satisf= ies + * any requested control. Ask without Request Redirect, so that neither + * pci_acs_rr_ineffective() nor the control word itself can deny isolation= and + * the read failure is the only thing left that can. + */ +static void pci_acs_flags_read_fails_test(struct kunit *test) +{ + u16 no_rr =3D ACS_REQ_FLAGS & ~PCI_ACS_RR; + struct acs_ctrl_cfg cfg =3D { .devfn =3D PCI_DEVFN(0, 0), + .cap =3D ACS_TEST_CAP, + .ctrl =3D ACS_REQ_FLAGS }; + + KUNIT_EXPECT_TRUE(test, acs_isolated(test, &cfg, ACS_TEST_CAP, no_rr)); + + cfg.fail_read =3D true; + KUNIT_EXPECT_FALSE(test, acs_isolated(test, &cfg, ACS_TEST_CAP, no_rr)); +} + /* * pci_acs_egress_ctrl_is_set(): drive the config-space reads with a fake = pci_ops * so the Egress Control Vector lookup is exercised without real hardware = -- @@ -591,6 +769,9 @@ static void acs_walk_thru_host_bridge_test(struct kunit= *test) static struct kunit_case pci_acs_test_cases[] =3D { KUNIT_CASE_PARAM(pci_acs_p2pdma_decision_test, acs_decision_gen_params), KUNIT_CASE_PARAM(pci_acs_egress_port_valid_test, egress_valid_gen_params), + KUNIT_CASE_PARAM(pci_acs_flags_enabled_test, acs_isolation_gen_params), + KUNIT_CASE(pci_acs_flags_no_cap_test), + KUNIT_CASE(pci_acs_flags_read_fails_test), KUNIT_CASE(acs_egress_vector_bit_set_test), KUNIT_CASE(acs_egress_vector_bit_clear_test), KUNIT_CASE(acs_egress_high_port_index_test), --=20 2.55.0