Documentation/PCI/endpoint/index.rst | 2 + .../PCI/endpoint/pci-dma-function.rst | 188 ++ Documentation/PCI/endpoint/pci-dma-howto.rst | 201 +++ MAINTAINERS | 1 + drivers/dma/dmaengine.c | 13 +- drivers/dma/dw-edma/dw-edma-core.c | 40 + drivers/dma/dw-edma/dw-edma-pcie.c | 394 +++- .../pci/controller/dwc/pcie-designware-ep.c | 184 +- drivers/pci/endpoint/functions/Kconfig | 13 + drivers/pci/endpoint/functions/Makefile | 1 + drivers/pci/endpoint/functions/pci-epf-dma.c | 1577 +++++++++++++++++ drivers/pci/endpoint/pci-epc-core.c | 70 + include/linux/dma/edma.h | 11 + include/linux/dmaengine.h | 20 + include/linux/pci-ep-dma.h | 175 ++ include/linux/pci-epc.h | 63 + 16 files changed, 2938 insertions(+), 15 deletions(-) create mode 100644 Documentation/PCI/endpoint/pci-dma-function.rst create mode 100644 Documentation/PCI/endpoint/pci-dma-howto.rst create mode 100644 drivers/pci/endpoint/functions/pci-epf-dma.c create mode 100644 include/linux/pci-ep-dma.h
Hi,
This is v7, the remaining patch set for PCI endpoint DMA.
Parts 2 and 3 were merged per Frank's suggestion.
This series defines an extensible endpoint DMA BAR metadata format and
controller-neutral EPC auxiliary-resource and delegation interfaces. It
then adds DesignWare eDMA/HDMA as the first implementation.
The infrastructure itself is not tied to dw-edma. Other endpoint DMA
controllers can use the same model by publishing their resources through
the EPC auxiliary-resource interface, adding their metadata-layout support
to pci-epf-dma, and providing the corresponding host-side driver.
The metadata lives in an endpoint BAR rather than a VSEC in PCIe extended
configuration space. This avoids requiring endpoint controllers to provide
writable configuration-space backing storage for an EPF-defined VSEC.
This series adds the host-side metadata parser, the pci-epf-dma
endpoint function driver, and documentation.
The endpoint function exposes selected endpoint-integrated DMA channels
as a separate PCI DMA controller function. The host-side dw-edma-pcie
driver discovers the BAR metadata and registers the exposed channels
with dmaengine. The endpoint function keeps the metadata BAR stable and
uses a separate DMA window BAR for resources that need dynamic subrange
mappings.
The endpoint function reserves each selected local channel through
dmaengine before asking the EPC backend to hand its hardware programming
ownership to the host. A new static channel-ID helper lets the endpoint
function select the exact channel without exposing a controller-specific
filter.
No fixed PCI ID is assigned. Users provide the PCI vendor/device ID
through configfs and bind dw-edma-pcie explicitly, for example with
driver_override.
One open question is how to support endpoint controllers with only one
PF. Keeping DMA in a separate EPF requires multi-function endpoint
support. Folding it into vNTB would work on single-function
controllers, but would also couple the two implementations. This series
keeps the separate EPF model.
I retested v7 with:
- small out-of-tree dmaengine client that uses delegated read channels
on SpacemiT K3 (HDMA).
- heavy load on delegated read channels on R-Car S4 (eDMA) with:
https://lore.kernel.org/r/20260810165136.2292436-1-den@valinux.co.jp/
Note: the entire read direction needs to be delegated because it's eDMA.
Part 1 v6 has landed in linux-next via dmaengine/next:
https://lore.kernel.org/all/20260721062815.4117887-1-den@valinux.co.jp/
v7 is based off of next-20260811.
Best regards,
Koichiro
---
Changes in v7:
- Merge parts 2 and 3. Part 3 was last posted as v5; no v6 was
sent. (Frank)
- Let DMA engine drivers assign static channel IDs. pci-epf-dma now
reserves exact channels itself, while EPC backends only perform the
hardware ownership handoff. (Frank)
- Put HOST_REQ in a dedicated host-request metadata word, avoiding a host
read-modify-write of endpoint-owned fields. (Frank)
- Restore automatic BAR selection when an EPF is rebound.
v6: https://lore.kernel.org/r/20260804033855.2115817-1-den@valinux.co.jp/
(part 2 only; part 3 was not posted as v6)
v5: https://lore.kernel.org/r/20260717050635.2145014-1-den@valinux.co.jp/
https://lore.kernel.org/r/20260717050953.2145851-1-den@valinux.co.jp/
v4: https://lore.kernel.org/r/20260710082156.2395844-1-den@valinux.co.jp/
https://lore.kernel.org/r/20260710082727.2397253-1-den@valinux.co.jp/
v3: https://lore.kernel.org/r/20260620170438.3756593-1-den@valinux.co.jp/
https://lore.kernel.org/r/20260620170844.3757241-1-den@valinux.co.jp/
v2: https://lore.kernel.org/r/20260525063129.3316894-1-den@valinux.co.jp/
https://lore.kernel.org/r/20260525063456.3317509-1-den@valinux.co.jp/
v1: https://lore.kernel.org/r/20260521063405.2842644-1-den@valinux.co.jp/
https://lore.kernel.org/r/20260521063638.2843021-1-den@valinux.co.jp/
Koichiro Den (10):
dmaengine: Allow drivers to assign static channel IDs
PCI: endpoint: Define endpoint DMA BAR metadata format
PCI: endpoint: Add DMA auxiliary resource metadata
PCI: endpoint: Add API to delegate EPC DMA channels to the host
dmaengine: dw-edma: Add channel delegation helpers
PCI: dwc: Implement endpoint DMA channel delegation
PCI: dwc: Expose endpoint DMA resources
dmaengine: dw-edma-pcie: Discover endpoint DMA metadata
PCI: endpoint: Add DMA endpoint function
Documentation: PCI: Add PCI DMA endpoint function documentation
Documentation/PCI/endpoint/index.rst | 2 +
.../PCI/endpoint/pci-dma-function.rst | 188 ++
Documentation/PCI/endpoint/pci-dma-howto.rst | 201 +++
MAINTAINERS | 1 +
drivers/dma/dmaengine.c | 13 +-
drivers/dma/dw-edma/dw-edma-core.c | 40 +
drivers/dma/dw-edma/dw-edma-pcie.c | 394 +++-
.../pci/controller/dwc/pcie-designware-ep.c | 184 +-
drivers/pci/endpoint/functions/Kconfig | 13 +
drivers/pci/endpoint/functions/Makefile | 1 +
drivers/pci/endpoint/functions/pci-epf-dma.c | 1577 +++++++++++++++++
drivers/pci/endpoint/pci-epc-core.c | 70 +
include/linux/dma/edma.h | 11 +
include/linux/dmaengine.h | 20 +
include/linux/pci-ep-dma.h | 175 ++
include/linux/pci-epc.h | 63 +
16 files changed, 2938 insertions(+), 15 deletions(-)
create mode 100644 Documentation/PCI/endpoint/pci-dma-function.rst
create mode 100644 Documentation/PCI/endpoint/pci-dma-howto.rst
create mode 100644 drivers/pci/endpoint/functions/pci-epf-dma.c
create mode 100644 include/linux/pci-ep-dma.h
base-commit: 5e6de6a2b522f659defacb1551d0465ba6ce13cf
--
2.51.0
Hello Koichiro,
On Thu, Aug 13, 2026 at 03:37:47PM +0900, Koichiro Den wrote:
> This is v7, the remaining patch set for PCI endpoint DMA.
> Parts 2 and 3 were merged per Frank's suggestion.
(snip)
> One open question is how to support endpoint controllers with only one
> PF. Keeping DMA in a separate EPF requires multi-function endpoint
> support. Folding it into vNTB would work on single-function
> controllers, but would also couple the two implementations. This series
> keeps the separate EPF model.
I see all the work you are putting in and I admire the effort.
This is now v7. I think it is time that we close the open question by
waiting for a reply from the PCI endpoint maintainers' opinion on the
design before continuing. (I am not a PCI endpoint maintainer.)
I understand that you want a common DMA abstraction, that can represent
different (embedded) DMA controllers on the endpoint side.
But if vNTB is the only consumer of this, then why not simply embed this
DMA abstraction in some BAR exposed by the vNTB EPF?
Looking at the host side driver that goes with the (v)NTB driver:
drivers/ntb/hw/epf/ntb_hw_epf.c
The BAR layouts are hard coded, and it only supports three different
layouts. Would it not be possible to add a fourth layout that has the
DMA abstraction somewhere in one of the BARs? ('BAR_DMA' ?)
Right now, I wonder if it is not a bit premature optimization to create a
DMA EPF, if vNTB will be the only (ever?) user.
I didn't follow all the details, but I know that you want to control the
DMA controller on the endpoint from the host side. Is this really a
normal use case outside of vNTB? I would imagine that most endpoints
will read some ring buffer of descriptors, perform some validation on
those descriptors, and then decide if it will do DMA to/from the host.
If the host side driver want to make use of your "generic DMA registers",
then you are basically creating another DMA controller? Shouldn't you
then create a new host side driver specifically for this "generic DMA
controller"? It would be nice if you could explain a bit better why you are
bothering to create a "generic DMA layout", but then you are reusing the
dw-edma-pcie driver. This seems a bit weird to me.
Right now you seem to "unpack" the "generic DMA layout" in a dw-edma specific
function: dw_edma_pcie_validate_ep_dma_metadata().
If you want this encapsulation, shouldn't the de-encapsulation be done by a
host side "DMA EPF" driver, and then this generic driver will then call
e.g. dw_edma_probe(). (Seems wrong to add de-encapsulation code in dw-edma
for your own made up format. And then all DMA drivers would need to do this
same de-encapsulation.)
Currently, I know R-Car 4 has an EPC controller that supports multi-function,
but I personally don't know any other. If you could embed your DMA abstraction
somewhere in one of the vNTB BARs, that would avoid the multi-function problem,
so your solution would not be limited to EPC controllers that only supports
multi-function.
Kind regards,
Niklas
On Thu, Aug 13, 2026 at 01:46:07PM +0200, Niklas Cassel wrote:
> Hello Koichiro,
>
> On Thu, Aug 13, 2026 at 03:37:47PM +0900, Koichiro Den wrote:
> > This is v7, the remaining patch set for PCI endpoint DMA.
> > Parts 2 and 3 were merged per Frank's suggestion.
>
> (snip)
>
> > One open question is how to support endpoint controllers with only one
> > PF. Keeping DMA in a separate EPF requires multi-function endpoint
> > support. Folding it into vNTB would work on single-function
> > controllers, but would also couple the two implementations. This series
> > keeps the separate EPF model.
>
> I see all the work you are putting in and I admire the effort.
>
> This is now v7. I think it is time that we close the open question by
> waiting for a reply from the PCI endpoint maintainers' opinion on the
> design before continuing. (I am not a PCI endpoint maintainer.)
>
> I understand that you want a common DMA abstraction, that can represent
> different (embedded) DMA controllers on the endpoint side.
>
> But if vNTB is the only consumer of this, then why not simply embed this
> DMA abstraction in some BAR exposed by the vNTB EPF?
>
> Looking at the host side driver that goes with the (v)NTB driver:
> drivers/ntb/hw/epf/ntb_hw_epf.c
>
> The BAR layouts are hard coded, and it only supports three different
> layouts. Would it not be possible to add a fourth layout that has the
> DMA abstraction somewhere in one of the BARs? ('BAR_DMA' ?)
>
>
> Right now, I wonder if it is not a bit premature optimization to create a
> DMA EPF, if vNTB will be the only (ever?) user.
>
Yeah, I feel the same. I haven't seen an usecase to program the DMA controller
from the host outside of vNTB. This feature is supported mostly because it
exists in hardware and someone wants to tick the checkbox.
Though, I'm not against doing it within vNTB as Niklas suggested, but
generalising it in the form of a new EPF driver just for the sake of a single
driver sounds like an overkill and maintenance burden.
Sorry for saying this in v7. I've been meaning to say it, but somehow ended up
procrastinating too much.
- Mani
--
மணிவண்ணன் சதாசிவம்
On Thu, Aug 13, 2026 at 02:50:59PM +0200, Manivannan Sadhasivam wrote:
> On Thu, Aug 13, 2026 at 01:46:07PM +0200, Niklas Cassel wrote:
> > Hello Koichiro,
> >
> > On Thu, Aug 13, 2026 at 03:37:47PM +0900, Koichiro Den wrote:
> > > This is v7, the remaining patch set for PCI endpoint DMA.
> > > Parts 2 and 3 were merged per Frank's suggestion.
> >
> > (snip)
> >
> > > One open question is how to support endpoint controllers with only one
> > > PF. Keeping DMA in a separate EPF requires multi-function endpoint
> > > support. Folding it into vNTB would work on single-function
> > > controllers, but would also couple the two implementations. This series
> > > keeps the separate EPF model.
> >
> > I see all the work you are putting in and I admire the effort.
> >
> > This is now v7. I think it is time that we close the open question by
> > waiting for a reply from the PCI endpoint maintainers' opinion on the
> > design before continuing. (I am not a PCI endpoint maintainer.)
> >
> > I understand that you want a common DMA abstraction, that can represent
> > different (embedded) DMA controllers on the endpoint side.
> >
> > But if vNTB is the only consumer of this, then why not simply embed this
> > DMA abstraction in some BAR exposed by the vNTB EPF?
> >
> > Looking at the host side driver that goes with the (v)NTB driver:
> > drivers/ntb/hw/epf/ntb_hw_epf.c
> >
> > The BAR layouts are hard coded, and it only supports three different
> > layouts. Would it not be possible to add a fourth layout that has the
> > DMA abstraction somewhere in one of the BARs? ('BAR_DMA' ?)
> >
> >
> > Right now, I wonder if it is not a bit premature optimization to create a
> > DMA EPF, if vNTB will be the only (ever?) user.
> >
>
> Yeah, I feel the same. I haven't seen an usecase to program the DMA controller
> from the host outside of vNTB. This feature is supported mostly because it
> exists in hardware and someone wants to tick the checkbox.
>
> Though, I'm not against doing it within vNTB as Niklas suggested, but
> generalising it in the form of a new EPF driver just for the sake of a single
> driver sounds like an overkill and maintenance burden.
>
> Sorry for saying this in v7. I've been meaning to say it, but somehow ended up
> procrastinating too much.
No bother. Since Frank has given a lot of feedback on this series, I'd also like
to hear his view.
I'm fine with either direction, and can revisit the earlier vNTB-embedded
approach:
https://lore.kernel.org/r/sn67hi7kljh7cgmgodatb3naz2astlaklqfobdbxyyzgoohxqb@4nnetbhqwba4/
Best regards,
Koichiro
>
> - Mani
>
> --
> மணிவண்ணன் சதாசிவம்
On Thu, Aug 13, 2026 at 11:15:31PM +0900, Koichiro Den wrote:
> On Thu, Aug 13, 2026 at 02:50:59PM +0200, Manivannan Sadhasivam wrote:
> > On Thu, Aug 13, 2026 at 01:46:07PM +0200, Niklas Cassel wrote:
> > > Hello Koichiro,
> > >
> > > On Thu, Aug 13, 2026 at 03:37:47PM +0900, Koichiro Den wrote:
> > > > This is v7, the remaining patch set for PCI endpoint DMA.
> > > > Parts 2 and 3 were merged per Frank's suggestion.
> > >
> > > (snip)
> > >
> > > > One open question is how to support endpoint controllers with only one
> > > > PF. Keeping DMA in a separate EPF requires multi-function endpoint
> > > > support. Folding it into vNTB would work on single-function
> > > > controllers, but would also couple the two implementations. This series
> > > > keeps the separate EPF model.
> > >
> > > I see all the work you are putting in and I admire the effort.
> > >
> > > This is now v7. I think it is time that we close the open question by
> > > waiting for a reply from the PCI endpoint maintainers' opinion on the
> > > design before continuing. (I am not a PCI endpoint maintainer.)
> > >
> > > I understand that you want a common DMA abstraction, that can represent
> > > different (embedded) DMA controllers on the endpoint side.
> > >
> > > But if vNTB is the only consumer of this, then why not simply embed this
> > > DMA abstraction in some BAR exposed by the vNTB EPF?
> > >
> > > Looking at the host side driver that goes with the (v)NTB driver:
> > > drivers/ntb/hw/epf/ntb_hw_epf.c
> > >
> > > The BAR layouts are hard coded, and it only supports three different
> > > layouts. Would it not be possible to add a fourth layout that has the
> > > DMA abstraction somewhere in one of the BARs? ('BAR_DMA' ?)
> > >
> > >
> > > Right now, I wonder if it is not a bit premature optimization to create a
> > > DMA EPF, if vNTB will be the only (ever?) user.
> > >
> >
> > Yeah, I feel the same. I haven't seen an usecase to program the DMA controller
> > from the host outside of vNTB. This feature is supported mostly because it
> > exists in hardware and someone wants to tick the checkbox.
> >
> > Though, I'm not against doing it within vNTB as Niklas suggested, but
> > generalising it in the form of a new EPF driver just for the sake of a single
> > driver sounds like an overkill and maintenance burden.
> >
> > Sorry for saying this in v7. I've been meaning to say it, but somehow ended up
> > procrastinating too much.
>
> No bother. Since Frank has given a lot of feedback on this series, I'd also like
> to hear his view.
>
> I'm fine with either direction, and can revisit the earlier vNTB-embedded
> approach:
> https://lore.kernel.org/r/sn67hi7kljh7cgmgodatb3naz2astlaklqfobdbxyyzgoohxqb@4nnetbhqwba4/
One of the important value is test dw-edma-pcie.c, which generally depend
on some fpga hardware. If there are epf driver work as fpga hardware, it
will help cover edma remote user case. So more user can test it.
Of course, this implement are over complex. I suggest update dma-engine
chan_id to support static allocate, which also need be fixed because
some drivers have such dependence, anyway need be fixed. After this fix,
this patches will become simpler.
VNTB case, it'd better put such informaiton into one BARs and work on
single-function.
I suggest split two things
1 - create simple epf driver to test dw-edma-pcie.c.
2 - vntb support DMA.
of course, if shared efforts, it will be great.
Frank
>
> Best regards,
> Koichiro
>
> >
> > - Mani
> >
> > --
> > மணிவண்ணன் சதாசிவம்
On Thu, Aug 13, 2026 at 10:59:58AM -0500, Frank Li wrote:
> On Thu, Aug 13, 2026 at 11:15:31PM +0900, Koichiro Den wrote:
> > On Thu, Aug 13, 2026 at 02:50:59PM +0200, Manivannan Sadhasivam wrote:
> > > On Thu, Aug 13, 2026 at 01:46:07PM +0200, Niklas Cassel wrote:
> > > > Hello Koichiro,
> > > >
> > > > On Thu, Aug 13, 2026 at 03:37:47PM +0900, Koichiro Den wrote:
> > > > > This is v7, the remaining patch set for PCI endpoint DMA.
> > > > > Parts 2 and 3 were merged per Frank's suggestion.
> > > >
> > > > (snip)
> > > >
> > > > > One open question is how to support endpoint controllers with only one
> > > > > PF. Keeping DMA in a separate EPF requires multi-function endpoint
> > > > > support. Folding it into vNTB would work on single-function
> > > > > controllers, but would also couple the two implementations. This series
> > > > > keeps the separate EPF model.
> > > >
> > > > I see all the work you are putting in and I admire the effort.
> > > >
> > > > This is now v7. I think it is time that we close the open question by
> > > > waiting for a reply from the PCI endpoint maintainers' opinion on the
> > > > design before continuing. (I am not a PCI endpoint maintainer.)
> > > >
> > > > I understand that you want a common DMA abstraction, that can represent
> > > > different (embedded) DMA controllers on the endpoint side.
> > > >
> > > > But if vNTB is the only consumer of this, then why not simply embed this
> > > > DMA abstraction in some BAR exposed by the vNTB EPF?
> > > >
> > > > Looking at the host side driver that goes with the (v)NTB driver:
> > > > drivers/ntb/hw/epf/ntb_hw_epf.c
> > > >
> > > > The BAR layouts are hard coded, and it only supports three different
> > > > layouts. Would it not be possible to add a fourth layout that has the
> > > > DMA abstraction somewhere in one of the BARs? ('BAR_DMA' ?)
> > > >
> > > >
> > > > Right now, I wonder if it is not a bit premature optimization to create a
> > > > DMA EPF, if vNTB will be the only (ever?) user.
> > > >
> > >
> > > Yeah, I feel the same. I haven't seen an usecase to program the DMA controller
> > > from the host outside of vNTB. This feature is supported mostly because it
> > > exists in hardware and someone wants to tick the checkbox.
> > >
> > > Though, I'm not against doing it within vNTB as Niklas suggested, but
> > > generalising it in the form of a new EPF driver just for the sake of a single
> > > driver sounds like an overkill and maintenance burden.
> > >
> > > Sorry for saying this in v7. I've been meaning to say it, but somehow ended up
> > > procrastinating too much.
> >
> > No bother. Since Frank has given a lot of feedback on this series, I'd also like
> > to hear his view.
> >
> > I'm fine with either direction, and can revisit the earlier vNTB-embedded
> > approach:
> > https://lore.kernel.org/r/sn67hi7kljh7cgmgodatb3naz2astlaklqfobdbxyyzgoohxqb@4nnetbhqwba4/
>
> One of the important value is test dw-edma-pcie.c, which generally depend
> on some fpga hardware. If there are epf driver work as fpga hardware, it
> will help cover edma remote user case. So more user can test it.
>
Testing is one thing, but using is what matters. Are there any products or
use case based on remote eDMA? Or even dw-edma-pcie.c?
Most of the time, dw-edma-pcie.c driver feels like a dead code to me which no
one actively uses or tests. So unless a real justification is put forward, I
don't want to add a whole new EPF driver for it.
- Mani
--
மணிவண்ணன் சதாசிவம்
On Fri, Aug 14, 2026 at 07:27:10AM +0200, Manivannan Sadhasivam wrote:
> On Thu, Aug 13, 2026 at 10:59:58AM -0500, Frank Li wrote:
> > On Thu, Aug 13, 2026 at 11:15:31PM +0900, Koichiro Den wrote:
> > > On Thu, Aug 13, 2026 at 02:50:59PM +0200, Manivannan Sadhasivam wrote:
> > > > On Thu, Aug 13, 2026 at 01:46:07PM +0200, Niklas Cassel wrote:
> > > > > Hello Koichiro,
> > > > >
> > > > > On Thu, Aug 13, 2026 at 03:37:47PM +0900, Koichiro Den wrote:
> > > > > > This is v7, the remaining patch set for PCI endpoint DMA.
> > > > > > Parts 2 and 3 were merged per Frank's suggestion.
> > > > >
> > > > > (snip)
> > > > >
> > > > > > One open question is how to support endpoint controllers with only one
> > > > > > PF. Keeping DMA in a separate EPF requires multi-function endpoint
> > > > > > support. Folding it into vNTB would work on single-function
> > > > > > controllers, but would also couple the two implementations. This series
> > > > > > keeps the separate EPF model.
> > > > >
> > > > > I see all the work you are putting in and I admire the effort.
> > > > >
> > > > > This is now v7. I think it is time that we close the open question by
> > > > > waiting for a reply from the PCI endpoint maintainers' opinion on the
> > > > > design before continuing. (I am not a PCI endpoint maintainer.)
> > > > >
> > > > > I understand that you want a common DMA abstraction, that can represent
> > > > > different (embedded) DMA controllers on the endpoint side.
> > > > >
> > > > > But if vNTB is the only consumer of this, then why not simply embed this
> > > > > DMA abstraction in some BAR exposed by the vNTB EPF?
> > > > >
> > > > > Looking at the host side driver that goes with the (v)NTB driver:
> > > > > drivers/ntb/hw/epf/ntb_hw_epf.c
> > > > >
> > > > > The BAR layouts are hard coded, and it only supports three different
> > > > > layouts. Would it not be possible to add a fourth layout that has the
> > > > > DMA abstraction somewhere in one of the BARs? ('BAR_DMA' ?)
> > > > >
> > > > >
> > > > > Right now, I wonder if it is not a bit premature optimization to create a
> > > > > DMA EPF, if vNTB will be the only (ever?) user.
> > > > >
> > > >
> > > > Yeah, I feel the same. I haven't seen an usecase to program the DMA controller
> > > > from the host outside of vNTB. This feature is supported mostly because it
> > > > exists in hardware and someone wants to tick the checkbox.
> > > >
> > > > Though, I'm not against doing it within vNTB as Niklas suggested, but
> > > > generalising it in the form of a new EPF driver just for the sake of a single
> > > > driver sounds like an overkill and maintenance burden.
> > > >
> > > > Sorry for saying this in v7. I've been meaning to say it, but somehow ended up
> > > > procrastinating too much.
> > >
> > > No bother. Since Frank has given a lot of feedback on this series, I'd also like
> > > to hear his view.
> > >
> > > I'm fine with either direction, and can revisit the earlier vNTB-embedded
> > > approach:
> > > https://lore.kernel.org/r/sn67hi7kljh7cgmgodatb3naz2astlaklqfobdbxyyzgoohxqb@4nnetbhqwba4/
> >
> > One of the important value is test dw-edma-pcie.c, which generally depend
> > on some fpga hardware. If there are epf driver work as fpga hardware, it
> > will help cover edma remote user case. So more user can test it.
> >
>
> Testing is one thing, but using is what matters. Are there any products or
> use case based on remote eDMA? Or even dw-edma-pcie.c?
My end goal for this work is this series:
https://lore.kernel.org/r/20260810165136.2292436-1-den@valinux.co.jp/
It now depends on the PCI DMA EPF. The host side controls the endpoint eDMA
through dw-edma-pcie for one direction. This is for an industrial use case, not
just testing.
The resulting ntb_netdev/ntb_transport improvement is substantial:
(unit: Gbps) (UL=EP->RC, DL=RC->EP)
UL UDP DL UDP UL TCP DL TCP
------- ------ ------- ------ ------
Before ~0.6 ~0.6 ~0.6 ~0.6
After ~19.5 ~17.3 ~12.3 ~10.8
(On R-Car S4, PCIe Gen4 x2, controller IP v5.20a, eDMA)
Best regards,
Koichiro
>
> Most of the time, dw-edma-pcie.c driver feels like a dead code to me which no
> one actively uses or tests. So unless a real justification is put forward, I
> don't want to add a whole new EPF driver for it.
>
> - Mani
>
> --
> மணிவண்ணன் சதாசிவம்
On Fri, Aug 14, 2026 at 02:57:12PM +0900, Koichiro Den wrote: > On Fri, Aug 14, 2026 at 07:27:10AM +0200, Manivannan Sadhasivam wrote: > > > > Testing is one thing, but using is what matters. Are there any products or > > use case based on remote eDMA? Or even dw-edma-pcie.c? > > My end goal for this work is this series: > > https://lore.kernel.org/r/20260810165136.2292436-1-den@valinux.co.jp/ > > It now depends on the PCI DMA EPF. The host side controls the endpoint eDMA > through dw-edma-pcie for one direction. This is for an industrial use case, not > just testing. > > The resulting ntb_netdev/ntb_transport improvement is substantial: > > (unit: Gbps) (UL=EP->RC, DL=RC->EP) > > UL UDP DL UDP UL TCP DL TCP > ------- ------ ------- ------ ------ > Before ~0.6 ~0.6 ~0.6 ~0.6 > After ~19.5 ~17.3 ~12.3 ~10.8 > > (On R-Car S4, PCIe Gen4 x2, controller IP v5.20a, eDMA) > You have an industrial use case, and your performance numbers show that remote eDMA can bring great performance gains for your use case. I don't think anyone is arguing about that. At least to me, the question is if you need a new PCI EPF driver to implement the code for this use case. I think the answer is: No, it is not strictly needed. You can extend vNTB EPF to support your use case. (As that was your original approach.) The question how you should test remote eDMA is a different question IMO. I'm not an expert, but from a testing perspective, does it really matter if it is the host or the endpoint itself that programs the eDMA hardware? I understand that you gain performance by having the host program the eDMA directly. But.. from a eDMA hardware verification standpoint, does it really matter which side that writes the eDMA registers? You should be able to test both dma directions, regardless of which side programs the eDMA hardware, no? I guess what you mentioned earlier, that the existing pci-epf-test tests are not pushing sufficient concurrent data to trigger certain driver bugs when multiple eDMA channels are used. I guess you could have a test suite that does whatever you did to uncover these bugs... vNTB + iperf? But I guess it could also be interesting to add tests that push more data concurrently, such that multiple eDMA channels are used. To me, that is basically what dmatest was designed for... Yes, we know that dmatest is currently not a great fit for DWC eDMA, because dmatest uses different dmaengine APIs. I think Vinod is best qualified to answer this question, but I guess the answer is either: A) Extend dmatest so that it can use the dmaengine APIs to fit DWC eDMA. or B) Write a copy of dmatest that is tailored to hardware that uses the dmaengine APIs in a similar way as DWC eDMA requires. Kind regards, Niklas
On Fri, Aug 21, 2026 at 05:55:17PM +0200, Niklas Cassel wrote: > On Fri, Aug 14, 2026 at 02:57:12PM +0900, Koichiro Den wrote: > > On Fri, Aug 14, 2026 at 07:27:10AM +0200, Manivannan Sadhasivam wrote: > > > > > > Testing is one thing, but using is what matters. Are there any products or > > > use case based on remote eDMA? Or even dw-edma-pcie.c? > > > > My end goal for this work is this series: > > > > https://lore.kernel.org/r/20260810165136.2292436-1-den@valinux.co.jp/ > > > > It now depends on the PCI DMA EPF. The host side controls the endpoint eDMA > > through dw-edma-pcie for one direction. This is for an industrial use case, not > > just testing. > > > > The resulting ntb_netdev/ntb_transport improvement is substantial: > > > > (unit: Gbps) (UL=EP->RC, DL=RC->EP) > > > > UL UDP DL UDP UL TCP DL TCP > > ------- ------ ------- ------ ------ > > Before ~0.6 ~0.6 ~0.6 ~0.6 > > After ~19.5 ~17.3 ~12.3 ~10.8 > > > > (On R-Car S4, PCIe Gen4 x2, controller IP v5.20a, eDMA) > > > > You have an industrial use case, and your performance numbers show that > remote eDMA can bring great performance gains for your use case. > > I don't think anyone is arguing about that. > > At least to me, the question is if you need a new PCI EPF driver to implement > the code for this use case. I think the answer is: No, it is not strictly > needed. You can extend vNTB EPF to support your use case. > (As that was your original approach.) > > > > The question how you should test remote eDMA is a different question IMO. > I'm not an expert, but from a testing perspective, does it really matter if > it is the host or the endpoint itself that programs the eDMA hardware? > > I understand that you gain performance by having the host program the eDMA > directly. But.. from a eDMA hardware verification standpoint, does it really > matter which side that writes the eDMA registers? > You should be able to test both dma directions, regardless of which side > programs the eDMA hardware, no? > > I guess what you mentioned earlier, that the existing pci-epf-test tests > are not pushing sufficient concurrent data to trigger certain driver bugs > when multiple eDMA channels are used. > > I guess you could have a test suite that does whatever you did to uncover > these bugs... vNTB + iperf? But I guess it could also be interesting to add > tests that push more data concurrently, such that multiple eDMA channels are > used. To me, that is basically what dmatest was designed for... > > Yes, we know that dmatest is currently not a great fit for DWC eDMA, because > dmatest uses different dmaengine APIs. > > I think Vinod is best qualified to answer this question, but I guess the answer > is either: > A) Extend dmatest so that it can use the dmaengine APIs to fit DWC eDMA. > or > B) Write a copy of dmatest that is tailored to hardware that uses the dmaengine > APIs in a similar way as DWC eDMA requires. As to dw-edma performance testing of dw-edma, we already discussed it here: https://lore.kernel.org/r/tau5svk3bcatzeapqeb6mun7dxi4ifk56g5ltkk366ljozjzit@vepneiac3f26/ I had already said in the first mail that dmatest did not seem to fit. Vinod agreed, and Mani suggested an eDMA/HDMA-specific kselftest instead. Which side programs the hardware may not matter for IP testing, but it does for driver coverage, specifically dw-edma-pcie. Whether that path should be exposed through a standalone EPF or vNTB should be a separate question, I suppose. P.S. I almost ready to start reviving the vNTB-embedded approach: https://lore.kernel.org/all/3ef3b7wdwpf364teperxcjc2leycxwke77cejjfzd5w4pnbmik@vlcbizvbxv4m/ but am waiting to hear back there in case the maintainers see something differently. Best regards, Koichiro > > > Kind regards, > Niklas
On Thu, Aug 13, 2026 at 10:59:58AM -0500, Frank Li wrote:
> On Thu, Aug 13, 2026 at 11:15:31PM +0900, Koichiro Den wrote:
> > On Thu, Aug 13, 2026 at 02:50:59PM +0200, Manivannan Sadhasivam wrote:
> > > On Thu, Aug 13, 2026 at 01:46:07PM +0200, Niklas Cassel wrote:
> > > > Hello Koichiro,
> > > >
> > > > On Thu, Aug 13, 2026 at 03:37:47PM +0900, Koichiro Den wrote:
> > > > > This is v7, the remaining patch set for PCI endpoint DMA.
> > > > > Parts 2 and 3 were merged per Frank's suggestion.
> > > >
> > > > (snip)
> > > >
> > > > > One open question is how to support endpoint controllers with only one
> > > > > PF. Keeping DMA in a separate EPF requires multi-function endpoint
> > > > > support. Folding it into vNTB would work on single-function
> > > > > controllers, but would also couple the two implementations. This series
> > > > > keeps the separate EPF model.
> > > >
> > > > I see all the work you are putting in and I admire the effort.
> > > >
> > > > This is now v7. I think it is time that we close the open question by
> > > > waiting for a reply from the PCI endpoint maintainers' opinion on the
> > > > design before continuing. (I am not a PCI endpoint maintainer.)
> > > >
> > > > I understand that you want a common DMA abstraction, that can represent
> > > > different (embedded) DMA controllers on the endpoint side.
> > > >
> > > > But if vNTB is the only consumer of this, then why not simply embed this
> > > > DMA abstraction in some BAR exposed by the vNTB EPF?
> > > >
> > > > Looking at the host side driver that goes with the (v)NTB driver:
> > > > drivers/ntb/hw/epf/ntb_hw_epf.c
> > > >
> > > > The BAR layouts are hard coded, and it only supports three different
> > > > layouts. Would it not be possible to add a fourth layout that has the
> > > > DMA abstraction somewhere in one of the BARs? ('BAR_DMA' ?)
> > > >
> > > >
> > > > Right now, I wonder if it is not a bit premature optimization to create a
> > > > DMA EPF, if vNTB will be the only (ever?) user.
> > > >
> > >
> > > Yeah, I feel the same. I haven't seen an usecase to program the DMA controller
> > > from the host outside of vNTB. This feature is supported mostly because it
> > > exists in hardware and someone wants to tick the checkbox.
> > >
> > > Though, I'm not against doing it within vNTB as Niklas suggested, but
> > > generalising it in the form of a new EPF driver just for the sake of a single
> > > driver sounds like an overkill and maintenance burden.
> > >
> > > Sorry for saying this in v7. I've been meaning to say it, but somehow ended up
> > > procrastinating too much.
> >
> > No bother. Since Frank has given a lot of feedback on this series, I'd also like
> > to hear his view.
> >
> > I'm fine with either direction, and can revisit the earlier vNTB-embedded
> > approach:
> > https://lore.kernel.org/r/sn67hi7kljh7cgmgodatb3naz2astlaklqfobdbxyyzgoohxqb@4nnetbhqwba4/
>
> One of the important value is test dw-edma-pcie.c, which generally depend
> on some fpga hardware. If there are epf driver work as fpga hardware, it
> will help cover edma remote user case. So more user can test it.
>
> Of course, this implement are over complex. I suggest update dma-engine
> chan_id to support static allocate, which also need be fixed because
> some drivers have such dependence, anyway need be fixed. After this fix,
> this patches will become simpler.
>
> VNTB case, it'd better put such informaiton into one BARs and work on
> single-function.
>
> I suggest split two things
>
> 1 - create simple epf driver to test dw-edma-pcie.c.
> 2 - vntb support DMA.
>
> of course, if shared efforts, it will be great.
Thanks for the feedback. v7 patch 1 adds that exact static chan_id support, so
you mean this v7 patch 2 ~ 9 can still be simplified enough to make the
maintenance burden acceptable, right? Please correct me if I misunderstood.
I think a large part of the remaining complexity comes from discovery. Ideally,
the EPF would create a VSEC like existing supported hardware, and put the DMA
description there. But I could not find a generic way for an EPF to provide
writable configuration space backing for such a VSEC on a DWC PCIe controller in
EP mode. That is why this series (sadly) ended up with a BAR protocol. BAR
subrange mapping use might look complicated, but is still needed on systems
without a fixed DMA BAR, but the discovery part would otherwise be much simpler.
I may be missing something I could simplify further.
Best regards,
Koichiro
>
> Frank
>
> >
> > Best regards,
> > Koichiro
> >
> > >
> > > - Mani
> > >
> > > --
> > > மணிவண்ணன் சதாசிவம்
On Fri, Aug 14, 2026 at 02:04:17AM +0900, Koichiro Den wrote:
> On Thu, Aug 13, 2026 at 10:59:58AM -0500, Frank Li wrote:
> > On Thu, Aug 13, 2026 at 11:15:31PM +0900, Koichiro Den wrote:
> > > On Thu, Aug 13, 2026 at 02:50:59PM +0200, Manivannan Sadhasivam wrote:
> > > > On Thu, Aug 13, 2026 at 01:46:07PM +0200, Niklas Cassel wrote:
> > > > > Hello Koichiro,
> > > > >
> > > > > On Thu, Aug 13, 2026 at 03:37:47PM +0900, Koichiro Den wrote:
> > > > > > This is v7, the remaining patch set for PCI endpoint DMA.
> > > > > > Parts 2 and 3 were merged per Frank's suggestion.
> > > > >
> > > > > (snip)
> > > > >
> > > > > > One open question is how to support endpoint controllers with only one
> > > > > > PF. Keeping DMA in a separate EPF requires multi-function endpoint
> > > > > > support. Folding it into vNTB would work on single-function
> > > > > > controllers, but would also couple the two implementations. This series
> > > > > > keeps the separate EPF model.
> > > > >
> > > > > I see all the work you are putting in and I admire the effort.
> > > > >
> > > > > This is now v7. I think it is time that we close the open question by
> > > > > waiting for a reply from the PCI endpoint maintainers' opinion on the
> > > > > design before continuing. (I am not a PCI endpoint maintainer.)
> > > > >
> > > > > I understand that you want a common DMA abstraction, that can represent
> > > > > different (embedded) DMA controllers on the endpoint side.
> > > > >
> > > > > But if vNTB is the only consumer of this, then why not simply embed this
> > > > > DMA abstraction in some BAR exposed by the vNTB EPF?
> > > > >
> > > > > Looking at the host side driver that goes with the (v)NTB driver:
> > > > > drivers/ntb/hw/epf/ntb_hw_epf.c
> > > > >
> > > > > The BAR layouts are hard coded, and it only supports three different
> > > > > layouts. Would it not be possible to add a fourth layout that has the
> > > > > DMA abstraction somewhere in one of the BARs? ('BAR_DMA' ?)
> > > > >
> > > > >
> > > > > Right now, I wonder if it is not a bit premature optimization to create a
> > > > > DMA EPF, if vNTB will be the only (ever?) user.
> > > > >
> > > >
> > > > Yeah, I feel the same. I haven't seen an usecase to program the DMA controller
> > > > from the host outside of vNTB. This feature is supported mostly because it
> > > > exists in hardware and someone wants to tick the checkbox.
> > > >
> > > > Though, I'm not against doing it within vNTB as Niklas suggested, but
> > > > generalising it in the form of a new EPF driver just for the sake of a single
> > > > driver sounds like an overkill and maintenance burden.
> > > >
> > > > Sorry for saying this in v7. I've been meaning to say it, but somehow ended up
> > > > procrastinating too much.
> > >
> > > No bother. Since Frank has given a lot of feedback on this series, I'd also like
> > > to hear his view.
> > >
> > > I'm fine with either direction, and can revisit the earlier vNTB-embedded
> > > approach:
> > > https://lore.kernel.org/r/sn67hi7kljh7cgmgodatb3naz2astlaklqfobdbxyyzgoohxqb@4nnetbhqwba4/
> >
> > One of the important value is test dw-edma-pcie.c, which generally depend
> > on some fpga hardware. If there are epf driver work as fpga hardware, it
> > will help cover edma remote user case. So more user can test it.
> >
> > Of course, this implement are over complex. I suggest update dma-engine
> > chan_id to support static allocate, which also need be fixed because
> > some drivers have such dependence, anyway need be fixed. After this fix,
> > this patches will become simpler.
> >
> > VNTB case, it'd better put such informaiton into one BARs and work on
> > single-function.
> >
> > I suggest split two things
> >
> > 1 - create simple epf driver to test dw-edma-pcie.c.
> > 2 - vntb support DMA.
> >
> > of course, if shared efforts, it will be great.
>
> Thanks for the feedback. v7 patch 1 adds that exact static chan_id support, so
> you mean this v7 patch 2 ~ 9 can still be simplified enough to make the
> maintenance burden acceptable, right? Please correct me if I misunderstood.
Sorry, I have not realized this new posted patches. let me check.
Frank
>
> I think a large part of the remaining complexity comes from discovery. Ideally,
> the EPF would create a VSEC like existing supported hardware, and put the DMA
> description there. But I could not find a generic way for an EPF to provide
> writable configuration space backing for such a VSEC on a DWC PCIe controller in
> EP mode. That is why this series (sadly) ended up with a BAR protocol. BAR
> subrange mapping use might look complicated, but is still needed on systems
> without a fixed DMA BAR, but the discovery part would otherwise be much simpler.
> I may be missing something I could simplify further.
>
> Best regards,
> Koichiro
>
> >
> > Frank
> >
> > >
> > > Best regards,
> > > Koichiro
> > >
> > > >
> > > > - Mani
> > > >
> > > > --
> > > > மணிவண்ணன் சதாசிவம்
© 2016 - 2026 Red Hat, Inc.