[PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver

Ekansh Gupta posted 15 patches 1 month, 1 week ago
Only 14 patches received!
Documentation/accel/index.rst          |    1 +
Documentation/accel/qda/index.rst      |   13 +
Documentation/accel/qda/qda.rst        |  191 ++++++
MAINTAINERS                            |   11 +
drivers/accel/Kconfig                  |    1 +
drivers/accel/Makefile                 |    2 +
drivers/accel/qda/Kconfig              |   34 ++
drivers/accel/qda/Makefile             |   19 +
drivers/accel/qda/qda_cb.c             |  125 ++++
drivers/accel/qda/qda_cb.h             |   32 +
drivers/accel/qda/qda_compute_bus.c    |   80 +++
drivers/accel/qda/qda_drv.c            |  144 +++++
drivers/accel/qda/qda_drv.h            |   90 +++
drivers/accel/qda/qda_fastrpc.c        | 1009 ++++++++++++++++++++++++++++++++
drivers/accel/qda/qda_fastrpc.h        |  367 ++++++++++++
drivers/accel/qda/qda_gem.c            |  155 +++++
drivers/accel/qda/qda_gem.h            |   60 ++
drivers/accel/qda/qda_ioctl.c          |  290 +++++++++
drivers/accel/qda/qda_ioctl.h          |   19 +
drivers/accel/qda/qda_memory_dma.c     |   82 +++
drivers/accel/qda/qda_memory_dma.h     |   17 +
drivers/accel/qda/qda_memory_manager.c |  369 ++++++++++++
drivers/accel/qda/qda_memory_manager.h |   85 +++
drivers/accel/qda/qda_prime.c          |  167 ++++++
drivers/accel/qda/qda_prime.h          |   18 +
drivers/accel/qda/qda_rpmsg.c          |  201 +++++++
drivers/accel/qda/qda_rpmsg.h          |   26 +
drivers/iommu/iommu.c                  |    4 +
include/linux/qda_compute_bus.h        |   33 ++
include/uapi/drm/qda_accel.h           |  242 ++++++++
30 files changed, 3887 insertions(+)
[PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver
Posted by Ekansh Gupta 1 month, 1 week ago
This patch series introduces the Qualcomm DSP Accelerator (QDA) driver,
a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a
standardized interface for offloading computational tasks to DSPs found
on Qualcomm SoCs, supporting all DSP domains.

The QDA driver implements the FastRPC protocol over the DRM accel
subsystem. It uses the same device-tree node structure as the existing
fastrpc driver in drivers/misc/. The approach for binding the QDA driver
to device-tree nodes while coexisting with the fastrpc driver is an open
item described below.

v1: https://lore.kernel.org/all/20260519-qda-series-v1-0-b2d984c297f8@oss.qualcomm.com/
RFC: https://lore.kernel.org/dri-devel/20260224-qda-firstpost-v1-0-fe46a9c1a046@oss.qualcomm.com/T/

Changes since v1
================

The v1 review raised two architectural objections and one correctness
issue; all three are resolved in v2:

* Christian König (dma-buf maintainer) pointed out that the imported-
  buffer path silently assumed the IOMMU maps every buffer as a single
  contiguous range, which is not guaranteed. v2 walks the scatterlist
  and cleanly rejects non-contiguous imports; contiguous imports (e.g.
  CMA DMA-buf heap) are accepted. (patch 11)

* Dmitry Baryshkov objected to three different buffer-passing formats
  in the invoke IOCTL (DMA-BUF fd, direct/inline, DMA handle). v2
  passes only GEM handles; userspace imports any fd to a GEM handle
  with DRM_IOCTL_PRIME_FD_TO_HANDLE before invoking. Packing and
  overlap handling are left to userspace. (patch 12)

* The memory manager (patch 07) used a fixed 16-entry array without
  justification and leaked the device descriptor on teardown. v2
  allocates the array from the DT context-bank count (as Dmitry
  suggested) and frees it correctly.

User-space staging branch
=========================
https://github.com/qualcomm/fastrpc/tree/accel/staging

Key Features
============

* Standard DRM accelerator interface via /dev/accel/accelN
* GEM-based buffer management with DMA-BUF import (PRIME)
* IOMMU-based memory isolation using per-process context banks
* FastRPC protocol implementation for DSP communication
* RPMsg transport layer for reliable message passing
* Support for all DSP domains (ADSP, CDSP, SDSP, GDSP)
* DRM IOCTL interface for DSP session management, buffer allocation,
  and remote procedure invocation

Architecture
============

1. DRM Accelerator Framework Integration
   The driver registers as a DRM accel device, exposing a standard
   /dev/accel/accelN character device node. This provides established
   DRM infrastructure for device management, file operations, and
   IOCTL dispatch.

2. Memory Management
   Buffers are managed as GEM objects with PRIME support for DMA-BUF
   import. This enables buffer sharing with other DRM drivers (GPU,
   camera, video) using standard kernel mechanisms. Only contiguous
   imports are accepted; the driver verifies contiguity at import time
   rather than assuming it.

3. IOMMU Context Bank Management
   IOMMU context banks (CBs) are represented as proper struct device
   instances on a custom virtual bus (qda-compute-cb). Each CB device
   is registered with the IOMMU subsystem and receives its own IOMMU
   domain, enabling per-session address space isolation. The custom
   bus was introduced because IOMMU context banks are synthetic
   constructs — not real platform devices — and to ensure CB device
   lifetime is strictly subordinate to the parent QDA device.
   See also: https://lore.kernel.org/all/245d602f-3037-4ae3-9af9-d98f37258aae@oss.qualcomm.com/

4. Memory Manager Architecture
   The memory manager maintains a registry of IOMMU devices in an
   array sized to the number of context banks described in the device
   tree, and coordinates per-process device assignment with reference-
   counted lifetime management. The DMA-coherent backend allocates
   buffers with SID-prefixed DMA addresses for DSP firmware
   compatibility.

5. Transport Layer
   RPMsg communication is handled in a dedicated transport layer
   (qda_rpmsg.c), separate from the core DRM driver logic.

6. Code Organization
   The driver is organized across multiple files (~4800 lines total):
   * qda_drv.c:            Core driver and DRM integration
   * qda_rpmsg.c:          RPMsg transport layer
   * qda_cb.c:             Context bank device management
   * qda_compute_bus.c:    Custom virtual bus for CB devices
   * qda_gem.c:            GEM object management
   * qda_prime.c:          DMA-BUF import (PRIME)
   * qda_memory_manager.c: IOMMU device registry and allocation
   * qda_memory_dma.c:     DMA-coherent allocation backend
   * qda_fastrpc.c:        FastRPC protocol implementation
   * qda_ioctl.c:          IOCTL dispatch

7. UAPI Design
   The driver exposes DRM-style IOCTLs defined in
   include/uapi/drm/qda_accel.h, following DRM UAPI conventions
   (__u32/__u64 types, C++ guard, GPL-2.0-only WITH Linux-syscall-note).
   Buffer arguments are identified by GEM handles; the driver never
   accepts DMA-BUF fds directly in any IOCTL.

Patch Series Organization
==========================

Patch 01:      MAINTAINERS entry
Patch 02:      Driver documentation (Documentation/accel/qda/)
Patches 03-04: Core driver skeleton and compute bus
Patch 05:      iommu: Register qda-compute-cb bus with IOMMU subsystem
Patches 06-07: CB device enumeration and memory manager
Patch 08:      QUERY IOCTL and UAPI header
Patches 09-11: GEM buffer management and PRIME import
Patches 12-15: FastRPC protocol (invoke, session create/release,
               map/unmap)

Open Items
===========

1. Device-Tree Compatible String
   The QDA driver uses the same device-tree node structure and
   properties as the existing fastrpc driver in drivers/misc/. A
   mechanism is needed to allow the QDA driver to bind to its device
   node independently of the fastrpc driver.

   The intended coexistence model is: platforms that require the
   complete fastrpc feature set continue to use "qcom,fastrpc"; new
   platforms where QDA's feature set is sufficient use a QDA-specific
   compatible string. New feature development is directed toward QDA.

   The options under consideration are:

   a) Add a new "qcom,qda" compatible string to the existing
      qcom,fastrpc.yaml binding, since the DT node structure and
      properties are identical.

   b) Introduce a separate qcom,qda.yaml binding that references or
      inherits the fastrpc binding properties.

   Seeking guidance from DT binding maintainers on the preferred
   approach.

2. Privilege Level Management
   Currently, daemon processes and user processes have the same access
   level as both use the same accel device node. Daemons attach to
   privileged DSP protection domains and require higher privilege
   levels for system-level operations. Seeking guidance on the best
   approach: separate device nodes, capability-based checks, or DRM
   master/authentication mechanisms.

3. Audio and Sensors PD Support
   The current series does not handle Audio PD and Sensors PD
   functionalities. These specialized protection domains require
   additional support for real-time constraints and power management.

Interface Compatibility
========================

The QDA driver uses the same device-tree node structure and child node
layout (including "qcom,fastrpc-compute-cb" child nodes) as the
existing fastrpc driver. The underlying FastRPC protocol and DSP
firmware interface are compatible with the existing fastrpc driver,
ensuring that DSP firmware and libraries continue to work without
modification.

References
==========

Previous discussions on this migration:
- https://lkml.org/lkml/2024/6/24/479
- https://lkml.org/lkml/2024/6/21/1252

Testing
=======

The driver has been tested on Qualcomm platforms with:
- Basic FastRPC attach/release operations
- DSP process creation and initialization
- Memory mapping/unmapping operations
- Dynamic invocation with various buffer types
- GEM buffer allocation and mmap
- PRIME buffer import from other subsystems (contiguous buffers)

Signed-off-by: Ekansh Gupta <ekansh.gupta@oss.qualcomm.com>
---
Ekansh Gupta (15):
      MAINTAINERS: Add entry for Qualcomm DSP Accelerator (QDA) driver
      accel/qda: Add QDA driver documentation
      accel/qda: Add initial QDA DRM accelerator driver
      accel/qda: Add compute bus for QDA context banks
      iommu: Add QDA compute context bank bus to iommu_buses
      accel/qda: Create compute context bank devices on QDA compute bus
      accel/qda: Add memory manager for CB devices
      accel/qda: Add QUERY IOCTL and QDA UAPI header
      accel/qda: Add DMA-backed GEM objects and memory manager integration
      accel/qda: Add GEM_CREATE and GEM_MMAP_OFFSET IOCTLs
      accel/qda: Add PRIME DMA-BUF import support
      accel/qda: Add FastRPC invocation support
      accel/qda: Add DSP process creation and release
      accel/qda: Add remote memory mapping to DSP address space
      accel/qda: Add remote memory unmap from DSP address space

 Documentation/accel/index.rst          |    1 +
 Documentation/accel/qda/index.rst      |   13 +
 Documentation/accel/qda/qda.rst        |  191 ++++++
 MAINTAINERS                            |   11 +
 drivers/accel/Kconfig                  |    1 +
 drivers/accel/Makefile                 |    2 +
 drivers/accel/qda/Kconfig              |   34 ++
 drivers/accel/qda/Makefile             |   19 +
 drivers/accel/qda/qda_cb.c             |  125 ++++
 drivers/accel/qda/qda_cb.h             |   32 +
 drivers/accel/qda/qda_compute_bus.c    |   80 +++
 drivers/accel/qda/qda_drv.c            |  144 +++++
 drivers/accel/qda/qda_drv.h            |   90 +++
 drivers/accel/qda/qda_fastrpc.c        | 1009 ++++++++++++++++++++++++++++++++
 drivers/accel/qda/qda_fastrpc.h        |  367 ++++++++++++
 drivers/accel/qda/qda_gem.c            |  155 +++++
 drivers/accel/qda/qda_gem.h            |   60 ++
 drivers/accel/qda/qda_ioctl.c          |  290 +++++++++
 drivers/accel/qda/qda_ioctl.h          |   19 +
 drivers/accel/qda/qda_memory_dma.c     |   82 +++
 drivers/accel/qda/qda_memory_dma.h     |   17 +
 drivers/accel/qda/qda_memory_manager.c |  369 ++++++++++++
 drivers/accel/qda/qda_memory_manager.h |   85 +++
 drivers/accel/qda/qda_prime.c          |  167 ++++++
 drivers/accel/qda/qda_prime.h          |   18 +
 drivers/accel/qda/qda_rpmsg.c          |  201 +++++++
 drivers/accel/qda/qda_rpmsg.h          |   26 +
 drivers/iommu/iommu.c                  |    4 +
 include/linux/qda_compute_bus.h        |   33 ++
 include/uapi/drm/qda_accel.h           |  242 ++++++++
 30 files changed, 3887 insertions(+)
---
base-commit: 5f07a0db7088b4ef4b9a48069a93b9f3e1a33379
change-id: 20260817-qda-v2-78e2d1f10529

Best regards,
-- 
Ekansh Gupta <ekansh.gupta@oss.qualcomm.com>

Re: [PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver
Posted by Krzysztof Kozlowski 1 month, 1 week ago
On 17/08/2026 06:47, Ekansh Gupta wrote:
> This patch series introduces the Qualcomm DSP Accelerator (QDA) driver,
> a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a
> standardized interface for offloading computational tasks to DSPs found
> on Qualcomm SoCs, supporting all DSP domains.
> 
> The QDA driver implements the FastRPC protocol over the DRM accel
> subsystem. It uses the same device-tree node structure as the existing
> fastrpc driver in drivers/misc/. The approach for binding the QDA driver
> to device-tree nodes while coexisting with the fastrpc driver is an open
> item described below.

No. Grow/replace/improve existing driver instead of coming with a duplicate.

That's a standard upstream requirement, basically given on every
upstreaming guide.

Please watch old talk from Greg - "I Don’t Want Your Code!".

> 
> v1: https://lore.kernel.org/all/20260519-qda-series-v1-0-b2d984c297f8@oss.qualcomm.com/
> RFC: https://lore.kernel.org/dri-devel/20260224-qda-firstpost-v1-0-fe46a9c1a046@oss.qualcomm.com/T/
> 
> Changes since v1
> ================
> 
> The v1 review raised two architectural objections and one correctness
> issue; all three are resolved in v2:
> 
> * Christian König (dma-buf maintainer) pointed out that the imported-
>   buffer path silently assumed the IOMMU maps every buffer as a single
>   contiguous range, which is not guaranteed. v2 walks the scatterlist
>   and cleanly rejects non-contiguous imports; contiguous imports (e.g.
>   CMA DMA-buf heap) are accepted. (patch 11)
> 
> * Dmitry Baryshkov objected to three different buffer-passing formats
>   in the invoke IOCTL (DMA-BUF fd, direct/inline, DMA handle). v2
>   passes only GEM handles; userspace imports any fd to a GEM handle
>   with DRM_IOCTL_PRIME_FD_TO_HANDLE before invoking. Packing and
>   overlap handling are left to userspace. (patch 12)
> 
> * The memory manager (patch 07) used a fixed 16-entry array without
>   justification and leaked the device descriptor on teardown. v2
>   allocates the array from the DT context-bank count (as Dmitry
>   suggested) and frees it correctly.
> 
> User-space staging branch
> =========================
> https://github.com/qualcomm/fastrpc/tree/accel/staging
> 
> Key Features
> ============
> 
> * Standard DRM accelerator interface via /dev/accel/accelN
> * GEM-based buffer management with DMA-BUF import (PRIME)
> * IOMMU-based memory isolation using per-process context banks
> * FastRPC protocol implementation for DSP communication
> * RPMsg transport layer for reliable message passing
> * Support for all DSP domains (ADSP, CDSP, SDSP, GDSP)
> * DRM IOCTL interface for DSP session management, buffer allocation,
>   and remote procedure invocation
> 
> Architecture
> ============
> 
> 1. DRM Accelerator Framework Integration
>    The driver registers as a DRM accel device, exposing a standard
>    /dev/accel/accelN character device node. This provides established
>    DRM infrastructure for device management, file operations, and
>    IOCTL dispatch.
> 
> 2. Memory Management
>    Buffers are managed as GEM objects with PRIME support for DMA-BUF
>    import. This enables buffer sharing with other DRM drivers (GPU,
>    camera, video) using standard kernel mechanisms. Only contiguous
>    imports are accepted; the driver verifies contiguity at import time
>    rather than assuming it.
> 
> 3. IOMMU Context Bank Management
>    IOMMU context banks (CBs) are represented as proper struct device
>    instances on a custom virtual bus (qda-compute-cb). Each CB device
>    is registered with the IOMMU subsystem and receives its own IOMMU
>    domain, enabling per-session address space isolation. The custom
>    bus was introduced because IOMMU context banks are synthetic
>    constructs — not real platform devices — and to ensure CB device
>    lifetime is strictly subordinate to the parent QDA device.
>    See also: https://lore.kernel.org/all/245d602f-3037-4ae3-9af9-d98f37258aae@oss.qualcomm.com/
> 
> 4. Memory Manager Architecture
>    The memory manager maintains a registry of IOMMU devices in an
>    array sized to the number of context banks described in the device
>    tree, and coordinates per-process device assignment with reference-
>    counted lifetime management. The DMA-coherent backend allocates
>    buffers with SID-prefixed DMA addresses for DSP firmware
>    compatibility.
> 
> 5. Transport Layer
>    RPMsg communication is handled in a dedicated transport layer
>    (qda_rpmsg.c), separate from the core DRM driver logic.
> 
> 6. Code Organization
>    The driver is organized across multiple files (~4800 lines total):
>    * qda_drv.c:            Core driver and DRM integration
>    * qda_rpmsg.c:          RPMsg transport layer
>    * qda_cb.c:             Context bank device management
>    * qda_compute_bus.c:    Custom virtual bus for CB devices
>    * qda_gem.c:            GEM object management
>    * qda_prime.c:          DMA-BUF import (PRIME)
>    * qda_memory_manager.c: IOMMU device registry and allocation
>    * qda_memory_dma.c:     DMA-coherent allocation backend
>    * qda_fastrpc.c:        FastRPC protocol implementation
>    * qda_ioctl.c:          IOCTL dispatch
> 
> 7. UAPI Design
>    The driver exposes DRM-style IOCTLs defined in
>    include/uapi/drm/qda_accel.h, following DRM UAPI conventions
>    (__u32/__u64 types, C++ guard, GPL-2.0-only WITH Linux-syscall-note).
>    Buffer arguments are identified by GEM handles; the driver never
>    accepts DMA-BUF fds directly in any IOCTL.
> 
> Patch Series Organization
> ==========================
> 
> Patch 01:      MAINTAINERS entry
> Patch 02:      Driver documentation (Documentation/accel/qda/)
> Patches 03-04: Core driver skeleton and compute bus
> Patch 05:      iommu: Register qda-compute-cb bus with IOMMU subsystem
> Patches 06-07: CB device enumeration and memory manager
> Patch 08:      QUERY IOCTL and UAPI header
> Patches 09-11: GEM buffer management and PRIME import
> Patches 12-15: FastRPC protocol (invoke, session create/release,
>                map/unmap)
> 
> Open Items
> ===========
> 
> 1. Device-Tree Compatible String
>    The QDA driver uses the same device-tree node structure and
>    properties as the existing fastrpc driver in drivers/misc/. A
>    mechanism is needed to allow the QDA driver to bind to its device
>    node independently of the fastrpc driver.
> 
>    The intended coexistence model is: platforms that require the
>    complete fastrpc feature set continue to use "qcom,fastrpc"; new
>    platforms where QDA's feature set is sufficient use a QDA-specific
>    compatible string. New feature development is directed toward QDA.
> 
>    The options under consideration are:
> 
>    a) Add a new "qcom,qda" compatible string to the existing
>       qcom,fastrpc.yaml binding, since the DT node structure and
>       properties are identical.
No

> 
>    b) Introduce a separate qcom,qda.yaml binding that references or
>       inherits the fastrpc binding properties.

No

> 
>    Seeking guidance from DT binding maintainers on the preferred
>    approach.

Grow existing driver. You do not get new driver, you do not get new
bindings.


Best regards,
Krzysztof
Re: [PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver
Posted by Ekansh Gupta 1 month, 1 week ago
On 19-08-2026 00:43, Krzysztof Kozlowski wrote:
> On 17/08/2026 06:47, Ekansh Gupta wrote:
>> This patch series introduces the Qualcomm DSP Accelerator (QDA) driver,
>> a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a
>> standardized interface for offloading computational tasks to DSPs found
>> on Qualcomm SoCs, supporting all DSP domains.
>>
>> The QDA driver implements the FastRPC protocol over the DRM accel
>> subsystem. It uses the same device-tree node structure as the existing
>> fastrpc driver in drivers/misc/. The approach for binding the QDA driver
>> to device-tree nodes while coexisting with the fastrpc driver is an open
>> item described below.
> 
> No. Grow/replace/improve existing driver instead of coming with a duplicate.
> 
> That's a standard upstream requirement, basically given on every
> upstreaming guide.
> 
> Please watch old talk from Greg - "I Don’t Want Your Code!".
Posted discussion threads here[1]. Would seek comments from Dmitry,
Srini as well.

[1]
https://lore.kernel.org/all/3476b5c3-7983-4994-a901-3d7d8bd75255@oss.qualcomm.com/

> 
>>
>> v1: https://lore.kernel.org/all/20260519-qda-series-v1-0-b2d984c297f8@oss.qualcomm.com/
>> RFC: https://lore.kernel.org/dri-devel/20260224-qda-firstpost-v1-0-fe46a9c1a046@oss.qualcomm.com/T/
>>
>> Changes since v1
>> ================
>>
>> The v1 review raised two architectural objections and one correctness
>> issue; all three are resolved in v2:
>>
>> * Christian König (dma-buf maintainer) pointed out that the imported-
>>   buffer path silently assumed the IOMMU maps every buffer as a single
>>   contiguous range, which is not guaranteed. v2 walks the scatterlist
>>   and cleanly rejects non-contiguous imports; contiguous imports (e.g.
>>   CMA DMA-buf heap) are accepted. (patch 11)
>>
>> * Dmitry Baryshkov objected to three different buffer-passing formats
>>   in the invoke IOCTL (DMA-BUF fd, direct/inline, DMA handle). v2
>>   passes only GEM handles; userspace imports any fd to a GEM handle
>>   with DRM_IOCTL_PRIME_FD_TO_HANDLE before invoking. Packing and
>>   overlap handling are left to userspace. (patch 12)
>>
>> * The memory manager (patch 07) used a fixed 16-entry array without
>>   justification and leaked the device descriptor on teardown. v2
>>   allocates the array from the DT context-bank count (as Dmitry
>>   suggested) and frees it correctly.
>>
>> User-space staging branch
>> =========================
>> https://github.com/qualcomm/fastrpc/tree/accel/staging
>>
>> Key Features
>> ============
>>
>> * Standard DRM accelerator interface via /dev/accel/accelN
>> * GEM-based buffer management with DMA-BUF import (PRIME)
>> * IOMMU-based memory isolation using per-process context banks
>> * FastRPC protocol implementation for DSP communication
>> * RPMsg transport layer for reliable message passing
>> * Support for all DSP domains (ADSP, CDSP, SDSP, GDSP)
>> * DRM IOCTL interface for DSP session management, buffer allocation,
>>   and remote procedure invocation
>>
>> Architecture
>> ============
>>
>> 1. DRM Accelerator Framework Integration
>>    The driver registers as a DRM accel device, exposing a standard
>>    /dev/accel/accelN character device node. This provides established
>>    DRM infrastructure for device management, file operations, and
>>    IOCTL dispatch.
>>
>> 2. Memory Management
>>    Buffers are managed as GEM objects with PRIME support for DMA-BUF
>>    import. This enables buffer sharing with other DRM drivers (GPU,
>>    camera, video) using standard kernel mechanisms. Only contiguous
>>    imports are accepted; the driver verifies contiguity at import time
>>    rather than assuming it.
>>
>> 3. IOMMU Context Bank Management
>>    IOMMU context banks (CBs) are represented as proper struct device
>>    instances on a custom virtual bus (qda-compute-cb). Each CB device
>>    is registered with the IOMMU subsystem and receives its own IOMMU
>>    domain, enabling per-session address space isolation. The custom
>>    bus was introduced because IOMMU context banks are synthetic
>>    constructs — not real platform devices — and to ensure CB device
>>    lifetime is strictly subordinate to the parent QDA device.
>>    See also: https://lore.kernel.org/all/245d602f-3037-4ae3-9af9-d98f37258aae@oss.qualcomm.com/
>>
>> 4. Memory Manager Architecture
>>    The memory manager maintains a registry of IOMMU devices in an
>>    array sized to the number of context banks described in the device
>>    tree, and coordinates per-process device assignment with reference-
>>    counted lifetime management. The DMA-coherent backend allocates
>>    buffers with SID-prefixed DMA addresses for DSP firmware
>>    compatibility.
>>
>> 5. Transport Layer
>>    RPMsg communication is handled in a dedicated transport layer
>>    (qda_rpmsg.c), separate from the core DRM driver logic.
>>
>> 6. Code Organization
>>    The driver is organized across multiple files (~4800 lines total):
>>    * qda_drv.c:            Core driver and DRM integration
>>    * qda_rpmsg.c:          RPMsg transport layer
>>    * qda_cb.c:             Context bank device management
>>    * qda_compute_bus.c:    Custom virtual bus for CB devices
>>    * qda_gem.c:            GEM object management
>>    * qda_prime.c:          DMA-BUF import (PRIME)
>>    * qda_memory_manager.c: IOMMU device registry and allocation
>>    * qda_memory_dma.c:     DMA-coherent allocation backend
>>    * qda_fastrpc.c:        FastRPC protocol implementation
>>    * qda_ioctl.c:          IOCTL dispatch
>>
>> 7. UAPI Design
>>    The driver exposes DRM-style IOCTLs defined in
>>    include/uapi/drm/qda_accel.h, following DRM UAPI conventions
>>    (__u32/__u64 types, C++ guard, GPL-2.0-only WITH Linux-syscall-note).
>>    Buffer arguments are identified by GEM handles; the driver never
>>    accepts DMA-BUF fds directly in any IOCTL.
>>
>> Patch Series Organization
>> ==========================
>>
>> Patch 01:      MAINTAINERS entry
>> Patch 02:      Driver documentation (Documentation/accel/qda/)
>> Patches 03-04: Core driver skeleton and compute bus
>> Patch 05:      iommu: Register qda-compute-cb bus with IOMMU subsystem
>> Patches 06-07: CB device enumeration and memory manager
>> Patch 08:      QUERY IOCTL and UAPI header
>> Patches 09-11: GEM buffer management and PRIME import
>> Patches 12-15: FastRPC protocol (invoke, session create/release,
>>                map/unmap)
>>
>> Open Items
>> ===========
>>
>> 1. Device-Tree Compatible String
>>    The QDA driver uses the same device-tree node structure and
>>    properties as the existing fastrpc driver in drivers/misc/. A
>>    mechanism is needed to allow the QDA driver to bind to its device
>>    node independently of the fastrpc driver.
>>
>>    The intended coexistence model is: platforms that require the
>>    complete fastrpc feature set continue to use "qcom,fastrpc"; new
>>    platforms where QDA's feature set is sufficient use a QDA-specific
>>    compatible string. New feature development is directed toward QDA.
>>
>>    The options under consideration are:
>>
>>    a) Add a new "qcom,qda" compatible string to the existing
>>       qcom,fastrpc.yaml binding, since the DT node structure and
>>       properties are identical.
> No
> 
>>
>>    b) Introduce a separate qcom,qda.yaml binding that references or
>>       inherits the fastrpc binding properties.
> 
> No
> 
>>
>>    Seeking guidance from DT binding maintainers on the preferred
>>    approach.
> 
> Grow existing driver. You do not get new driver, you do not get new
> bindings.
> 
> 
> Best regards,
> Krzysztof

Re: [PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver
Posted by Krzysztof Kozlowski 1 month, 1 week ago
On 19/08/2026 15:26, Ekansh Gupta wrote:
> On 19-08-2026 00:43, Krzysztof Kozlowski wrote:
>> On 17/08/2026 06:47, Ekansh Gupta wrote:
>>> This patch series introduces the Qualcomm DSP Accelerator (QDA) driver,
>>> a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a
>>> standardized interface for offloading computational tasks to DSPs found
>>> on Qualcomm SoCs, supporting all DSP domains.
>>>
>>> The QDA driver implements the FastRPC protocol over the DRM accel
>>> subsystem. It uses the same device-tree node structure as the existing
>>> fastrpc driver in drivers/misc/. The approach for binding the QDA driver
>>> to device-tree nodes while coexisting with the fastrpc driver is an open
>>> item described below.
>>
>> No. Grow/replace/improve existing driver instead of coming with a duplicate.
>>
>> That's a standard upstream requirement, basically given on every
>> upstreaming guide.
>>
>> Please watch old talk from Greg - "I Don’t Want Your Code!".
> Posted discussion threads here[1]. Would seek comments from Dmitry,
> Srini as well.
> 
> [1]
> https://lore.kernel.org/all/3476b5c3-7983-4994-a901-3d7d8bd75255@oss.qualcomm.com/

The rest of the comments is still valid even if you did not acknowledge
them.

Anyway, regarding above - again, watch the talk from Greg.

You have ONE driver. Not two.

Best regards,
Krzysztof
Re: [PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver
Posted by Rob Clark 1 month, 1 week ago
On Wed, Aug 19, 2026 at 7:21 AM Krzysztof Kozlowski <krzk@kernel.org> wrote:
>
> On 19/08/2026 15:26, Ekansh Gupta wrote:
> > On 19-08-2026 00:43, Krzysztof Kozlowski wrote:
> >> On 17/08/2026 06:47, Ekansh Gupta wrote:
> >>> This patch series introduces the Qualcomm DSP Accelerator (QDA) driver,
> >>> a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a
> >>> standardized interface for offloading computational tasks to DSPs found
> >>> on Qualcomm SoCs, supporting all DSP domains.
> >>>
> >>> The QDA driver implements the FastRPC protocol over the DRM accel
> >>> subsystem. It uses the same device-tree node structure as the existing
> >>> fastrpc driver in drivers/misc/. The approach for binding the QDA driver
> >>> to device-tree nodes while coexisting with the fastrpc driver is an open
> >>> item described below.
> >>
> >> No. Grow/replace/improve existing driver instead of coming with a duplicate.
> >>
> >> That's a standard upstream requirement, basically given on every
> >> upstreaming guide.
> >>
> >> Please watch old talk from Greg - "I Don’t Want Your Code!".
> > Posted discussion threads here[1]. Would seek comments from Dmitry,
> > Srini as well.
> >
> > [1]
> > https://lore.kernel.org/all/3476b5c3-7983-4994-a901-3d7d8bd75255@oss.qualcomm.com/
>
> The rest of the comments is still valid even if you did not acknowledge
> them.
>
> Anyway, regarding above - again, watch the talk from Greg.
>
> You have ONE driver. Not two.

Long term, moving to the common driver framework (which did not exist
when fastrpc was first created) seems like a good thing.  But does
that not allow for some transition period?  How can we get from here
to there without otherwise breaking userspace?  Is there some other
precedent elsewhere in other driver subsystems?

I suppose drm exposing legacy fbdev on top of drm drivers is _sort of_
a precedent if you squint a bit?  I'm not really familiar enough to
say if that would be reasonable/possible in this case.

BR,
-R

> Best regards,
> Krzysztof
Re: [PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver
Posted by Krzysztof Kozlowski 1 month, 1 week ago
On 19/08/2026 16:38, Rob Clark wrote:
> On Wed, Aug 19, 2026 at 7:21 AM Krzysztof Kozlowski <krzk@kernel.org> wrote:
>>
>> On 19/08/2026 15:26, Ekansh Gupta wrote:
>>> On 19-08-2026 00:43, Krzysztof Kozlowski wrote:
>>>> On 17/08/2026 06:47, Ekansh Gupta wrote:
>>>>> This patch series introduces the Qualcomm DSP Accelerator (QDA) driver,
>>>>> a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a
>>>>> standardized interface for offloading computational tasks to DSPs found
>>>>> on Qualcomm SoCs, supporting all DSP domains.
>>>>>
>>>>> The QDA driver implements the FastRPC protocol over the DRM accel
>>>>> subsystem. It uses the same device-tree node structure as the existing
>>>>> fastrpc driver in drivers/misc/. The approach for binding the QDA driver
>>>>> to device-tree nodes while coexisting with the fastrpc driver is an open
>>>>> item described below.
>>>>
>>>> No. Grow/replace/improve existing driver instead of coming with a duplicate.
>>>>
>>>> That's a standard upstream requirement, basically given on every
>>>> upstreaming guide.
>>>>
>>>> Please watch old talk from Greg - "I Don’t Want Your Code!".
>>> Posted discussion threads here[1]. Would seek comments from Dmitry,
>>> Srini as well.
>>>
>>> [1]
>>> https://lore.kernel.org/all/3476b5c3-7983-4994-a901-3d7d8bd75255@oss.qualcomm.com/
>>
>> The rest of the comments is still valid even if you did not acknowledge
>> them.
>>
>> Anyway, regarding above - again, watch the talk from Greg.
>>
>> You have ONE driver. Not two.
> 
> Long term, moving to the common driver framework (which did not exist
> when fastrpc was first created) seems like a good thing.  But does
> that not allow for some transition period?  How can we get from here
> to there without otherwise breaking userspace?  Is there some other
> precedent elsewhere in other driver subsystems?

Yes, Iris and Venus where we agreed for an exception (two drivers) as
long as new driver supports old hardware / features.

This is not the case here, right?

> 
> I suppose drm exposing legacy fbdev on top of drm drivers is _sort of_
> a precedent if you squint a bit?  I'm not really familiar enough to
> say if that would be reasonable/possible in this case.

Growing old driver does not look complicated itself. The only a bit
tricky thing is to present somehow exclusive interface to user-space,
like usage of one disables the second etc. Depending on actual
differences in that interface.

But having a duplicated driver is a clear no go and it is well known
upstream requirement. Nothing new here.

Best regards,
Krzysztof
Re: [PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver
Posted by Rob Clark 1 month, 1 week ago
On Wed, Aug 19, 2026 at 7:43 AM Krzysztof Kozlowski <krzk@kernel.org> wrote:
>
> On 19/08/2026 16:38, Rob Clark wrote:
> > On Wed, Aug 19, 2026 at 7:21 AM Krzysztof Kozlowski <krzk@kernel.org> wrote:
> >>
> >> On 19/08/2026 15:26, Ekansh Gupta wrote:
> >>> On 19-08-2026 00:43, Krzysztof Kozlowski wrote:
> >>>> On 17/08/2026 06:47, Ekansh Gupta wrote:
> >>>>> This patch series introduces the Qualcomm DSP Accelerator (QDA) driver,
> >>>>> a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a
> >>>>> standardized interface for offloading computational tasks to DSPs found
> >>>>> on Qualcomm SoCs, supporting all DSP domains.
> >>>>>
> >>>>> The QDA driver implements the FastRPC protocol over the DRM accel
> >>>>> subsystem. It uses the same device-tree node structure as the existing
> >>>>> fastrpc driver in drivers/misc/. The approach for binding the QDA driver
> >>>>> to device-tree nodes while coexisting with the fastrpc driver is an open
> >>>>> item described below.
> >>>>
> >>>> No. Grow/replace/improve existing driver instead of coming with a duplicate.
> >>>>
> >>>> That's a standard upstream requirement, basically given on every
> >>>> upstreaming guide.
> >>>>
> >>>> Please watch old talk from Greg - "I Don’t Want Your Code!".
> >>> Posted discussion threads here[1]. Would seek comments from Dmitry,
> >>> Srini as well.
> >>>
> >>> [1]
> >>> https://lore.kernel.org/all/3476b5c3-7983-4994-a901-3d7d8bd75255@oss.qualcomm.com/
> >>
> >> The rest of the comments is still valid even if you did not acknowledge
> >> them.
> >>
> >> Anyway, regarding above - again, watch the talk from Greg.
> >>
> >> You have ONE driver. Not two.
> >
> > Long term, moving to the common driver framework (which did not exist
> > when fastrpc was first created) seems like a good thing.  But does
> > that not allow for some transition period?  How can we get from here
> > to there without otherwise breaking userspace?  Is there some other
> > precedent elsewhere in other driver subsystems?
>
> Yes, Iris and Venus where we agreed for an exception (two drivers) as
> long as new driver supports old hardware / features.
>
> This is not the case here, right?
>
> >
> > I suppose drm exposing legacy fbdev on top of drm drivers is _sort of_
> > a precedent if you squint a bit?  I'm not really familiar enough to
> > say if that would be reasonable/possible in this case.
>
> Growing old driver does not look complicated itself. The only a bit
> tricky thing is to present somehow exclusive interface to user-space,
> like usage of one disables the second etc. Depending on actual
> differences in that interface.
>
> But having a duplicated driver is a clear no go and it is well known
> upstream requirement. Nothing new here.

Would it be acceptable as a first step to, by default (ie. when no
cmdline override/etc) for the new driver to bind to new compatibles
and the existing driver to old?  Ie. have both drivers but only one
binds on a given platform?

BR,
-R
Re: [PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver
Posted by Krzysztof Kozlowski 1 month, 1 week ago
On 19/08/2026 16:49, Rob Clark wrote:
> On Wed, Aug 19, 2026 at 7:43 AM Krzysztof Kozlowski <krzk@kernel.org> wrote:
>>
>> On 19/08/2026 16:38, Rob Clark wrote:
>>> On Wed, Aug 19, 2026 at 7:21 AM Krzysztof Kozlowski <krzk@kernel.org> wrote:
>>>>
>>>> On 19/08/2026 15:26, Ekansh Gupta wrote:
>>>>> On 19-08-2026 00:43, Krzysztof Kozlowski wrote:
>>>>>> On 17/08/2026 06:47, Ekansh Gupta wrote:
>>>>>>> This patch series introduces the Qualcomm DSP Accelerator (QDA) driver,
>>>>>>> a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a
>>>>>>> standardized interface for offloading computational tasks to DSPs found
>>>>>>> on Qualcomm SoCs, supporting all DSP domains.
>>>>>>>
>>>>>>> The QDA driver implements the FastRPC protocol over the DRM accel
>>>>>>> subsystem. It uses the same device-tree node structure as the existing
>>>>>>> fastrpc driver in drivers/misc/. The approach for binding the QDA driver
>>>>>>> to device-tree nodes while coexisting with the fastrpc driver is an open
>>>>>>> item described below.
>>>>>>
>>>>>> No. Grow/replace/improve existing driver instead of coming with a duplicate.
>>>>>>
>>>>>> That's a standard upstream requirement, basically given on every
>>>>>> upstreaming guide.
>>>>>>
>>>>>> Please watch old talk from Greg - "I Don’t Want Your Code!".
>>>>> Posted discussion threads here[1]. Would seek comments from Dmitry,
>>>>> Srini as well.
>>>>>
>>>>> [1]
>>>>> https://lore.kernel.org/all/3476b5c3-7983-4994-a901-3d7d8bd75255@oss.qualcomm.com/
>>>>
>>>> The rest of the comments is still valid even if you did not acknowledge
>>>> them.
>>>>
>>>> Anyway, regarding above - again, watch the talk from Greg.
>>>>
>>>> You have ONE driver. Not two.
>>>
>>> Long term, moving to the common driver framework (which did not exist
>>> when fastrpc was first created) seems like a good thing.  But does
>>> that not allow for some transition period?  How can we get from here
>>> to there without otherwise breaking userspace?  Is there some other
>>> precedent elsewhere in other driver subsystems?
>>
>> Yes, Iris and Venus where we agreed for an exception (two drivers) as
>> long as new driver supports old hardware / features.
>>
>> This is not the case here, right?
>>
>>>
>>> I suppose drm exposing legacy fbdev on top of drm drivers is _sort of_
>>> a precedent if you squint a bit?  I'm not really familiar enough to
>>> say if that would be reasonable/possible in this case.
>>
>> Growing old driver does not look complicated itself. The only a bit
>> tricky thing is to present somehow exclusive interface to user-space,
>> like usage of one disables the second etc. Depending on actual
>> differences in that interface.
>>
>> But having a duplicated driver is a clear no go and it is well known
>> upstream requirement. Nothing new here.
> 
> Would it be acceptable as a first step to, by default (ie. when no
> cmdline override/etc) for the new driver to bind to new compatibles

You have one compatible.

> and the existing driver to old?  Ie. have both drivers but only one
> binds on a given platform?

I don't see how is it possible to write such DTS, because - repeating my
question - how many hardware blocks is there? I believe only one per
given DSP, so how could you have two device nodes?


Best regards,
Krzysztof
Re: [PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver
Posted by Rob Clark 1 month, 1 week ago
On Wed, Aug 19, 2026 at 7:53 AM Krzysztof Kozlowski <krzk@kernel.org> wrote:
>
> On 19/08/2026 16:49, Rob Clark wrote:
> > On Wed, Aug 19, 2026 at 7:43 AM Krzysztof Kozlowski <krzk@kernel.org> wrote:
> >>
> >> On 19/08/2026 16:38, Rob Clark wrote:
> >>> On Wed, Aug 19, 2026 at 7:21 AM Krzysztof Kozlowski <krzk@kernel.org> wrote:
> >>>>
> >>>> On 19/08/2026 15:26, Ekansh Gupta wrote:
> >>>>> On 19-08-2026 00:43, Krzysztof Kozlowski wrote:
> >>>>>> On 17/08/2026 06:47, Ekansh Gupta wrote:
> >>>>>>> This patch series introduces the Qualcomm DSP Accelerator (QDA) driver,
> >>>>>>> a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a
> >>>>>>> standardized interface for offloading computational tasks to DSPs found
> >>>>>>> on Qualcomm SoCs, supporting all DSP domains.
> >>>>>>>
> >>>>>>> The QDA driver implements the FastRPC protocol over the DRM accel
> >>>>>>> subsystem. It uses the same device-tree node structure as the existing
> >>>>>>> fastrpc driver in drivers/misc/. The approach for binding the QDA driver
> >>>>>>> to device-tree nodes while coexisting with the fastrpc driver is an open
> >>>>>>> item described below.
> >>>>>>
> >>>>>> No. Grow/replace/improve existing driver instead of coming with a duplicate.
> >>>>>>
> >>>>>> That's a standard upstream requirement, basically given on every
> >>>>>> upstreaming guide.
> >>>>>>
> >>>>>> Please watch old talk from Greg - "I Don’t Want Your Code!".
> >>>>> Posted discussion threads here[1]. Would seek comments from Dmitry,
> >>>>> Srini as well.
> >>>>>
> >>>>> [1]
> >>>>> https://lore.kernel.org/all/3476b5c3-7983-4994-a901-3d7d8bd75255@oss.qualcomm.com/
> >>>>
> >>>> The rest of the comments is still valid even if you did not acknowledge
> >>>> them.
> >>>>
> >>>> Anyway, regarding above - again, watch the talk from Greg.
> >>>>
> >>>> You have ONE driver. Not two.
> >>>
> >>> Long term, moving to the common driver framework (which did not exist
> >>> when fastrpc was first created) seems like a good thing.  But does
> >>> that not allow for some transition period?  How can we get from here
> >>> to there without otherwise breaking userspace?  Is there some other
> >>> precedent elsewhere in other driver subsystems?
> >>
> >> Yes, Iris and Venus where we agreed for an exception (two drivers) as
> >> long as new driver supports old hardware / features.
> >>
> >> This is not the case here, right?
> >>
> >>>
> >>> I suppose drm exposing legacy fbdev on top of drm drivers is _sort of_
> >>> a precedent if you squint a bit?  I'm not really familiar enough to
> >>> say if that would be reasonable/possible in this case.
> >>
> >> Growing old driver does not look complicated itself. The only a bit
> >> tricky thing is to present somehow exclusive interface to user-space,
> >> like usage of one disables the second etc. Depending on actual
> >> differences in that interface.
> >>
> >> But having a duplicated driver is a clear no go and it is well known
> >> upstream requirement. Nothing new here.
> >
> > Would it be acceptable as a first step to, by default (ie. when no
> > cmdline override/etc) for the new driver to bind to new compatibles
>
> You have one compatible.
>
> > and the existing driver to old?  Ie. have both drivers but only one
> > binds on a given platform?
>
> I don't see how is it possible to write such DTS, because - repeating my
> question - how many hardware blocks is there? I believe only one per
> given DSP, so how could you have two device nodes?

Maybe I'm misunderstanding something here.. I don't see any new
bindings with this series so my assumption is that the bindings are
the same.  What I meant was something more like "qcom,glymur-fastrpc"
would bind to new driver but "qcom,fastrpc" (which seems to be what is
used on older platforms) would bind to the old driver.

So single dts node, but different compatible strings picking which
driver is used.

BR,
-R
Re: [PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver
Posted by Krzysztof Kozlowski 1 month, 1 week ago
On 19/08/2026 17:23, Rob Clark wrote:
>>>>
>>>> Growing old driver does not look complicated itself. The only a bit
>>>> tricky thing is to present somehow exclusive interface to user-space,
>>>> like usage of one disables the second etc. Depending on actual
>>>> differences in that interface.
>>>>
>>>> But having a duplicated driver is a clear no go and it is well known
>>>> upstream requirement. Nothing new here.
>>>
>>> Would it be acceptable as a first step to, by default (ie. when no
>>> cmdline override/etc) for the new driver to bind to new compatibles
>>
>> You have one compatible.
>>
>>> and the existing driver to old?  Ie. have both drivers but only one
>>> binds on a given platform?
>>
>> I don't see how is it possible to write such DTS, because - repeating my
>> question - how many hardware blocks is there? I believe only one per
>> given DSP, so how could you have two device nodes?
> 
> Maybe I'm misunderstanding something here.. I don't see any new
> bindings with this series so my assumption is that the bindings are
> the same.  What I meant was something more like "qcom,glymur-fastrpc"
> would bind to new driver but "qcom,fastrpc" (which seems to be what is
> used on older platforms) would bind to the old driver.
> 
> So single dts node, but different compatible strings picking which
> driver is used.

Glymur is already done, so imagining we talk about next/future SoC then
we would be at point of duplicating drivers for the same hardware. So
back to square one of my comments.

The rule of usptream development is that we do not accept duplicated
code, just because a vendor wants to write something new. This is
basically the concept applied all over the drivers tree, where we pushed
back against all sorts of duplications all over the vendors.

What I miss in this thread is why would there be any exception here. We
do not grant exceptions from standard practices on "I want" reasons.

Best regards,
Krzysztof
Re: [PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver
Posted by Rob Clark 1 month, 1 week ago
On Wed, Aug 19, 2026 at 8:27 AM Krzysztof Kozlowski <krzk@kernel.org> wrote:
>
> On 19/08/2026 17:23, Rob Clark wrote:
> >>>>
> >>>> Growing old driver does not look complicated itself. The only a bit
> >>>> tricky thing is to present somehow exclusive interface to user-space,
> >>>> like usage of one disables the second etc. Depending on actual
> >>>> differences in that interface.
> >>>>
> >>>> But having a duplicated driver is a clear no go and it is well known
> >>>> upstream requirement. Nothing new here.
> >>>
> >>> Would it be acceptable as a first step to, by default (ie. when no
> >>> cmdline override/etc) for the new driver to bind to new compatibles
> >>
> >> You have one compatible.
> >>
> >>> and the existing driver to old?  Ie. have both drivers but only one
> >>> binds on a given platform?
> >>
> >> I don't see how is it possible to write such DTS, because - repeating my
> >> question - how many hardware blocks is there? I believe only one per
> >> given DSP, so how could you have two device nodes?
> >
> > Maybe I'm misunderstanding something here.. I don't see any new
> > bindings with this series so my assumption is that the bindings are
> > the same.  What I meant was something more like "qcom,glymur-fastrpc"
> > would bind to new driver but "qcom,fastrpc" (which seems to be what is
> > used on older platforms) would bind to the old driver.
> >
> > So single dts node, but different compatible strings picking which
> > driver is used.
>
> Glymur is already done, so imagining we talk about next/future SoC then
> we would be at point of duplicating drivers for the same hardware. So
> back to square one of my comments.

I just picked glymur as an example, because I needed to pick
_something_, and toplevel dts for devices available outside of qcom is
only just starting to land.  Maybe it wasn't the right example. Let's
just say "qcom,supercoolfuturething-fastrpc" instead.

I think my example otherwise still stands.

(In either case, kernel cmdline or similar override to get the new
driver instead of old would be a good idea.)

> The rule of usptream development is that we do not accept duplicated
> code, just because a vendor wants to write something new. This is
> basically the concept applied all over the drivers tree, where we pushed
> back against all sorts of duplications all over the vendors.
>
> What I miss in this thread is why would there be any exception here. We
> do not grant exceptions from standard practices on "I want" reasons.

I agree that we should not have duplicated drivers just for vendor
lolz.  But when it comes to adopting common frameworks and integrating
better into the ecosystem, this doesn't seem like something we should
actively discourage.  I don't think this is a case of vendor lolz, but
rather reacting to drm/accel emerging as the standard framework for
this sort of driver.

So how do we get from here to there?

BR,
-R
Re: [PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver
Posted by Krzysztof Kozlowski 1 month, 1 week ago
On 19/08/2026 17:48, Rob Clark wrote:
> On Wed, Aug 19, 2026 at 8:27 AM Krzysztof Kozlowski <krzk@kernel.org> wrote:
>> The rule of usptream development is that we do not accept duplicated
>> code, just because a vendor wants to write something new. This is
>> basically the concept applied all over the drivers tree, where we pushed
>> back against all sorts of duplications all over the vendors.
>>
>> What I miss in this thread is why would there be any exception here. We
>> do not grant exceptions from standard practices on "I want" reasons.
> 
> I agree that we should not have duplicated drivers just for vendor
> lolz.  But when it comes to adopting common frameworks and integrating
> better into the ecosystem, this doesn't seem like something we should
> actively discourage.  I don't think this is a case of vendor lolz, but

No one discourages it. Following standard Linux kernel practices and
requirements is not discouraging, do not twist the narrative here.
Again, it is standard upstream review telling that we do not duplicate
drivers. Ever, unless there is serious exception needed.

I asked why there should be an exception granted? Is the reason for
exception following:
"We want to adopt common framework"
?

> rather reacting to drm/accel emerging as the standard framework for
> this sort of driver.
> 
> So how do we get from here to there?

What is wrong with my proposal?

Best regards,
Krzysztof
Re: [PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver
Posted by Rob Clark 1 month, 1 week ago
On Wed, Aug 19, 2026 at 11:15 PM Krzysztof Kozlowski <krzk@kernel.org> wrote:
>
> On 19/08/2026 17:48, Rob Clark wrote:
> > On Wed, Aug 19, 2026 at 8:27 AM Krzysztof Kozlowski <krzk@kernel.org> wrote:
> >> The rule of usptream development is that we do not accept duplicated
> >> code, just because a vendor wants to write something new. This is
> >> basically the concept applied all over the drivers tree, where we pushed
> >> back against all sorts of duplications all over the vendors.
> >>
> >> What I miss in this thread is why would there be any exception here. We
> >> do not grant exceptions from standard practices on "I want" reasons.
> >
> > I agree that we should not have duplicated drivers just for vendor
> > lolz.  But when it comes to adopting common frameworks and integrating
> > better into the ecosystem, this doesn't seem like something we should
> > actively discourage.  I don't think this is a case of vendor lolz, but
>
> No one discourages it. Following standard Linux kernel practices and
> requirements is not discouraging, do not twist the narrative here.
> Again, it is standard upstream review telling that we do not duplicate
> drivers. Ever, unless there is serious exception needed.

I wasn't trying to twist the narrative, just trying to come up with a
path forward that isn't "no" or "improve existing driver", since
neither of those gets us towards a future using common frameworks.

> I asked why there should be an exception granted? Is the reason for
> exception following:
> "We want to adopt common framework"
> ?

Possibly?  But I don't think we want two drivers to be any sort of
long term solution.  (Ie. as long as venus/iris have co-exist.)

>
> > rather reacting to drm/accel emerging as the standard framework for
> > this sort of driver.
> >
> > So how do we get from here to there?
>
> What is wrong with my proposal?

Maybe I missed something, my understanding was your proposal was
"Grow/replace/improve existing driver instead of coming with a
duplicate"..  grow or improve doesn't move us toward common
frameworks.  Maybe "replace" is a valid option.  If there is something
I missed, then I apologize.

Options I can think of are:

1. Hardware cutoff.. new hw gets new driver, existing hw gets existing
   driver
2. Backwards compat chardev registered by new driver, providing existing
   UABI.  I'm not 100% sure about the feasibility/drawbacks of this..
   AFAIU the fastrpc folks where planning a backwards compat layer in
   userspace, so maybe it is possible.
3. exception?

I'd like to know what the feasibility of #2 is, since at a high level
that sounds like the best option.  Possibly limit exposure of legacy
UABI to existing hw so we don't get into a place of needing to extend
the legacy UABI for new hw?

But #1 sounds like a non-controversial place to start regardless.
Possibly with #2 coming as followup and necessary step before eventual
migration to new driver for existing hw?

Even if we start with #2, how do we handle first-merge-window
bugs/regressions without reverting addition of new driver and removal
of old?  It seems like we'd need a window of a couple release cycles
where both drivers exist?

Maybe others have other/better options in mind?

BR,
-R
Re: [PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver
Posted by Ekansh Gupta 1 month ago
On 20-08-2026 20:17, Rob Clark wrote:
> On Wed, Aug 19, 2026 at 11:15 PM Krzysztof Kozlowski <krzk@kernel.org> wrote:
>>
>> On 19/08/2026 17:48, Rob Clark wrote:
>>> On Wed, Aug 19, 2026 at 8:27 AM Krzysztof Kozlowski <krzk@kernel.org> wrote:
>>>> The rule of usptream development is that we do not accept duplicated
>>>> code, just because a vendor wants to write something new. This is
>>>> basically the concept applied all over the drivers tree, where we pushed
>>>> back against all sorts of duplications all over the vendors.
>>>>
>>>> What I miss in this thread is why would there be any exception here. We
>>>> do not grant exceptions from standard practices on "I want" reasons.
>>>
>>> I agree that we should not have duplicated drivers just for vendor
>>> lolz.  But when it comes to adopting common frameworks and integrating
>>> better into the ecosystem, this doesn't seem like something we should
>>> actively discourage.  I don't think this is a case of vendor lolz, but
>>
>> No one discourages it. Following standard Linux kernel practices and
>> requirements is not discouraging, do not twist the narrative here.
>> Again, it is standard upstream review telling that we do not duplicate
>> drivers. Ever, unless there is serious exception needed.
> 
> I wasn't trying to twist the narrative, just trying to come up with a
> path forward that isn't "no" or "improve existing driver", since
> neither of those gets us towards a future using common frameworks.
> 
>> I asked why there should be an exception granted? Is the reason for
>> exception following:
>> "We want to adopt common framework"
>> ?
> 
> Possibly?  But I don't think we want two drivers to be any sort of
> long term solution.  (Ie. as long as venus/iris have co-exist.)
> 
>>
>>> rather reacting to drm/accel emerging as the standard framework for
>>> this sort of driver.
>>>
>>> So how do we get from here to there?
>>
>> What is wrong with my proposal?
> 
> Maybe I missed something, my understanding was your proposal was
> "Grow/replace/improve existing driver instead of coming with a
> duplicate"..  grow or improve doesn't move us toward common
> frameworks.  Maybe "replace" is a valid option.  If there is something
> I missed, then I apologize.
> 
> Options I can think of are:
> 
> 1. Hardware cutoff.. new hw gets new driver, existing hw gets existing
>    driver
> 2. Backwards compat chardev registered by new driver, providing existing
>    UABI.  I'm not 100% sure about the feasibility/drawbacks of this..
>    AFAIU the fastrpc folks where planning a backwards compat layer in
>    userspace, so maybe it is possible.
> 3. exception?
> 
> I'd like to know what the feasibility of #2 is, since at a high level
> that sounds like the best option.  Possibly limit exposure of legacy
> UABI to existing hw so we don't get into a place of needing to extend
> the legacy UABI for new hw?
> 
> But #1 sounds like a non-controversial place to start regardless.
> Possibly with #2 coming as followup and necessary step before eventual
> migration to new driver for existing hw?
> 
> Even if we start with #2, how do we handle first-merge-window
> bugs/regressions without reverting addition of new driver and removal
> of old?  It seems like we'd need a window of a couple release cycles
> where both drivers exist?
> 
> Maybe others have other/better options in mind?
To all, I'm seeking on the approach I should follow to go ahead here. I
can work on implementing #1(as per Rob's list) with hw specific
compatible for v4 if it's acceptable.

#2(compat driver) is something that we are still exploring as we
couldn't find any standard way to achieve it. We might start a separate
discussion for that once we have few possible designs with us.

Happy to take any other suggestion also.>
> BR,
> -R

Re: [PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver
Posted by Bjorn Andersson 2 weeks, 5 days ago
On Wed, Aug 26, 2026 at 06:37:00PM +0530, Ekansh Gupta wrote:
> On 20-08-2026 20:17, Rob Clark wrote:
> > On Wed, Aug 19, 2026 at 11:15 PM Krzysztof Kozlowski <krzk@kernel.org> wrote:
> >>
> >> On 19/08/2026 17:48, Rob Clark wrote:
> >>> On Wed, Aug 19, 2026 at 8:27 AM Krzysztof Kozlowski <krzk@kernel.org> wrote:
> >>>> The rule of usptream development is that we do not accept duplicated
> >>>> code, just because a vendor wants to write something new. This is
> >>>> basically the concept applied all over the drivers tree, where we pushed
> >>>> back against all sorts of duplications all over the vendors.
> >>>>
> >>>> What I miss in this thread is why would there be any exception here. We
> >>>> do not grant exceptions from standard practices on "I want" reasons.
> >>>
> >>> I agree that we should not have duplicated drivers just for vendor
> >>> lolz.  But when it comes to adopting common frameworks and integrating
> >>> better into the ecosystem, this doesn't seem like something we should
> >>> actively discourage.  I don't think this is a case of vendor lolz, but
> >>
> >> No one discourages it. Following standard Linux kernel practices and
> >> requirements is not discouraging, do not twist the narrative here.
> >> Again, it is standard upstream review telling that we do not duplicate
> >> drivers. Ever, unless there is serious exception needed.
> > 
> > I wasn't trying to twist the narrative, just trying to come up with a
> > path forward that isn't "no" or "improve existing driver", since
> > neither of those gets us towards a future using common frameworks.
> > 
> >> I asked why there should be an exception granted? Is the reason for
> >> exception following:
> >> "We want to adopt common framework"
> >> ?
> > 
> > Possibly?  But I don't think we want two drivers to be any sort of
> > long term solution.  (Ie. as long as venus/iris have co-exist.)
> > 
> >>
> >>> rather reacting to drm/accel emerging as the standard framework for
> >>> this sort of driver.
> >>>
> >>> So how do we get from here to there?
> >>
> >> What is wrong with my proposal?
> > 
> > Maybe I missed something, my understanding was your proposal was
> > "Grow/replace/improve existing driver instead of coming with a
> > duplicate"..  grow or improve doesn't move us toward common
> > frameworks.  Maybe "replace" is a valid option.  If there is something
> > I missed, then I apologize.
> > 
> > Options I can think of are:
> > 
> > 1. Hardware cutoff.. new hw gets new driver, existing hw gets existing
> >    driver
> > 2. Backwards compat chardev registered by new driver, providing existing
> >    UABI.  I'm not 100% sure about the feasibility/drawbacks of this..
> >    AFAIU the fastrpc folks where planning a backwards compat layer in
> >    userspace, so maybe it is possible.
> > 3. exception?
> > 
> > I'd like to know what the feasibility of #2 is, since at a high level
> > that sounds like the best option.  Possibly limit exposure of legacy
> > UABI to existing hw so we don't get into a place of needing to extend
> > the legacy UABI for new hw?
> > 
> > But #1 sounds like a non-controversial place to start regardless.
> > Possibly with #2 coming as followup and necessary step before eventual
> > migration to new driver for existing hw?
> > 
> > Even if we start with #2, how do we handle first-merge-window
> > bugs/regressions without reverting addition of new driver and removal
> > of old?  It seems like we'd need a window of a couple release cycles
> > where both drivers exist?
> > 
> > Maybe others have other/better options in mind?
> To all, I'm seeking on the approach I should follow to go ahead here. I
> can work on implementing #1(as per Rob's list) with hw specific
> compatible for v4 if it's acceptable.
> 

I don't see any reason for you to define a "hw specific compatible",
because as you have shown in this series (and as Rob point out), there's
no difference in the "hardware".

The only reason for your "hw specific compatible" is to make a software
selection in Linux - and that's not what DeviceTree is for.


As such, I don't see that you have a DeviceTree problem at all, because
this is a Linux-internal problem.

> #2(compat driver) is something that we are still exploring as we
> couldn't find any standard way to achieve it. We might start a separate
> discussion for that once we have few possible designs with us.
> 

This is the actual problem!

We have existing user space that depends on the ioctl interface exposed
by the current misc driver. You must not break these.

Hardware cutoff is not a viable solution, because that's just a
declaration that we'll let the old platforms rotten - or alternatively
you commit to maintain two drivers to the very same feature and quality
level.

So the only reasonable solution is #2; from there it's a valid question
if you reach that point my stepwise migrating the current misc driver
that solution, or if you present a new driver with the fully backwards
compatible interface, alongside the new ABI.


But this does bring to a question which the cover letter should explain
- but doesn't: what problem does this patch series actually solve?

Regards,
Bjorn

> Happy to take any other suggestion also.>
> > BR,
> > -R
> 
> 
Re: [PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver
Posted by Dmitry Baryshkov 2 weeks, 5 days ago
On Tue, Sep 08, 2026 at 05:44:10PM -0500, Bjorn Andersson wrote:
> On Wed, Aug 26, 2026 at 06:37:00PM +0530, Ekansh Gupta wrote:
> > On 20-08-2026 20:17, Rob Clark wrote:
> > > On Wed, Aug 19, 2026 at 11:15 PM Krzysztof Kozlowski <krzk@kernel.org> wrote:
> > >>
> > >> On 19/08/2026 17:48, Rob Clark wrote:
> > >>> On Wed, Aug 19, 2026 at 8:27 AM Krzysztof Kozlowski <krzk@kernel.org> wrote:
> > >>>> The rule of usptream development is that we do not accept duplicated
> > >>>> code, just because a vendor wants to write something new. This is
> > >>>> basically the concept applied all over the drivers tree, where we pushed
> > >>>> back against all sorts of duplications all over the vendors.
> > >>>>
> > >>>> What I miss in this thread is why would there be any exception here. We
> > >>>> do not grant exceptions from standard practices on "I want" reasons.
> > >>>
> > >>> I agree that we should not have duplicated drivers just for vendor
> > >>> lolz.  But when it comes to adopting common frameworks and integrating
> > >>> better into the ecosystem, this doesn't seem like something we should
> > >>> actively discourage.  I don't think this is a case of vendor lolz, but
> > >>
> > >> No one discourages it. Following standard Linux kernel practices and
> > >> requirements is not discouraging, do not twist the narrative here.
> > >> Again, it is standard upstream review telling that we do not duplicate
> > >> drivers. Ever, unless there is serious exception needed.
> > > 
> > > I wasn't trying to twist the narrative, just trying to come up with a
> > > path forward that isn't "no" or "improve existing driver", since
> > > neither of those gets us towards a future using common frameworks.
> > > 
> > >> I asked why there should be an exception granted? Is the reason for
> > >> exception following:
> > >> "We want to adopt common framework"
> > >> ?
> > > 
> > > Possibly?  But I don't think we want two drivers to be any sort of
> > > long term solution.  (Ie. as long as venus/iris have co-exist.)
> > > 
> > >>
> > >>> rather reacting to drm/accel emerging as the standard framework for
> > >>> this sort of driver.
> > >>>
> > >>> So how do we get from here to there?
> > >>
> > >> What is wrong with my proposal?
> > > 
> > > Maybe I missed something, my understanding was your proposal was
> > > "Grow/replace/improve existing driver instead of coming with a
> > > duplicate"..  grow or improve doesn't move us toward common
> > > frameworks.  Maybe "replace" is a valid option.  If there is something
> > > I missed, then I apologize.
> > > 
> > > Options I can think of are:
> > > 
> > > 1. Hardware cutoff.. new hw gets new driver, existing hw gets existing
> > >    driver
> > > 2. Backwards compat chardev registered by new driver, providing existing
> > >    UABI.  I'm not 100% sure about the feasibility/drawbacks of this..
> > >    AFAIU the fastrpc folks where planning a backwards compat layer in
> > >    userspace, so maybe it is possible.
> > > 3. exception?
> > > 
> > > I'd like to know what the feasibility of #2 is, since at a high level
> > > that sounds like the best option.  Possibly limit exposure of legacy
> > > UABI to existing hw so we don't get into a place of needing to extend
> > > the legacy UABI for new hw?
> > > 
> > > But #1 sounds like a non-controversial place to start regardless.
> > > Possibly with #2 coming as followup and necessary step before eventual
> > > migration to new driver for existing hw?
> > > 
> > > Even if we start with #2, how do we handle first-merge-window
> > > bugs/regressions without reverting addition of new driver and removal
> > > of old?  It seems like we'd need a window of a couple release cycles
> > > where both drivers exist?
> > > 
> > > Maybe others have other/better options in mind?
> > To all, I'm seeking on the approach I should follow to go ahead here. I
> > can work on implementing #1(as per Rob's list) with hw specific
> > compatible for v4 if it's acceptable.
> > 
> 
> I don't see any reason for you to define a "hw specific compatible",
> because as you have shown in this series (and as Rob point out), there's
> no difference in the "hardware".
> 
> The only reason for your "hw specific compatible" is to make a software
> selection in Linux - and that's not what DeviceTree is for.

That's not exactly true. There are protocol differences. For example,
polling mode is supported only since a certain timeline in the history.
Likewise other interface features are not supported on all the FastRPC
devices. For the polling mode support we were already beaten by the lack
of SoC-specific compats.

> As such, I don't see that you have a DeviceTree problem at all, because
> this is a Linux-internal problem.
> 
> > #2(compat driver) is something that we are still exploring as we
> > couldn't find any standard way to achieve it. We might start a separate
> > discussion for that once we have few possible designs with us.
> > 
> 
> This is the actual problem!
> 
> We have existing user space that depends on the ioctl interface exposed
> by the current misc driver. You must not break these.

This is clear.

> Hardware cutoff is not a viable solution, because that's just a
> declaration that we'll let the old platforms rotten - or alternatively
> you commit to maintain two drivers to the very same feature and quality
> level.
> 
> So the only reasonable solution is #2; from there it's a valid question
> if you reach that point my stepwise migrating the current misc driver
> that solution, or if you present a new driver with the fully backwards
> compatible interface, alongside the new ABI.

I think the general direction was #3 (or #2.1): implement a shim layer
on top of the QDA driver as a separate module. Put all the historical
over-complicated solutions into that shim module and let it die at some
point. Current fastrpc driver lets userspace specify buffers in several
different ways, forcing the kernel driver to perform a lot of work
with buffer addresses. I don't think that this legacy code should be a
part of the QDA. It can go to qda-fastrpc-shim, keeping it out of the
kernel memory if it's not necessary.

> But this does bring to a question which the cover letter should explain
> - but doesn't: what problem does this patch series actually solve?

I agree that it should be a part of the cover letter.

As a person who triggered this work, I can propose my reasons:

- Current driver has over-complicated memory manager (both on the
  userspace and on the kernel side). Correspondng uAPI is not really
  suitable for virtualization. Using GEM simplifies both the kernel
  driver and uAPI. Also using handle-offset-length to specify the
  buffers makes it easy to support virtual QDA devices.

- Current driver predates the accel subsystem. Using common subsystem
  simplifies reviews. The QDA driver has gotten several comments about
  the usage of the DMA-BUFs. It'not unlikely that the same issues are
  present in the current FastRPC driver, just being unnoticed.

- The ideas present in the current FastRPC driver also predate the
  current design practices. The uAPI was created in the ad-hoc way, just
  following the momentary needs. Driver code also shows the result of
  that, having enough of the spaghetti code.

Given all of that, yes, it is possible to provide an evolution of the
FastRPC driver into the accel+shim, improve the code quality meanwhile
and end up with the good enough split. However I think that the path
taken would be longer and the net result might be worse.

With all of that in mind, my suggestion is to continue working on the
QDA driver, get the core of it to integrate nicely with the accel and
DMA-BUF usage requirements and implement the (minimal?) fastrpc shim on
top of it.


-- 
With best wishes
Dmitry
Re: [PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver
Posted by Ekansh Gupta 2 weeks, 3 days ago
On 09-09-2026 17:18, Dmitry Baryshkov wrote:
> On Tue, Sep 08, 2026 at 05:44:10PM -0500, Bjorn Andersson wrote:
>> On Wed, Aug 26, 2026 at 06:37:00PM +0530, Ekansh Gupta wrote:
>>> On 20-08-2026 20:17, Rob Clark wrote:
>>>> On Wed, Aug 19, 2026 at 11:15 PM Krzysztof Kozlowski <krzk@kernel.org> wrote:
>>>>>
>>>>> On 19/08/2026 17:48, Rob Clark wrote:
>>>>>> On Wed, Aug 19, 2026 at 8:27 AM Krzysztof Kozlowski <krzk@kernel.org> wrote:
>>>>>>> The rule of usptream development is that we do not accept duplicated
>>>>>>> code, just because a vendor wants to write something new. This is
>>>>>>> basically the concept applied all over the drivers tree, where we pushed
>>>>>>> back against all sorts of duplications all over the vendors.
>>>>>>>
>>>>>>> What I miss in this thread is why would there be any exception here. We
>>>>>>> do not grant exceptions from standard practices on "I want" reasons.
>>>>>>
>>>>>> I agree that we should not have duplicated drivers just for vendor
>>>>>> lolz.  But when it comes to adopting common frameworks and integrating
>>>>>> better into the ecosystem, this doesn't seem like something we should
>>>>>> actively discourage.  I don't think this is a case of vendor lolz, but
>>>>>
>>>>> No one discourages it. Following standard Linux kernel practices and
>>>>> requirements is not discouraging, do not twist the narrative here.
>>>>> Again, it is standard upstream review telling that we do not duplicate
>>>>> drivers. Ever, unless there is serious exception needed.
>>>>
>>>> I wasn't trying to twist the narrative, just trying to come up with a
>>>> path forward that isn't "no" or "improve existing driver", since
>>>> neither of those gets us towards a future using common frameworks.
>>>>
>>>>> I asked why there should be an exception granted? Is the reason for
>>>>> exception following:
>>>>> "We want to adopt common framework"
>>>>> ?
>>>>
>>>> Possibly?  But I don't think we want two drivers to be any sort of
>>>> long term solution.  (Ie. as long as venus/iris have co-exist.)
>>>>
>>>>>
>>>>>> rather reacting to drm/accel emerging as the standard framework for
>>>>>> this sort of driver.
>>>>>>
>>>>>> So how do we get from here to there?
>>>>>
>>>>> What is wrong with my proposal?
>>>>
>>>> Maybe I missed something, my understanding was your proposal was
>>>> "Grow/replace/improve existing driver instead of coming with a
>>>> duplicate"..  grow or improve doesn't move us toward common
>>>> frameworks.  Maybe "replace" is a valid option.  If there is something
>>>> I missed, then I apologize.
>>>>
>>>> Options I can think of are:
>>>>
>>>> 1. Hardware cutoff.. new hw gets new driver, existing hw gets existing
>>>>    driver
>>>> 2. Backwards compat chardev registered by new driver, providing existing
>>>>    UABI.  I'm not 100% sure about the feasibility/drawbacks of this..
>>>>    AFAIU the fastrpc folks where planning a backwards compat layer in
>>>>    userspace, so maybe it is possible.
>>>> 3. exception?
>>>>
>>>> I'd like to know what the feasibility of #2 is, since at a high level
>>>> that sounds like the best option.  Possibly limit exposure of legacy
>>>> UABI to existing hw so we don't get into a place of needing to extend
>>>> the legacy UABI for new hw?
>>>>
>>>> But #1 sounds like a non-controversial place to start regardless.
>>>> Possibly with #2 coming as followup and necessary step before eventual
>>>> migration to new driver for existing hw?
>>>>
>>>> Even if we start with #2, how do we handle first-merge-window
>>>> bugs/regressions without reverting addition of new driver and removal
>>>> of old?  It seems like we'd need a window of a couple release cycles
>>>> where both drivers exist?
>>>>
>>>> Maybe others have other/better options in mind?
>>> To all, I'm seeking on the approach I should follow to go ahead here. I
>>> can work on implementing #1(as per Rob's list) with hw specific
>>> compatible for v4 if it's acceptable.
>>>
>>
>> I don't see any reason for you to define a "hw specific compatible",
>> because as you have shown in this series (and as Rob point out), there's
>> no difference in the "hardware".
>>
>> The only reason for your "hw specific compatible" is to make a software
>> selection in Linux - and that's not what DeviceTree is for.
> 
> That's not exactly true. There are protocol differences. For example,
> polling mode is supported only since a certain timeline in the history.
> Likewise other interface features are not supported on all the FastRPC
> devices. For the polling mode support we were already beaten by the lack
> of SoC-specific compats.
> 
>> As such, I don't see that you have a DeviceTree problem at all, because
>> this is a Linux-internal problem.
>>
>>> #2(compat driver) is something that we are still exploring as we
>>> couldn't find any standard way to achieve it. We might start a separate
>>> discussion for that once we have few possible designs with us.
>>>
>>
>> This is the actual problem!
>>
>> We have existing user space that depends on the ioctl interface exposed
>> by the current misc driver. You must not break these.
> 
> This is clear.
> 
>> Hardware cutoff is not a viable solution, because that's just a
>> declaration that we'll let the old platforms rotten - or alternatively
>> you commit to maintain two drivers to the very same feature and quality
>> level.
>>
>> So the only reasonable solution is #2; from there it's a valid question
>> if you reach that point my stepwise migrating the current misc driver
>> that solution, or if you present a new driver with the fully backwards
>> compatible interface, alongside the new ABI.
> 
> I think the general direction was #3 (or #2.1): implement a shim layer
> on top of the QDA driver as a separate module. Put all the historical
> over-complicated solutions into that shim module and let it die at some
> point. Current fastrpc driver lets userspace specify buffers in several
> different ways, forcing the kernel driver to perform a lot of work
> with buffer addresses. I don't think that this legacy code should be a
> part of the QDA. It can go to qda-fastrpc-shim, keeping it out of the
> kernel memory if it's not necessary.
> 
>> But this does bring to a question which the cover letter should explain
>> - but doesn't: what problem does this patch series actually solve?
> 
> I agree that it should be a part of the cover letter.
> 
> As a person who triggered this work, I can propose my reasons:
> 
> - Current driver has over-complicated memory manager (both on the
>   userspace and on the kernel side). Correspondng uAPI is not really
>   suitable for virtualization. Using GEM simplifies both the kernel
>   driver and uAPI. Also using handle-offset-length to specify the
>   buffers makes it easy to support virtual QDA devices.
> 
> - Current driver predates the accel subsystem. Using common subsystem
>   simplifies reviews. The QDA driver has gotten several comments about
>   the usage of the DMA-BUFs. It'not unlikely that the same issues are
>   present in the current FastRPC driver, just being unnoticed.
> 
> - The ideas present in the current FastRPC driver also predate the
>   current design practices. The uAPI was created in the ad-hoc way, just
>   following the momentary needs. Driver code also shows the result of
>   that, having enough of the spaghetti code.
> 
> Given all of that, yes, it is possible to provide an evolution of the
> FastRPC driver into the accel+shim, improve the code quality meanwhile
> and end up with the good enough split. However I think that the path
> taken would be longer and the net result might be worse.
> 
> With all of that in mind, my suggestion is to continue working on the
> QDA driver, get the core of it to integrate nicely with the accel and
> DMA-BUF usage requirements and implement the (minimal?) fastrpc shim on
> top of it.
> 
Agreed on all the above points. We discussed this internally and are
exploring a shim module(something like qda_fastrpc_shim.ko) along these
lines, sitting on top of the QDA driver rather than growing fastrpc.c in
place. Still working through some of the details (e.g. who owns the
rpmsg binding, and what "minimal" means precisely given feature parity
is expected for the compat layer's behaviour). Will follow up with a
design doc here once that's settled.

To all, one separate question on daemon-attach security for feature
parity: would running the daemon with CAP_SYS_ADMIN, and having the
driver reject attach requests from any process without that capability,
be an acceptable approach? The concern is preventing a
malicious/unprivileged application from attaching to a privileged DSP PD
and disrupting it.

Thanks for the detailed feedback.

//Ekansh

Re: [PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver
Posted by Bjorn Andersson 2 weeks, 4 days ago
On Wed, Sep 09, 2026 at 02:48:17PM +0300, Dmitry Baryshkov wrote:
> On Tue, Sep 08, 2026 at 05:44:10PM -0500, Bjorn Andersson wrote:
> > On Wed, Aug 26, 2026 at 06:37:00PM +0530, Ekansh Gupta wrote:
> > > On 20-08-2026 20:17, Rob Clark wrote:
> > > > On Wed, Aug 19, 2026 at 11:15 PM Krzysztof Kozlowski <krzk@kernel.org> wrote:
> > > >>
> > > >> On 19/08/2026 17:48, Rob Clark wrote:
> > > >>> On Wed, Aug 19, 2026 at 8:27 AM Krzysztof Kozlowski <krzk@kernel.org> wrote:
> > > >>>> The rule of usptream development is that we do not accept duplicated
> > > >>>> code, just because a vendor wants to write something new. This is
> > > >>>> basically the concept applied all over the drivers tree, where we pushed
> > > >>>> back against all sorts of duplications all over the vendors.
> > > >>>>
> > > >>>> What I miss in this thread is why would there be any exception here. We
> > > >>>> do not grant exceptions from standard practices on "I want" reasons.
> > > >>>
> > > >>> I agree that we should not have duplicated drivers just for vendor
> > > >>> lolz.  But when it comes to adopting common frameworks and integrating
> > > >>> better into the ecosystem, this doesn't seem like something we should
> > > >>> actively discourage.  I don't think this is a case of vendor lolz, but
> > > >>
> > > >> No one discourages it. Following standard Linux kernel practices and
> > > >> requirements is not discouraging, do not twist the narrative here.
> > > >> Again, it is standard upstream review telling that we do not duplicate
> > > >> drivers. Ever, unless there is serious exception needed.
> > > > 
> > > > I wasn't trying to twist the narrative, just trying to come up with a
> > > > path forward that isn't "no" or "improve existing driver", since
> > > > neither of those gets us towards a future using common frameworks.
> > > > 
> > > >> I asked why there should be an exception granted? Is the reason for
> > > >> exception following:
> > > >> "We want to adopt common framework"
> > > >> ?
> > > > 
> > > > Possibly?  But I don't think we want two drivers to be any sort of
> > > > long term solution.  (Ie. as long as venus/iris have co-exist.)
> > > > 
> > > >>
> > > >>> rather reacting to drm/accel emerging as the standard framework for
> > > >>> this sort of driver.
> > > >>>
> > > >>> So how do we get from here to there?
> > > >>
> > > >> What is wrong with my proposal?
> > > > 
> > > > Maybe I missed something, my understanding was your proposal was
> > > > "Grow/replace/improve existing driver instead of coming with a
> > > > duplicate"..  grow or improve doesn't move us toward common
> > > > frameworks.  Maybe "replace" is a valid option.  If there is something
> > > > I missed, then I apologize.
> > > > 
> > > > Options I can think of are:
> > > > 
> > > > 1. Hardware cutoff.. new hw gets new driver, existing hw gets existing
> > > >    driver
> > > > 2. Backwards compat chardev registered by new driver, providing existing
> > > >    UABI.  I'm not 100% sure about the feasibility/drawbacks of this..
> > > >    AFAIU the fastrpc folks where planning a backwards compat layer in
> > > >    userspace, so maybe it is possible.
> > > > 3. exception?
> > > > 
> > > > I'd like to know what the feasibility of #2 is, since at a high level
> > > > that sounds like the best option.  Possibly limit exposure of legacy
> > > > UABI to existing hw so we don't get into a place of needing to extend
> > > > the legacy UABI for new hw?
> > > > 
> > > > But #1 sounds like a non-controversial place to start regardless.
> > > > Possibly with #2 coming as followup and necessary step before eventual
> > > > migration to new driver for existing hw?
> > > > 
> > > > Even if we start with #2, how do we handle first-merge-window
> > > > bugs/regressions without reverting addition of new driver and removal
> > > > of old?  It seems like we'd need a window of a couple release cycles
> > > > where both drivers exist?
> > > > 
> > > > Maybe others have other/better options in mind?
> > > To all, I'm seeking on the approach I should follow to go ahead here. I
> > > can work on implementing #1(as per Rob's list) with hw specific
> > > compatible for v4 if it's acceptable.
> > > 
> > 
> > I don't see any reason for you to define a "hw specific compatible",
> > because as you have shown in this series (and as Rob point out), there's
> > no difference in the "hardware".
> > 
> > The only reason for your "hw specific compatible" is to make a software
> > selection in Linux - and that's not what DeviceTree is for.
> 
> That's not exactly true. There are protocol differences. For example,
> polling mode is supported only since a certain timeline in the history.
> Likewise other interface features are not supported on all the FastRPC
> devices. For the polling mode support we were already beaten by the lack
> of SoC-specific compats.
> 

I can see the benefit of capturing some of the generational features in
a compatible, like the changes related to address width. But for pure
software features that doesn't have an actual bearing in the hardware,
I'd prefer if we relied on dynamic discovery.

But none of that applies to the question of "can I use compatible to
select if we should use the new or old Linux driver".

> > As such, I don't see that you have a DeviceTree problem at all, because
> > this is a Linux-internal problem.
> > 
> > > #2(compat driver) is something that we are still exploring as we
> > > couldn't find any standard way to achieve it. We might start a separate
> > > discussion for that once we have few possible designs with us.
> > > 
> > 
> > This is the actual problem!
> > 
> > We have existing user space that depends on the ioctl interface exposed
> > by the current misc driver. You must not break these.
> 
> This is clear.
> 
> > Hardware cutoff is not a viable solution, because that's just a
> > declaration that we'll let the old platforms rotten - or alternatively
> > you commit to maintain two drivers to the very same feature and quality
> > level.
> > 
> > So the only reasonable solution is #2; from there it's a valid question
> > if you reach that point my stepwise migrating the current misc driver
> > that solution, or if you present a new driver with the fully backwards
> > compatible interface, alongside the new ABI.
> 
> I think the general direction was #3 (or #2.1): implement a shim layer
> on top of the QDA driver as a separate module. Put all the historical
> over-complicated solutions into that shim module and let it die at some
> point. Current fastrpc driver lets userspace specify buffers in several
> different ways, forcing the kernel driver to perform a lot of work
> with buffer addresses. I don't think that this legacy code should be a
> part of the QDA. It can go to qda-fastrpc-shim, keeping it out of the
> kernel memory if it's not necessary.
> 
> > But this does bring to a question which the cover letter should explain
> > - but doesn't: what problem does this patch series actually solve?
> 
> I agree that it should be a part of the cover letter.
> 
> As a person who triggered this work, I can propose my reasons:
> 
> - Current driver has over-complicated memory manager (both on the
>   userspace and on the kernel side).

The userspace library is spaghetti, but when I wrote my own it turned
out quite succinct. The ioctl structs certainly could have been cleaner,
and documented, but it seems to me that a fair amount of the complexity
comes from the different use cases - such as SMMU vs XPU, secure and
unsecure buffers, remnants of now unsupported options.

I'm presuming that QDA will need to adopt most of these, and that QDA
support will be bolted into the spaghetti library.

> Correspondng uAPI is not really
>   suitable for virtualization. Using GEM simplifies both the kernel
>   driver and uAPI. Also using handle-offset-length to specify the
>   buffers makes it easy to support virtual QDA devices.

I'm looking forward to learn more about this!

> 
> - Current driver predates the accel subsystem. Using common subsystem
>   simplifies reviews. The QDA driver has gotten several comments about
>   the usage of the DMA-BUFs. It'not unlikely that the same issues are
>   present in the current FastRPC driver, just being unnoticed.

Yeah, this is unfortunate. It would certainly be nice to have a
documented and clean ABI.

> 
> - The ideas present in the current FastRPC driver also predate the
>   current design practices. The uAPI was created in the ad-hoc way, just
>   following the momentary needs. Driver code also shows the result of
>   that, having enough of the spaghetti code.
> 

Yeah, again, this isn't desirable.

> Given all of that, yes, it is possible to provide an evolution of the
> FastRPC driver into the accel+shim, improve the code quality meanwhile
> and end up with the good enough split. However I think that the path
> taken would be longer and the net result might be worse.
> 
> With all of that in mind, my suggestion is to continue working on the
> QDA driver, get the core of it to integrate nicely with the accel and
> DMA-BUF usage requirements and implement the (minimal?) fastrpc shim on
> top of it.
> 

If you believe this is the way to reach the proper design, then I won't
object. My requirement is that my userspace continues to work when my
distro suddenly switches FASTRPC=n/QDA=m.

I have no problems with dropping the fastrpc driver once the QDA is
drop-in-compatible. I'm also open to marking the compat layer deprecated
and eventually drop it once we're certain that users have moved to a
userspace that used the accel interface.

But we can't merge the QDA driver as long as we believe that
implementing a compat layer will be hard/impossible.

Regards,
Bjorn

> 
> -- 
> With best wishes
> Dmitry
Re: [PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver
Posted by Rob Clark 2 weeks, 4 days ago
On Thu, Sep 10, 2026 at 7:25 AM Bjorn Andersson <andersson@kernel.org> wrote:
>
> On Wed, Sep 09, 2026 at 02:48:17PM +0300, Dmitry Baryshkov wrote:
> > On Tue, Sep 08, 2026 at 05:44:10PM -0500, Bjorn Andersson wrote:
> > > On Wed, Aug 26, 2026 at 06:37:00PM +0530, Ekansh Gupta wrote:
> > > > On 20-08-2026 20:17, Rob Clark wrote:
> > > > > On Wed, Aug 19, 2026 at 11:15 PM Krzysztof Kozlowski <krzk@kernel.org> wrote:
> > > > >>
> > > > >> On 19/08/2026 17:48, Rob Clark wrote:
> > > > >>> On Wed, Aug 19, 2026 at 8:27 AM Krzysztof Kozlowski <krzk@kernel.org> wrote:
> > > > >>>> The rule of usptream development is that we do not accept duplicated
> > > > >>>> code, just because a vendor wants to write something new. This is
> > > > >>>> basically the concept applied all over the drivers tree, where we pushed
> > > > >>>> back against all sorts of duplications all over the vendors.
> > > > >>>>
> > > > >>>> What I miss in this thread is why would there be any exception here. We
> > > > >>>> do not grant exceptions from standard practices on "I want" reasons.
> > > > >>>
> > > > >>> I agree that we should not have duplicated drivers just for vendor
> > > > >>> lolz.  But when it comes to adopting common frameworks and integrating
> > > > >>> better into the ecosystem, this doesn't seem like something we should
> > > > >>> actively discourage.  I don't think this is a case of vendor lolz, but
> > > > >>
> > > > >> No one discourages it. Following standard Linux kernel practices and
> > > > >> requirements is not discouraging, do not twist the narrative here.
> > > > >> Again, it is standard upstream review telling that we do not duplicate
> > > > >> drivers. Ever, unless there is serious exception needed.
> > > > >
> > > > > I wasn't trying to twist the narrative, just trying to come up with a
> > > > > path forward that isn't "no" or "improve existing driver", since
> > > > > neither of those gets us towards a future using common frameworks.
> > > > >
> > > > >> I asked why there should be an exception granted? Is the reason for
> > > > >> exception following:
> > > > >> "We want to adopt common framework"
> > > > >> ?
> > > > >
> > > > > Possibly?  But I don't think we want two drivers to be any sort of
> > > > > long term solution.  (Ie. as long as venus/iris have co-exist.)
> > > > >
> > > > >>
> > > > >>> rather reacting to drm/accel emerging as the standard framework for
> > > > >>> this sort of driver.
> > > > >>>
> > > > >>> So how do we get from here to there?
> > > > >>
> > > > >> What is wrong with my proposal?
> > > > >
> > > > > Maybe I missed something, my understanding was your proposal was
> > > > > "Grow/replace/improve existing driver instead of coming with a
> > > > > duplicate"..  grow or improve doesn't move us toward common
> > > > > frameworks.  Maybe "replace" is a valid option.  If there is something
> > > > > I missed, then I apologize.
> > > > >
> > > > > Options I can think of are:
> > > > >
> > > > > 1. Hardware cutoff.. new hw gets new driver, existing hw gets existing
> > > > >    driver
> > > > > 2. Backwards compat chardev registered by new driver, providing existing
> > > > >    UABI.  I'm not 100% sure about the feasibility/drawbacks of this..
> > > > >    AFAIU the fastrpc folks where planning a backwards compat layer in
> > > > >    userspace, so maybe it is possible.
> > > > > 3. exception?
> > > > >
> > > > > I'd like to know what the feasibility of #2 is, since at a high level
> > > > > that sounds like the best option.  Possibly limit exposure of legacy
> > > > > UABI to existing hw so we don't get into a place of needing to extend
> > > > > the legacy UABI for new hw?
> > > > >
> > > > > But #1 sounds like a non-controversial place to start regardless.
> > > > > Possibly with #2 coming as followup and necessary step before eventual
> > > > > migration to new driver for existing hw?
> > > > >
> > > > > Even if we start with #2, how do we handle first-merge-window
> > > > > bugs/regressions without reverting addition of new driver and removal
> > > > > of old?  It seems like we'd need a window of a couple release cycles
> > > > > where both drivers exist?
> > > > >
> > > > > Maybe others have other/better options in mind?
> > > > To all, I'm seeking on the approach I should follow to go ahead here. I
> > > > can work on implementing #1(as per Rob's list) with hw specific
> > > > compatible for v4 if it's acceptable.
> > > >
> > >
> > > I don't see any reason for you to define a "hw specific compatible",
> > > because as you have shown in this series (and as Rob point out), there's
> > > no difference in the "hardware".
> > >
> > > The only reason for your "hw specific compatible" is to make a software
> > > selection in Linux - and that's not what DeviceTree is for.
> >
> > That's not exactly true. There are protocol differences. For example,
> > polling mode is supported only since a certain timeline in the history.
> > Likewise other interface features are not supported on all the FastRPC
> > devices. For the polling mode support we were already beaten by the lack
> > of SoC-specific compats.
> >
>
> I can see the benefit of capturing some of the generational features in
> a compatible, like the changes related to address width. But for pure
> software features that doesn't have an actual bearing in the hardware,
> I'd prefer if we relied on dynamic discovery.
>
> But none of that applies to the question of "can I use compatible to
> select if we should use the new or old Linux driver".
>

Nit, driver could still have an allow/deny-list of machine
compatibles.  Might be something to keep in mind if a phased
depreciation strategy was desirable..

BR,
-R
Re: [PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver
Posted by Ekansh Gupta 2 weeks, 5 days ago
On 09-09-2026 04:14, Bjorn Andersson wrote:
> On Wed, Aug 26, 2026 at 06:37:00PM +0530, Ekansh Gupta wrote:
>> On 20-08-2026 20:17, Rob Clark wrote:
>>> On Wed, Aug 19, 2026 at 11:15 PM Krzysztof Kozlowski <krzk@kernel.org> wrote:
>>>>
>>>> On 19/08/2026 17:48, Rob Clark wrote:
>>>>> On Wed, Aug 19, 2026 at 8:27 AM Krzysztof Kozlowski <krzk@kernel.org> wrote:
>>>>>> The rule of usptream development is that we do not accept duplicated
>>>>>> code, just because a vendor wants to write something new. This is
>>>>>> basically the concept applied all over the drivers tree, where we pushed
>>>>>> back against all sorts of duplications all over the vendors.
>>>>>>
>>>>>> What I miss in this thread is why would there be any exception here. We
>>>>>> do not grant exceptions from standard practices on "I want" reasons.
>>>>>
>>>>> I agree that we should not have duplicated drivers just for vendor
>>>>> lolz.  But when it comes to adopting common frameworks and integrating
>>>>> better into the ecosystem, this doesn't seem like something we should
>>>>> actively discourage.  I don't think this is a case of vendor lolz, but
>>>>
>>>> No one discourages it. Following standard Linux kernel practices and
>>>> requirements is not discouraging, do not twist the narrative here.
>>>> Again, it is standard upstream review telling that we do not duplicate
>>>> drivers. Ever, unless there is serious exception needed.
>>>
>>> I wasn't trying to twist the narrative, just trying to come up with a
>>> path forward that isn't "no" or "improve existing driver", since
>>> neither of those gets us towards a future using common frameworks.
>>>
>>>> I asked why there should be an exception granted? Is the reason for
>>>> exception following:
>>>> "We want to adopt common framework"
>>>> ?
>>>
>>> Possibly?  But I don't think we want two drivers to be any sort of
>>> long term solution.  (Ie. as long as venus/iris have co-exist.)
>>>
>>>>
>>>>> rather reacting to drm/accel emerging as the standard framework for
>>>>> this sort of driver.
>>>>>
>>>>> So how do we get from here to there?
>>>>
>>>> What is wrong with my proposal?
>>>
>>> Maybe I missed something, my understanding was your proposal was
>>> "Grow/replace/improve existing driver instead of coming with a
>>> duplicate"..  grow or improve doesn't move us toward common
>>> frameworks.  Maybe "replace" is a valid option.  If there is something
>>> I missed, then I apologize.
>>>
>>> Options I can think of are:
>>>
>>> 1. Hardware cutoff.. new hw gets new driver, existing hw gets existing
>>>    driver
>>> 2. Backwards compat chardev registered by new driver, providing existing
>>>    UABI.  I'm not 100% sure about the feasibility/drawbacks of this..
>>>    AFAIU the fastrpc folks where planning a backwards compat layer in
>>>    userspace, so maybe it is possible.
>>> 3. exception?
>>>
>>> I'd like to know what the feasibility of #2 is, since at a high level
>>> that sounds like the best option.  Possibly limit exposure of legacy
>>> UABI to existing hw so we don't get into a place of needing to extend
>>> the legacy UABI for new hw?
>>>
>>> But #1 sounds like a non-controversial place to start regardless.
>>> Possibly with #2 coming as followup and necessary step before eventual
>>> migration to new driver for existing hw?
>>>
>>> Even if we start with #2, how do we handle first-merge-window
>>> bugs/regressions without reverting addition of new driver and removal
>>> of old?  It seems like we'd need a window of a couple release cycles
>>> where both drivers exist?
>>>
>>> Maybe others have other/better options in mind?
>> To all, I'm seeking on the approach I should follow to go ahead here. I
>> can work on implementing #1(as per Rob's list) with hw specific
>> compatible for v4 if it's acceptable.
>>
> 
> I don't see any reason for you to define a "hw specific compatible",
> because as you have shown in this series (and as Rob point out), there's
> no difference in the "hardware".
> 
> The only reason for your "hw specific compatible" is to make a software
> selection in Linux - and that's not what DeviceTree is for.
> 
> 
> As such, I don't see that you have a DeviceTree problem at all, because
> this is a Linux-internal problem.

Agreed.

> 
>> #2(compat driver) is something that we are still exploring as we
>> couldn't find any standard way to achieve it. We might start a separate
>> discussion for that once we have few possible designs with us.
>>
> 
> This is the actual problem!
> 
> We have existing user space that depends on the ioctl interface exposed
> by the current misc driver. You must not break these.
> 
> Hardware cutoff is not a viable solution, because that's just a
> declaration that we'll let the old platforms rotten - or alternatively
> you commit to maintain two drivers to the very same feature and quality
> level.
> 
> So the only reasonable solution is #2; from there it's a valid question
> if you reach that point my stepwise migrating the current misc driver
> that solution, or if you present a new driver with the fully backwards
> compatible interface, alongside the new ABI.
> 
> 
> But this does bring to a question which the cover letter should explain
> - but doesn't: what problem does this patch series actually solve?
> 
> Regards,
> Bjorn

Agreed. I'll target #2: QDA implementing the existing fastrpc UABI
alongside the new one, rather than a driver split by platform.

On how to get there: the blocker we hit in v1 was that legacy fastrpc
buffer semantics appeared to need a drm_file, and there's no exported
way to construct one outside the DRM core. I want to re-examine that
constraint rather than treat it as final, since the legacy interface's
own buffer model (a dma_buf fd as the buffer identity, no GEM
involved) is not inherently tied to drm_file, that's how the existing
misc driver implements it today. I don't have a concrete design yet
and would rather work through it here than commit to one prematurely.
If anyone has thoughts on how the legacy UABI could be served without
requiring a drm_file per session or if I can somehow bind drm_file with
chardev by exposing some APIs from DRM core, I'd welcome them.

On the cover letter: fair point, and I'll fix it. The problem this
series solves is that a miscdevice interface requires us to hand-roll
what the accel/DRM subsystem already provides as common
infrastructure: GEM for buffer lifecycle and reference counting,
PRIME for cross-driver import/export, per-file (per-open) context and
handle-namespace isolation, and the existing debug and lifecycle
tooling the DRM core already ships. Every accelerator driver added to
drivers/accel (habanalabs, ivpu, qaic, rocket) has taken this path for
the same reason, rather than each maintaining its own equivalent
inside drivers/misc. Building QDA directly on this shared
infrastructure, instead of extending fastrpc's own ad hoc buffer and
session tracking to cover the same ground, avoids that duplication
going forward.

I'll put this plainly in the cover letter rather than leaving it
implicit.

/Ekansh

> 
>> Happy to take any other suggestion also.>
>>> BR,
>>> -R
>>
>>

Re: [PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver
Posted by Srinivas Kandagatla 2 weeks, 5 days ago
On 9/9/26 8:53 AM, Ekansh Gupta wrote:
>> This is the actual problem!
>>
>> We have existing user space that depends on the ioctl interface exposed
>> by the current misc driver. You must not break these.
>>
>> Hardware cutoff is not a viable solution, because that's just a
>> declaration that we'll let the old platforms rotten - or alternatively
>> you commit to maintain two drivers to the very same feature and quality
>> level.
>>
>> So the only reasonable solution is #2; from there it's a valid question
>> if you reach that point my stepwise migrating the current misc driver
>> that solution, or if you present a new driver with the fully backwards
>> compatible interface, alongside the new ABI.
>>
>>
>> But this does bring to a question which the cover letter should explain
>> - but doesn't: what problem does this patch series actually solve?
>>
>> Regards,
>> Bjorn
> Agreed. I'll target #2: QDA implementing the existing fastrpc UABI
> alongside the new one, rather than a driver split by platform.
> 
> On how to get there: the blocker we hit in v1 was that legacy fastrpc
> buffer semantics appeared to need a drm_file, and there's no exported
> way to construct one outside the DRM core. I want to re-examine that
> constraint rather than treat it as final, since the legacy interface's
> own buffer model (a dma_buf fd as the buffer identity, no GEM
> involved) is not inherently tied to drm_file, that's how the existing
> misc driver implements it today. I don't have a concrete design yet
> and would rather work through it here than commit to one prematurely.
> If anyone has thoughts on how the legacy UABI could be served without
> requiring a drm_file per session or if I can somehow bind drm_file with
> chardev by exposing some APIs from DRM core, I'd welcome them.
> 
> On the cover letter: fair point, and I'll fix it. The problem this
> series solves is that a miscdevice interface requires us to hand-roll
> what the accel/DRM subsystem already provides as common
> infrastructure: GEM for buffer lifecycle and reference counting,
> PRIME for cross-driver import/export, per-file (per-open) context and
> handle-namespace isolation, and the existing debug and lifecycle
> tooling the DRM core already ships. Every accelerator driver added to
> drivers/accel (habanalabs, ivpu, qaic, rocket) has taken this path for
> the same reason, rather than each maintaining its own equivalent
> inside drivers/misc. Building QDA directly on this shared
> infrastructure, instead of extending fastrpc's own ad hoc buffer and
> session tracking to cover the same ground, avoids that duplication
Am sure you must have already tried this, but Can you not migrate
existing fastrpc driver to this shared infrastructure under the hood?

Can you elaborate on what are the blockers you hit in doing so?

This will ensure that UAPI is retained and still get benefit of QDA.

--srini
> going forward.
Re: [PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver
Posted by Ekansh Gupta 2 weeks, 5 days ago
On 09-09-2026 13:32, Srinivas Kandagatla wrote:
> On 9/9/26 8:53 AM, Ekansh Gupta wrote:
>>> This is the actual problem!
>>>
>>> We have existing user space that depends on the ioctl interface exposed
>>> by the current misc driver. You must not break these.
>>>
>>> Hardware cutoff is not a viable solution, because that's just a
>>> declaration that we'll let the old platforms rotten - or alternatively
>>> you commit to maintain two drivers to the very same feature and quality
>>> level.
>>>
>>> So the only reasonable solution is #2; from there it's a valid question
>>> if you reach that point my stepwise migrating the current misc driver
>>> that solution, or if you present a new driver with the fully backwards
>>> compatible interface, alongside the new ABI.
>>>
>>>
>>> But this does bring to a question which the cover letter should explain
>>> - but doesn't: what problem does this patch series actually solve?
>>>
>>> Regards,
>>> Bjorn
>> Agreed. I'll target #2: QDA implementing the existing fastrpc UABI
>> alongside the new one, rather than a driver split by platform.
>>
>> On how to get there: the blocker we hit in v1 was that legacy fastrpc
>> buffer semantics appeared to need a drm_file, and there's no exported
>> way to construct one outside the DRM core. I want to re-examine that
>> constraint rather than treat it as final, since the legacy interface's
>> own buffer model (a dma_buf fd as the buffer identity, no GEM
>> involved) is not inherently tied to drm_file, that's how the existing
>> misc driver implements it today. I don't have a concrete design yet
>> and would rather work through it here than commit to one prematurely.
>> If anyone has thoughts on how the legacy UABI could be served without
>> requiring a drm_file per session or if I can somehow bind drm_file with
>> chardev by exposing some APIs from DRM core, I'd welcome them.
>>
>> On the cover letter: fair point, and I'll fix it. The problem this
>> series solves is that a miscdevice interface requires us to hand-roll
>> what the accel/DRM subsystem already provides as common
>> infrastructure: GEM for buffer lifecycle and reference counting,
>> PRIME for cross-driver import/export, per-file (per-open) context and
>> handle-namespace isolation, and the existing debug and lifecycle
>> tooling the DRM core already ships. Every accelerator driver added to
>> drivers/accel (habanalabs, ivpu, qaic, rocket) has taken this path for
>> the same reason, rather than each maintaining its own equivalent
>> inside drivers/misc. Building QDA directly on this shared
>> infrastructure, instead of extending fastrpc's own ad hoc buffer and
>> session tracking to cover the same ground, avoids that duplication
> Am sure you must have already tried this, but Can you not migrate
> existing fastrpc driver to this shared infrastructure under the hood?
> 
> Can you elaborate on what are the blockers you hit in doing so?
> 
> This will ensure that UAPI is retained and still get benefit of QDA.
> 

QDA's session model is built around drm_file, and fastrpc's chardev has
no drm_file associated with it. I'm still checking whether that
dependency is unavoidable for fastrpc's case, or whether there's a way
to use the underlying buffer-management infrastructure without it.

Please correct me if you are asking/suggesting a different approach?

> --srini
>> going forward.
>
Re: [PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver
Posted by Konrad Dybcio 1 month, 1 week ago
On 8/19/26 4:38 PM, Rob Clark wrote:
> On Wed, Aug 19, 2026 at 7:21 AM Krzysztof Kozlowski <krzk@kernel.org> wrote:
>>
>> On 19/08/2026 15:26, Ekansh Gupta wrote:
>>> On 19-08-2026 00:43, Krzysztof Kozlowski wrote:
>>>> On 17/08/2026 06:47, Ekansh Gupta wrote:
>>>>> This patch series introduces the Qualcomm DSP Accelerator (QDA) driver,
>>>>> a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a
>>>>> standardized interface for offloading computational tasks to DSPs found
>>>>> on Qualcomm SoCs, supporting all DSP domains.
>>>>>
>>>>> The QDA driver implements the FastRPC protocol over the DRM accel
>>>>> subsystem. It uses the same device-tree node structure as the existing
>>>>> fastrpc driver in drivers/misc/. The approach for binding the QDA driver
>>>>> to device-tree nodes while coexisting with the fastrpc driver is an open
>>>>> item described below.
>>>>
>>>> No. Grow/replace/improve existing driver instead of coming with a duplicate.
>>>>
>>>> That's a standard upstream requirement, basically given on every
>>>> upstreaming guide.
>>>>
>>>> Please watch old talk from Greg - "I Don’t Want Your Code!".
>>> Posted discussion threads here[1]. Would seek comments from Dmitry,
>>> Srini as well.
>>>
>>> [1]
>>> https://lore.kernel.org/all/3476b5c3-7983-4994-a901-3d7d8bd75255@oss.qualcomm.com/
>>
>> The rest of the comments is still valid even if you did not acknowledge
>> them.
>>
>> Anyway, regarding above - again, watch the talk from Greg.
>>
>> You have ONE driver. Not two.
> 
> Long term, moving to the common driver framework (which did not exist
> when fastrpc was first created) seems like a good thing.  But does
> that not allow for some transition period?  How can we get from here
> to there without otherwise breaking userspace?  Is there some other
> precedent elsewhere in other driver subsystems?
> 
> I suppose drm exposing legacy fbdev on top of drm drivers is _sort of_
> a precedent if you squint a bit?  I'm not really familiar enough to
> say if that would be reasonable/possible in this case.

Also, venus/iris, mdp5/dpu1

Konrad
Re: [PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver
Posted by Krzysztof Kozlowski 1 month, 1 week ago
On 18/08/2026 21:13, Krzysztof Kozlowski wrote:
> On 17/08/2026 06:47, Ekansh Gupta wrote:
>> This patch series introduces the Qualcomm DSP Accelerator (QDA) driver,
>> a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a
>> standardized interface for offloading computational tasks to DSPs found
>> on Qualcomm SoCs, supporting all DSP domains.
>>
>> The QDA driver implements the FastRPC protocol over the DRM accel
>> subsystem. It uses the same device-tree node structure as the existing
>> fastrpc driver in drivers/misc/. The approach for binding the QDA driver
>> to device-tree nodes while coexisting with the fastrpc driver is an open
>> item described below.
> 
> No. Grow/replace/improve existing driver instead of coming with a duplicate.
> 
> That's a standard upstream requirement, basically given on every
> upstreaming guide.
> 
> Please watch old talk from Greg - "I Don’t Want Your Code!".
> 
>>
>> v1: https://lore.kernel.org/all/20260519-qda-series-v1-0-b2d984c297f8@oss.qualcomm.com/
>> RFC: https://lore.kernel.org/dri-devel/20260224-qda-firstpost-v1-0-fe46a9c1a046@oss.qualcomm.com/T/
>>
>> Changes since v1
>> ================
>>
>> The v1 review raised two architectural objections and one correctness
>> issue; all three are resolved in v2:
>>
>> * Christian König (dma-buf maintainer) pointed out that the imported-
>>   buffer path silently assumed the IOMMU maps every buffer as a single
>>   contiguous range, which is not guaranteed. v2 walks the scatterlist
>>   and cleanly rejects non-contiguous imports; contiguous imports (e.g.
>>   CMA DMA-buf heap) are accepted. (patch 11)
>>
>> * Dmitry Baryshkov objected to three different buffer-passing formats
>>   in the invoke IOCTL (DMA-BUF fd, direct/inline, DMA handle). v2
>>   passes only GEM handles; userspace imports any fd to a GEM handle
>>   with DRM_IOCTL_PRIME_FD_TO_HANDLE before invoking. Packing and
>>   overlap handling are left to userspace. (patch 12)
>>
>> * The memory manager (patch 07) used a fixed 16-entry array without
>>   justification and leaked the device descriptor on teardown. v2
>>   allocates the array from the DT context-bank count (as Dmitry
>>   suggested) and frees it correctly.
>>
>> User-space staging branch
>> =========================
>> https://github.com/qualcomm/fastrpc/tree/accel/staging
>>
>> Key Features
>> ============
>>
>> * Standard DRM accelerator interface via /dev/accel/accelN
>> * GEM-based buffer management with DMA-BUF import (PRIME)
>> * IOMMU-based memory isolation using per-process context banks
>> * FastRPC protocol implementation for DSP communication
>> * RPMsg transport layer for reliable message passing
>> * Support for all DSP domains (ADSP, CDSP, SDSP, GDSP)
>> * DRM IOCTL interface for DSP session management, buffer allocation,
>>   and remote procedure invocation
>>
>> Architecture
>> ============
>>
>> 1. DRM Accelerator Framework Integration
>>    The driver registers as a DRM accel device, exposing a standard
>>    /dev/accel/accelN character device node. This provides established
>>    DRM infrastructure for device management, file operations, and
>>    IOCTL dispatch.
>>
>> 2. Memory Management
>>    Buffers are managed as GEM objects with PRIME support for DMA-BUF
>>    import. This enables buffer sharing with other DRM drivers (GPU,
>>    camera, video) using standard kernel mechanisms. Only contiguous
>>    imports are accepted; the driver verifies contiguity at import time
>>    rather than assuming it.
>>
>> 3. IOMMU Context Bank Management
>>    IOMMU context banks (CBs) are represented as proper struct device
>>    instances on a custom virtual bus (qda-compute-cb). Each CB device
>>    is registered with the IOMMU subsystem and receives its own IOMMU
>>    domain, enabling per-session address space isolation. The custom
>>    bus was introduced because IOMMU context banks are synthetic
>>    constructs — not real platform devices — and to ensure CB device
>>    lifetime is strictly subordinate to the parent QDA device.
>>    See also: https://lore.kernel.org/all/245d602f-3037-4ae3-9af9-d98f37258aae@oss.qualcomm.com/
>>
>> 4. Memory Manager Architecture
>>    The memory manager maintains a registry of IOMMU devices in an
>>    array sized to the number of context banks described in the device
>>    tree, and coordinates per-process device assignment with reference-
>>    counted lifetime management. The DMA-coherent backend allocates
>>    buffers with SID-prefixed DMA addresses for DSP firmware
>>    compatibility.
>>
>> 5. Transport Layer
>>    RPMsg communication is handled in a dedicated transport layer
>>    (qda_rpmsg.c), separate from the core DRM driver logic.
>>
>> 6. Code Organization
>>    The driver is organized across multiple files (~4800 lines total):
>>    * qda_drv.c:            Core driver and DRM integration
>>    * qda_rpmsg.c:          RPMsg transport layer
>>    * qda_cb.c:             Context bank device management
>>    * qda_compute_bus.c:    Custom virtual bus for CB devices
>>    * qda_gem.c:            GEM object management
>>    * qda_prime.c:          DMA-BUF import (PRIME)
>>    * qda_memory_manager.c: IOMMU device registry and allocation
>>    * qda_memory_dma.c:     DMA-coherent allocation backend
>>    * qda_fastrpc.c:        FastRPC protocol implementation
>>    * qda_ioctl.c:          IOCTL dispatch
>>
>> 7. UAPI Design
>>    The driver exposes DRM-style IOCTLs defined in
>>    include/uapi/drm/qda_accel.h, following DRM UAPI conventions
>>    (__u32/__u64 types, C++ guard, GPL-2.0-only WITH Linux-syscall-note).
>>    Buffer arguments are identified by GEM handles; the driver never
>>    accepts DMA-BUF fds directly in any IOCTL.
>>
>> Patch Series Organization
>> ==========================
>>
>> Patch 01:      MAINTAINERS entry
>> Patch 02:      Driver documentation (Documentation/accel/qda/)
>> Patches 03-04: Core driver skeleton and compute bus
>> Patch 05:      iommu: Register qda-compute-cb bus with IOMMU subsystem
>> Patches 06-07: CB device enumeration and memory manager
>> Patch 08:      QUERY IOCTL and UAPI header
>> Patches 09-11: GEM buffer management and PRIME import
>> Patches 12-15: FastRPC protocol (invoke, session create/release,
>>                map/unmap)
>>
>> Open Items
>> ===========
>>
>> 1. Device-Tree Compatible String
>>    The QDA driver uses the same device-tree node structure and
>>    properties as the existing fastrpc driver in drivers/misc/. A
>>    mechanism is needed to allow the QDA driver to bind to its device
>>    node independently of the fastrpc driver.
>>
>>    The intended coexistence model is: platforms that require the
>>    complete fastrpc feature set continue to use "qcom,fastrpc"; new
>>    platforms where QDA's feature set is sufficient use a QDA-specific
>>    compatible string. New feature development is directed toward QDA.
>>
>>    The options under consideration are:
>>
>>    a) Add a new "qcom,qda" compatible string to the existing
>>       qcom,fastrpc.yaml binding, since the DT node structure and
>>       properties are identical.
> No
> 
>>
>>    b) Introduce a separate qcom,qda.yaml binding that references or
>>       inherits the fastrpc binding properties.
> 
> No
> 
>>
>>    Seeking guidance from DT binding maintainers on the preferred
>>    approach.
> 
> Grow existing driver. You do not get new driver, you do not get new
> bindings.
> 

And this was already questioned at v1 (the true v1, not v1+1) but you
ignored the comment.

Great, so here goes away trust.

NAK

Best regards,
Krzysztof
Re: [PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver
Posted by Ekansh Gupta 1 month, 1 week ago
On 19-08-2026 00:51, Krzysztof Kozlowski wrote:
> On 18/08/2026 21:13, Krzysztof Kozlowski wrote:
>> On 17/08/2026 06:47, Ekansh Gupta wrote:
>>> This patch series introduces the Qualcomm DSP Accelerator (QDA) driver,
>>> a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a
>>> standardized interface for offloading computational tasks to DSPs found
>>> on Qualcomm SoCs, supporting all DSP domains.
>>>
>>> The QDA driver implements the FastRPC protocol over the DRM accel
>>> subsystem. It uses the same device-tree node structure as the existing
>>> fastrpc driver in drivers/misc/. The approach for binding the QDA driver
>>> to device-tree nodes while coexisting with the fastrpc driver is an open
>>> item described below.
>>
>> No. Grow/replace/improve existing driver instead of coming with a duplicate.
>>
>> That's a standard upstream requirement, basically given on every
>> upstreaming guide.
>>
>> Please watch old talk from Greg - "I Don’t Want Your Code!".
>>
>>>
>>> v1: https://lore.kernel.org/all/20260519-qda-series-v1-0-b2d984c297f8@oss.qualcomm.com/
>>> RFC: https://lore.kernel.org/dri-devel/20260224-qda-firstpost-v1-0-fe46a9c1a046@oss.qualcomm.com/T/
>>>
>>> Changes since v1
>>> ================
>>>
>>> The v1 review raised two architectural objections and one correctness
>>> issue; all three are resolved in v2:
>>>
>>> * Christian König (dma-buf maintainer) pointed out that the imported-
>>>   buffer path silently assumed the IOMMU maps every buffer as a single
>>>   contiguous range, which is not guaranteed. v2 walks the scatterlist
>>>   and cleanly rejects non-contiguous imports; contiguous imports (e.g.
>>>   CMA DMA-buf heap) are accepted. (patch 11)
>>>
>>> * Dmitry Baryshkov objected to three different buffer-passing formats
>>>   in the invoke IOCTL (DMA-BUF fd, direct/inline, DMA handle). v2
>>>   passes only GEM handles; userspace imports any fd to a GEM handle
>>>   with DRM_IOCTL_PRIME_FD_TO_HANDLE before invoking. Packing and
>>>   overlap handling are left to userspace. (patch 12)
>>>
>>> * The memory manager (patch 07) used a fixed 16-entry array without
>>>   justification and leaked the device descriptor on teardown. v2
>>>   allocates the array from the DT context-bank count (as Dmitry
>>>   suggested) and frees it correctly.
>>>
>>> User-space staging branch
>>> =========================
>>> https://github.com/qualcomm/fastrpc/tree/accel/staging
>>>
>>> Key Features
>>> ============
>>>
>>> * Standard DRM accelerator interface via /dev/accel/accelN
>>> * GEM-based buffer management with DMA-BUF import (PRIME)
>>> * IOMMU-based memory isolation using per-process context banks
>>> * FastRPC protocol implementation for DSP communication
>>> * RPMsg transport layer for reliable message passing
>>> * Support for all DSP domains (ADSP, CDSP, SDSP, GDSP)
>>> * DRM IOCTL interface for DSP session management, buffer allocation,
>>>   and remote procedure invocation
>>>
>>> Architecture
>>> ============
>>>
>>> 1. DRM Accelerator Framework Integration
>>>    The driver registers as a DRM accel device, exposing a standard
>>>    /dev/accel/accelN character device node. This provides established
>>>    DRM infrastructure for device management, file operations, and
>>>    IOCTL dispatch.
>>>
>>> 2. Memory Management
>>>    Buffers are managed as GEM objects with PRIME support for DMA-BUF
>>>    import. This enables buffer sharing with other DRM drivers (GPU,
>>>    camera, video) using standard kernel mechanisms. Only contiguous
>>>    imports are accepted; the driver verifies contiguity at import time
>>>    rather than assuming it.
>>>
>>> 3. IOMMU Context Bank Management
>>>    IOMMU context banks (CBs) are represented as proper struct device
>>>    instances on a custom virtual bus (qda-compute-cb). Each CB device
>>>    is registered with the IOMMU subsystem and receives its own IOMMU
>>>    domain, enabling per-session address space isolation. The custom
>>>    bus was introduced because IOMMU context banks are synthetic
>>>    constructs — not real platform devices — and to ensure CB device
>>>    lifetime is strictly subordinate to the parent QDA device.
>>>    See also: https://lore.kernel.org/all/245d602f-3037-4ae3-9af9-d98f37258aae@oss.qualcomm.com/
>>>
>>> 4. Memory Manager Architecture
>>>    The memory manager maintains a registry of IOMMU devices in an
>>>    array sized to the number of context banks described in the device
>>>    tree, and coordinates per-process device assignment with reference-
>>>    counted lifetime management. The DMA-coherent backend allocates
>>>    buffers with SID-prefixed DMA addresses for DSP firmware
>>>    compatibility.
>>>
>>> 5. Transport Layer
>>>    RPMsg communication is handled in a dedicated transport layer
>>>    (qda_rpmsg.c), separate from the core DRM driver logic.
>>>
>>> 6. Code Organization
>>>    The driver is organized across multiple files (~4800 lines total):
>>>    * qda_drv.c:            Core driver and DRM integration
>>>    * qda_rpmsg.c:          RPMsg transport layer
>>>    * qda_cb.c:             Context bank device management
>>>    * qda_compute_bus.c:    Custom virtual bus for CB devices
>>>    * qda_gem.c:            GEM object management
>>>    * qda_prime.c:          DMA-BUF import (PRIME)
>>>    * qda_memory_manager.c: IOMMU device registry and allocation
>>>    * qda_memory_dma.c:     DMA-coherent allocation backend
>>>    * qda_fastrpc.c:        FastRPC protocol implementation
>>>    * qda_ioctl.c:          IOCTL dispatch
>>>
>>> 7. UAPI Design
>>>    The driver exposes DRM-style IOCTLs defined in
>>>    include/uapi/drm/qda_accel.h, following DRM UAPI conventions
>>>    (__u32/__u64 types, C++ guard, GPL-2.0-only WITH Linux-syscall-note).
>>>    Buffer arguments are identified by GEM handles; the driver never
>>>    accepts DMA-BUF fds directly in any IOCTL.
>>>
>>> Patch Series Organization
>>> ==========================
>>>
>>> Patch 01:      MAINTAINERS entry
>>> Patch 02:      Driver documentation (Documentation/accel/qda/)
>>> Patches 03-04: Core driver skeleton and compute bus
>>> Patch 05:      iommu: Register qda-compute-cb bus with IOMMU subsystem
>>> Patches 06-07: CB device enumeration and memory manager
>>> Patch 08:      QUERY IOCTL and UAPI header
>>> Patches 09-11: GEM buffer management and PRIME import
>>> Patches 12-15: FastRPC protocol (invoke, session create/release,
>>>                map/unmap)
>>>
>>> Open Items
>>> ===========
>>>
>>> 1. Device-Tree Compatible String
>>>    The QDA driver uses the same device-tree node structure and
>>>    properties as the existing fastrpc driver in drivers/misc/. A
>>>    mechanism is needed to allow the QDA driver to bind to its device
>>>    node independently of the fastrpc driver.
>>>
>>>    The intended coexistence model is: platforms that require the
>>>    complete fastrpc feature set continue to use "qcom,fastrpc"; new
>>>    platforms where QDA's feature set is sufficient use a QDA-specific
>>>    compatible string. New feature development is directed toward QDA.
>>>
>>>    The options under consideration are:
>>>
>>>    a) Add a new "qcom,qda" compatible string to the existing
>>>       qcom,fastrpc.yaml binding, since the DT node structure and
>>>       properties are identical.
>> No
>>
>>>
>>>    b) Introduce a separate qcom,qda.yaml binding that references or
>>>       inherits the fastrpc binding properties.
>>
>> No
>>
>>>
>>>    Seeking guidance from DT binding maintainers on the preferred
>>>    approach.
>>
>> Grow existing driver. You do not get new driver, you do not get new
>> bindings.
>>
> 
> And this was already questioned at v1 (the true v1, not v1+1) but you
> ignored the comment.
The discussion were around compat layers in v1 patch (which is not yet
concluded) and on whether this driver is going to be an alternative or a
replacement. Please excuse be of overlooking any NAK on the true v1
series, but I can't still find it.

Thanks for your review, Krzysztof.

//Ekansh>
> Great, so here goes away trust.
> 
> NAK
> 
> Best regards,
> Krzysztof

Re: [PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver
Posted by Krzysztof Kozlowski 1 month, 1 week ago
On 19/08/2026 15:32, Ekansh Gupta wrote:
> On 19-08-2026 00:51, Krzysztof Kozlowski wrote:
>> On 18/08/2026 21:13, Krzysztof Kozlowski wrote:
>>> On 17/08/2026 06:47, Ekansh Gupta wrote:
>>>> This patch series introduces the Qualcomm DSP Accelerator (QDA) driver,
>>>> a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a
>>>> standardized interface for offloading computational tasks to DSPs found
>>>> on Qualcomm SoCs, supporting all DSP domains.
>>>>
>>>> The QDA driver implements the FastRPC protocol over the DRM accel
>>>> subsystem. It uses the same device-tree node structure as the existing
>>>> fastrpc driver in drivers/misc/. The approach for binding the QDA driver
>>>> to device-tree nodes while coexisting with the fastrpc driver is an open
>>>> item described below.
>>>
>>> No. Grow/replace/improve existing driver instead of coming with a duplicate.
>>>
>>> That's a standard upstream requirement, basically given on every
>>> upstreaming guide.
>>>
>>> Please watch old talk from Greg - "I Don’t Want Your Code!".
>>>
>>>>
>>>> v1: https://lore.kernel.org/all/20260519-qda-series-v1-0-b2d984c297f8@oss.qualcomm.com/
>>>> RFC: https://lore.kernel.org/dri-devel/20260224-qda-firstpost-v1-0-fe46a9c1a046@oss.qualcomm.com/T/
>>>>
>>>> Changes since v1
>>>> ================
>>>>
>>>> The v1 review raised two architectural objections and one correctness
>>>> issue; all three are resolved in v2:
>>>>
>>>> * Christian König (dma-buf maintainer) pointed out that the imported-
>>>>   buffer path silently assumed the IOMMU maps every buffer as a single
>>>>   contiguous range, which is not guaranteed. v2 walks the scatterlist
>>>>   and cleanly rejects non-contiguous imports; contiguous imports (e.g.
>>>>   CMA DMA-buf heap) are accepted. (patch 11)
>>>>
>>>> * Dmitry Baryshkov objected to three different buffer-passing formats
>>>>   in the invoke IOCTL (DMA-BUF fd, direct/inline, DMA handle). v2
>>>>   passes only GEM handles; userspace imports any fd to a GEM handle
>>>>   with DRM_IOCTL_PRIME_FD_TO_HANDLE before invoking. Packing and
>>>>   overlap handling are left to userspace. (patch 12)
>>>>
>>>> * The memory manager (patch 07) used a fixed 16-entry array without
>>>>   justification and leaked the device descriptor on teardown. v2
>>>>   allocates the array from the DT context-bank count (as Dmitry
>>>>   suggested) and frees it correctly.
>>>>
>>>> User-space staging branch
>>>> =========================
>>>> https://github.com/qualcomm/fastrpc/tree/accel/staging
>>>>
>>>> Key Features
>>>> ============
>>>>
>>>> * Standard DRM accelerator interface via /dev/accel/accelN
>>>> * GEM-based buffer management with DMA-BUF import (PRIME)
>>>> * IOMMU-based memory isolation using per-process context banks
>>>> * FastRPC protocol implementation for DSP communication
>>>> * RPMsg transport layer for reliable message passing
>>>> * Support for all DSP domains (ADSP, CDSP, SDSP, GDSP)
>>>> * DRM IOCTL interface for DSP session management, buffer allocation,
>>>>   and remote procedure invocation
>>>>
>>>> Architecture
>>>> ============
>>>>
>>>> 1. DRM Accelerator Framework Integration
>>>>    The driver registers as a DRM accel device, exposing a standard
>>>>    /dev/accel/accelN character device node. This provides established
>>>>    DRM infrastructure for device management, file operations, and
>>>>    IOCTL dispatch.
>>>>
>>>> 2. Memory Management
>>>>    Buffers are managed as GEM objects with PRIME support for DMA-BUF
>>>>    import. This enables buffer sharing with other DRM drivers (GPU,
>>>>    camera, video) using standard kernel mechanisms. Only contiguous
>>>>    imports are accepted; the driver verifies contiguity at import time
>>>>    rather than assuming it.
>>>>
>>>> 3. IOMMU Context Bank Management
>>>>    IOMMU context banks (CBs) are represented as proper struct device
>>>>    instances on a custom virtual bus (qda-compute-cb). Each CB device
>>>>    is registered with the IOMMU subsystem and receives its own IOMMU
>>>>    domain, enabling per-session address space isolation. The custom
>>>>    bus was introduced because IOMMU context banks are synthetic
>>>>    constructs — not real platform devices — and to ensure CB device
>>>>    lifetime is strictly subordinate to the parent QDA device.
>>>>    See also: https://lore.kernel.org/all/245d602f-3037-4ae3-9af9-d98f37258aae@oss.qualcomm.com/
>>>>
>>>> 4. Memory Manager Architecture
>>>>    The memory manager maintains a registry of IOMMU devices in an
>>>>    array sized to the number of context banks described in the device
>>>>    tree, and coordinates per-process device assignment with reference-
>>>>    counted lifetime management. The DMA-coherent backend allocates
>>>>    buffers with SID-prefixed DMA addresses for DSP firmware
>>>>    compatibility.
>>>>
>>>> 5. Transport Layer
>>>>    RPMsg communication is handled in a dedicated transport layer
>>>>    (qda_rpmsg.c), separate from the core DRM driver logic.
>>>>
>>>> 6. Code Organization
>>>>    The driver is organized across multiple files (~4800 lines total):
>>>>    * qda_drv.c:            Core driver and DRM integration
>>>>    * qda_rpmsg.c:          RPMsg transport layer
>>>>    * qda_cb.c:             Context bank device management
>>>>    * qda_compute_bus.c:    Custom virtual bus for CB devices
>>>>    * qda_gem.c:            GEM object management
>>>>    * qda_prime.c:          DMA-BUF import (PRIME)
>>>>    * qda_memory_manager.c: IOMMU device registry and allocation
>>>>    * qda_memory_dma.c:     DMA-coherent allocation backend
>>>>    * qda_fastrpc.c:        FastRPC protocol implementation
>>>>    * qda_ioctl.c:          IOCTL dispatch
>>>>
>>>> 7. UAPI Design
>>>>    The driver exposes DRM-style IOCTLs defined in
>>>>    include/uapi/drm/qda_accel.h, following DRM UAPI conventions
>>>>    (__u32/__u64 types, C++ guard, GPL-2.0-only WITH Linux-syscall-note).
>>>>    Buffer arguments are identified by GEM handles; the driver never
>>>>    accepts DMA-BUF fds directly in any IOCTL.
>>>>
>>>> Patch Series Organization
>>>> ==========================
>>>>
>>>> Patch 01:      MAINTAINERS entry
>>>> Patch 02:      Driver documentation (Documentation/accel/qda/)
>>>> Patches 03-04: Core driver skeleton and compute bus
>>>> Patch 05:      iommu: Register qda-compute-cb bus with IOMMU subsystem
>>>> Patches 06-07: CB device enumeration and memory manager
>>>> Patch 08:      QUERY IOCTL and UAPI header
>>>> Patches 09-11: GEM buffer management and PRIME import
>>>> Patches 12-15: FastRPC protocol (invoke, session create/release,
>>>>                map/unmap)
>>>>
>>>> Open Items
>>>> ===========
>>>>
>>>> 1. Device-Tree Compatible String
>>>>    The QDA driver uses the same device-tree node structure and
>>>>    properties as the existing fastrpc driver in drivers/misc/. A
>>>>    mechanism is needed to allow the QDA driver to bind to its device
>>>>    node independently of the fastrpc driver.
>>>>
>>>>    The intended coexistence model is: platforms that require the
>>>>    complete fastrpc feature set continue to use "qcom,fastrpc"; new
>>>>    platforms where QDA's feature set is sufficient use a QDA-specific
>>>>    compatible string. New feature development is directed toward QDA.
>>>>
>>>>    The options under consideration are:
>>>>
>>>>    a) Add a new "qcom,qda" compatible string to the existing
>>>>       qcom,fastrpc.yaml binding, since the DT node structure and
>>>>       properties are identical.
>>> No
>>>
>>>>
>>>>    b) Introduce a separate qcom,qda.yaml binding that references or
>>>>       inherits the fastrpc binding properties.
>>>
>>> No
>>>
>>>>
>>>>    Seeking guidance from DT binding maintainers on the preferred
>>>>    approach.
>>>
>>> Grow existing driver. You do not get new driver, you do not get new
>>> bindings.
>>>
>>
>> And this was already questioned at v1 (the true v1, not v1+1) but you
>> ignored the comment.
> The discussion were around compat layers in v1 patch (which is not yet
> concluded) and on whether this driver is going to be an alternative or a

No, you got comment, from Trilok I think, asking what is the plan in
respect of existing fastrpc driver.


Best regards,
Krzysztof
Re: [PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver
Posted by Krzysztof Kozlowski 1 month, 1 week ago
On 17/08/2026 06:47, Ekansh Gupta wrote:
> This patch series introduces the Qualcomm DSP Accelerator (QDA) driver,
> a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a
> standardized interface for offloading computational tasks to DSPs found
> on Qualcomm SoCs, supporting all DSP domains.
> 
> The QDA driver implements the FastRPC protocol over the DRM accel
> subsystem. It uses the same device-tree node structure as the existing
> fastrpc driver in drivers/misc/. The approach for binding the QDA driver
> to device-tree nodes while coexisting with the fastrpc driver is an open
> item described below.
> 
> v1: https://lore.kernel.org/all/20260519-qda-series-v1-0-b2d984c297f8@oss.qualcomm.com/
> RFC: https://lore.kernel.org/dri-devel/20260224-qda-firstpost-v1-0-fe46a9c1a046@oss.qualcomm.com/T/

So this is a v3, not v2. Please start using b4 correctly, so versioning
will be kept instead of faking the numbers.

Best regards,
Krzysztof
Re: [PATCH v2 00/15] accel/qda: Qualcomm DSP Accelerator driver
Posted by Ekansh Gupta 1 month, 1 week ago
On 19-08-2026 00:48, Krzysztof Kozlowski wrote:
> On 17/08/2026 06:47, Ekansh Gupta wrote:
>> This patch series introduces the Qualcomm DSP Accelerator (QDA) driver,
>> a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a
>> standardized interface for offloading computational tasks to DSPs found
>> on Qualcomm SoCs, supporting all DSP domains.
>>
>> The QDA driver implements the FastRPC protocol over the DRM accel
>> subsystem. It uses the same device-tree node structure as the existing
>> fastrpc driver in drivers/misc/. The approach for binding the QDA driver
>> to device-tree nodes while coexisting with the fastrpc driver is an open
>> item described below.
>>
>> v1: https://lore.kernel.org/all/20260519-qda-series-v1-0-b2d984c297f8@oss.qualcomm.com/
>> RFC: https://lore.kernel.org/dri-devel/20260224-qda-firstpost-v1-0-fe46a9c1a046@oss.qualcomm.com/T/
> 
> So this is a v3, not v2. Please start using b4 correctly, so versioning
> will be kept instead of faking the numbers.
I'll correct the versioning in the next post. Thank you for pointing this.
> 
> Best regards,
> Krzysztof