Documentation/accel/index.rst | 1 + Documentation/accel/qda/index.rst | 13 + Documentation/accel/qda/qda.rst | 191 ++++++ MAINTAINERS | 11 + drivers/accel/Kconfig | 1 + drivers/accel/Makefile | 2 + drivers/accel/qda/Kconfig | 34 ++ drivers/accel/qda/Makefile | 19 + drivers/accel/qda/qda_cb.c | 125 ++++ drivers/accel/qda/qda_cb.h | 32 + drivers/accel/qda/qda_compute_bus.c | 80 +++ drivers/accel/qda/qda_drv.c | 144 +++++ drivers/accel/qda/qda_drv.h | 90 +++ drivers/accel/qda/qda_fastrpc.c | 1009 ++++++++++++++++++++++++++++++++ drivers/accel/qda/qda_fastrpc.h | 367 ++++++++++++ drivers/accel/qda/qda_gem.c | 155 +++++ drivers/accel/qda/qda_gem.h | 60 ++ drivers/accel/qda/qda_ioctl.c | 290 +++++++++ drivers/accel/qda/qda_ioctl.h | 19 + drivers/accel/qda/qda_memory_dma.c | 82 +++ drivers/accel/qda/qda_memory_dma.h | 17 + drivers/accel/qda/qda_memory_manager.c | 369 ++++++++++++ drivers/accel/qda/qda_memory_manager.h | 85 +++ drivers/accel/qda/qda_prime.c | 167 ++++++ drivers/accel/qda/qda_prime.h | 18 + drivers/accel/qda/qda_rpmsg.c | 201 +++++++ drivers/accel/qda/qda_rpmsg.h | 26 + drivers/iommu/iommu.c | 4 + include/linux/qda_compute_bus.h | 33 ++ include/uapi/drm/qda_accel.h | 242 ++++++++ 30 files changed, 3887 insertions(+)
This patch series introduces the Qualcomm DSP Accelerator (QDA) driver,
a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a
standardized interface for offloading computational tasks to DSPs found
on Qualcomm SoCs, supporting all DSP domains.
The QDA driver implements the FastRPC protocol over the DRM accel
subsystem. It uses the same device-tree node structure as the existing
fastrpc driver in drivers/misc/. The approach for binding the QDA driver
to device-tree nodes while coexisting with the fastrpc driver is an open
item described below.
v1: https://lore.kernel.org/all/20260519-qda-series-v1-0-b2d984c297f8@oss.qualcomm.com/
RFC: https://lore.kernel.org/dri-devel/20260224-qda-firstpost-v1-0-fe46a9c1a046@oss.qualcomm.com/T/
Changes since v1
================
The v1 review raised two architectural objections and one correctness
issue; all three are resolved in v2:
* Christian König (dma-buf maintainer) pointed out that the imported-
buffer path silently assumed the IOMMU maps every buffer as a single
contiguous range, which is not guaranteed. v2 walks the scatterlist
and cleanly rejects non-contiguous imports; contiguous imports (e.g.
CMA DMA-buf heap) are accepted. (patch 11)
* Dmitry Baryshkov objected to three different buffer-passing formats
in the invoke IOCTL (DMA-BUF fd, direct/inline, DMA handle). v2
passes only GEM handles; userspace imports any fd to a GEM handle
with DRM_IOCTL_PRIME_FD_TO_HANDLE before invoking. Packing and
overlap handling are left to userspace. (patch 12)
* The memory manager (patch 07) used a fixed 16-entry array without
justification and leaked the device descriptor on teardown. v2
allocates the array from the DT context-bank count (as Dmitry
suggested) and frees it correctly.
User-space staging branch
=========================
https://github.com/qualcomm/fastrpc/tree/accel/staging
Key Features
============
* Standard DRM accelerator interface via /dev/accel/accelN
* GEM-based buffer management with DMA-BUF import (PRIME)
* IOMMU-based memory isolation using per-process context banks
* FastRPC protocol implementation for DSP communication
* RPMsg transport layer for reliable message passing
* Support for all DSP domains (ADSP, CDSP, SDSP, GDSP)
* DRM IOCTL interface for DSP session management, buffer allocation,
and remote procedure invocation
Architecture
============
1. DRM Accelerator Framework Integration
The driver registers as a DRM accel device, exposing a standard
/dev/accel/accelN character device node. This provides established
DRM infrastructure for device management, file operations, and
IOCTL dispatch.
2. Memory Management
Buffers are managed as GEM objects with PRIME support for DMA-BUF
import. This enables buffer sharing with other DRM drivers (GPU,
camera, video) using standard kernel mechanisms. Only contiguous
imports are accepted; the driver verifies contiguity at import time
rather than assuming it.
3. IOMMU Context Bank Management
IOMMU context banks (CBs) are represented as proper struct device
instances on a custom virtual bus (qda-compute-cb). Each CB device
is registered with the IOMMU subsystem and receives its own IOMMU
domain, enabling per-session address space isolation. The custom
bus was introduced because IOMMU context banks are synthetic
constructs — not real platform devices — and to ensure CB device
lifetime is strictly subordinate to the parent QDA device.
See also: https://lore.kernel.org/all/245d602f-3037-4ae3-9af9-d98f37258aae@oss.qualcomm.com/
4. Memory Manager Architecture
The memory manager maintains a registry of IOMMU devices in an
array sized to the number of context banks described in the device
tree, and coordinates per-process device assignment with reference-
counted lifetime management. The DMA-coherent backend allocates
buffers with SID-prefixed DMA addresses for DSP firmware
compatibility.
5. Transport Layer
RPMsg communication is handled in a dedicated transport layer
(qda_rpmsg.c), separate from the core DRM driver logic.
6. Code Organization
The driver is organized across multiple files (~4800 lines total):
* qda_drv.c: Core driver and DRM integration
* qda_rpmsg.c: RPMsg transport layer
* qda_cb.c: Context bank device management
* qda_compute_bus.c: Custom virtual bus for CB devices
* qda_gem.c: GEM object management
* qda_prime.c: DMA-BUF import (PRIME)
* qda_memory_manager.c: IOMMU device registry and allocation
* qda_memory_dma.c: DMA-coherent allocation backend
* qda_fastrpc.c: FastRPC protocol implementation
* qda_ioctl.c: IOCTL dispatch
7. UAPI Design
The driver exposes DRM-style IOCTLs defined in
include/uapi/drm/qda_accel.h, following DRM UAPI conventions
(__u32/__u64 types, C++ guard, GPL-2.0-only WITH Linux-syscall-note).
Buffer arguments are identified by GEM handles; the driver never
accepts DMA-BUF fds directly in any IOCTL.
Patch Series Organization
==========================
Patch 01: MAINTAINERS entry
Patch 02: Driver documentation (Documentation/accel/qda/)
Patches 03-04: Core driver skeleton and compute bus
Patch 05: iommu: Register qda-compute-cb bus with IOMMU subsystem
Patches 06-07: CB device enumeration and memory manager
Patch 08: QUERY IOCTL and UAPI header
Patches 09-11: GEM buffer management and PRIME import
Patches 12-15: FastRPC protocol (invoke, session create/release,
map/unmap)
Open Items
===========
1. Device-Tree Compatible String
The QDA driver uses the same device-tree node structure and
properties as the existing fastrpc driver in drivers/misc/. A
mechanism is needed to allow the QDA driver to bind to its device
node independently of the fastrpc driver.
The intended coexistence model is: platforms that require the
complete fastrpc feature set continue to use "qcom,fastrpc"; new
platforms where QDA's feature set is sufficient use a QDA-specific
compatible string. New feature development is directed toward QDA.
The options under consideration are:
a) Add a new "qcom,qda" compatible string to the existing
qcom,fastrpc.yaml binding, since the DT node structure and
properties are identical.
b) Introduce a separate qcom,qda.yaml binding that references or
inherits the fastrpc binding properties.
Seeking guidance from DT binding maintainers on the preferred
approach.
2. Privilege Level Management
Currently, daemon processes and user processes have the same access
level as both use the same accel device node. Daemons attach to
privileged DSP protection domains and require higher privilege
levels for system-level operations. Seeking guidance on the best
approach: separate device nodes, capability-based checks, or DRM
master/authentication mechanisms.
3. Audio and Sensors PD Support
The current series does not handle Audio PD and Sensors PD
functionalities. These specialized protection domains require
additional support for real-time constraints and power management.
Interface Compatibility
========================
The QDA driver uses the same device-tree node structure and child node
layout (including "qcom,fastrpc-compute-cb" child nodes) as the
existing fastrpc driver. The underlying FastRPC protocol and DSP
firmware interface are compatible with the existing fastrpc driver,
ensuring that DSP firmware and libraries continue to work without
modification.
References
==========
Previous discussions on this migration:
- https://lkml.org/lkml/2024/6/24/479
- https://lkml.org/lkml/2024/6/21/1252
Testing
=======
The driver has been tested on Qualcomm platforms with:
- Basic FastRPC attach/release operations
- DSP process creation and initialization
- Memory mapping/unmapping operations
- Dynamic invocation with various buffer types
- GEM buffer allocation and mmap
- PRIME buffer import from other subsystems (contiguous buffers)
Signed-off-by: Ekansh Gupta <ekansh.gupta@oss.qualcomm.com>
---
Ekansh Gupta (15):
MAINTAINERS: Add entry for Qualcomm DSP Accelerator (QDA) driver
accel/qda: Add QDA driver documentation
accel/qda: Add initial QDA DRM accelerator driver
accel/qda: Add compute bus for QDA context banks
iommu: Add QDA compute context bank bus to iommu_buses
accel/qda: Create compute context bank devices on QDA compute bus
accel/qda: Add memory manager for CB devices
accel/qda: Add QUERY IOCTL and QDA UAPI header
accel/qda: Add DMA-backed GEM objects and memory manager integration
accel/qda: Add GEM_CREATE and GEM_MMAP_OFFSET IOCTLs
accel/qda: Add PRIME DMA-BUF import support
accel/qda: Add FastRPC invocation support
accel/qda: Add DSP process creation and release
accel/qda: Add remote memory mapping to DSP address space
accel/qda: Add remote memory unmap from DSP address space
Documentation/accel/index.rst | 1 +
Documentation/accel/qda/index.rst | 13 +
Documentation/accel/qda/qda.rst | 191 ++++++
MAINTAINERS | 11 +
drivers/accel/Kconfig | 1 +
drivers/accel/Makefile | 2 +
drivers/accel/qda/Kconfig | 34 ++
drivers/accel/qda/Makefile | 19 +
drivers/accel/qda/qda_cb.c | 125 ++++
drivers/accel/qda/qda_cb.h | 32 +
drivers/accel/qda/qda_compute_bus.c | 80 +++
drivers/accel/qda/qda_drv.c | 144 +++++
drivers/accel/qda/qda_drv.h | 90 +++
drivers/accel/qda/qda_fastrpc.c | 1009 ++++++++++++++++++++++++++++++++
drivers/accel/qda/qda_fastrpc.h | 367 ++++++++++++
drivers/accel/qda/qda_gem.c | 155 +++++
drivers/accel/qda/qda_gem.h | 60 ++
drivers/accel/qda/qda_ioctl.c | 290 +++++++++
drivers/accel/qda/qda_ioctl.h | 19 +
drivers/accel/qda/qda_memory_dma.c | 82 +++
drivers/accel/qda/qda_memory_dma.h | 17 +
drivers/accel/qda/qda_memory_manager.c | 369 ++++++++++++
drivers/accel/qda/qda_memory_manager.h | 85 +++
drivers/accel/qda/qda_prime.c | 167 ++++++
drivers/accel/qda/qda_prime.h | 18 +
drivers/accel/qda/qda_rpmsg.c | 201 +++++++
drivers/accel/qda/qda_rpmsg.h | 26 +
drivers/iommu/iommu.c | 4 +
include/linux/qda_compute_bus.h | 33 ++
include/uapi/drm/qda_accel.h | 242 ++++++++
30 files changed, 3887 insertions(+)
---
base-commit: 5f07a0db7088b4ef4b9a48069a93b9f3e1a33379
change-id: 20260817-qda-v2-78e2d1f10529
Best regards,
--
Ekansh Gupta <ekansh.gupta@oss.qualcomm.com>
On 17/08/2026 06:47, Ekansh Gupta wrote: > This patch series introduces the Qualcomm DSP Accelerator (QDA) driver, > a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a > standardized interface for offloading computational tasks to DSPs found > on Qualcomm SoCs, supporting all DSP domains. > > The QDA driver implements the FastRPC protocol over the DRM accel > subsystem. It uses the same device-tree node structure as the existing > fastrpc driver in drivers/misc/. The approach for binding the QDA driver > to device-tree nodes while coexisting with the fastrpc driver is an open > item described below. No. Grow/replace/improve existing driver instead of coming with a duplicate. That's a standard upstream requirement, basically given on every upstreaming guide. Please watch old talk from Greg - "I Don’t Want Your Code!". > > v1: https://lore.kernel.org/all/20260519-qda-series-v1-0-b2d984c297f8@oss.qualcomm.com/ > RFC: https://lore.kernel.org/dri-devel/20260224-qda-firstpost-v1-0-fe46a9c1a046@oss.qualcomm.com/T/ > > Changes since v1 > ================ > > The v1 review raised two architectural objections and one correctness > issue; all three are resolved in v2: > > * Christian König (dma-buf maintainer) pointed out that the imported- > buffer path silently assumed the IOMMU maps every buffer as a single > contiguous range, which is not guaranteed. v2 walks the scatterlist > and cleanly rejects non-contiguous imports; contiguous imports (e.g. > CMA DMA-buf heap) are accepted. (patch 11) > > * Dmitry Baryshkov objected to three different buffer-passing formats > in the invoke IOCTL (DMA-BUF fd, direct/inline, DMA handle). v2 > passes only GEM handles; userspace imports any fd to a GEM handle > with DRM_IOCTL_PRIME_FD_TO_HANDLE before invoking. Packing and > overlap handling are left to userspace. (patch 12) > > * The memory manager (patch 07) used a fixed 16-entry array without > justification and leaked the device descriptor on teardown. v2 > allocates the array from the DT context-bank count (as Dmitry > suggested) and frees it correctly. > > User-space staging branch > ========================= > https://github.com/qualcomm/fastrpc/tree/accel/staging > > Key Features > ============ > > * Standard DRM accelerator interface via /dev/accel/accelN > * GEM-based buffer management with DMA-BUF import (PRIME) > * IOMMU-based memory isolation using per-process context banks > * FastRPC protocol implementation for DSP communication > * RPMsg transport layer for reliable message passing > * Support for all DSP domains (ADSP, CDSP, SDSP, GDSP) > * DRM IOCTL interface for DSP session management, buffer allocation, > and remote procedure invocation > > Architecture > ============ > > 1. DRM Accelerator Framework Integration > The driver registers as a DRM accel device, exposing a standard > /dev/accel/accelN character device node. This provides established > DRM infrastructure for device management, file operations, and > IOCTL dispatch. > > 2. Memory Management > Buffers are managed as GEM objects with PRIME support for DMA-BUF > import. This enables buffer sharing with other DRM drivers (GPU, > camera, video) using standard kernel mechanisms. Only contiguous > imports are accepted; the driver verifies contiguity at import time > rather than assuming it. > > 3. IOMMU Context Bank Management > IOMMU context banks (CBs) are represented as proper struct device > instances on a custom virtual bus (qda-compute-cb). Each CB device > is registered with the IOMMU subsystem and receives its own IOMMU > domain, enabling per-session address space isolation. The custom > bus was introduced because IOMMU context banks are synthetic > constructs — not real platform devices — and to ensure CB device > lifetime is strictly subordinate to the parent QDA device. > See also: https://lore.kernel.org/all/245d602f-3037-4ae3-9af9-d98f37258aae@oss.qualcomm.com/ > > 4. Memory Manager Architecture > The memory manager maintains a registry of IOMMU devices in an > array sized to the number of context banks described in the device > tree, and coordinates per-process device assignment with reference- > counted lifetime management. The DMA-coherent backend allocates > buffers with SID-prefixed DMA addresses for DSP firmware > compatibility. > > 5. Transport Layer > RPMsg communication is handled in a dedicated transport layer > (qda_rpmsg.c), separate from the core DRM driver logic. > > 6. Code Organization > The driver is organized across multiple files (~4800 lines total): > * qda_drv.c: Core driver and DRM integration > * qda_rpmsg.c: RPMsg transport layer > * qda_cb.c: Context bank device management > * qda_compute_bus.c: Custom virtual bus for CB devices > * qda_gem.c: GEM object management > * qda_prime.c: DMA-BUF import (PRIME) > * qda_memory_manager.c: IOMMU device registry and allocation > * qda_memory_dma.c: DMA-coherent allocation backend > * qda_fastrpc.c: FastRPC protocol implementation > * qda_ioctl.c: IOCTL dispatch > > 7. UAPI Design > The driver exposes DRM-style IOCTLs defined in > include/uapi/drm/qda_accel.h, following DRM UAPI conventions > (__u32/__u64 types, C++ guard, GPL-2.0-only WITH Linux-syscall-note). > Buffer arguments are identified by GEM handles; the driver never > accepts DMA-BUF fds directly in any IOCTL. > > Patch Series Organization > ========================== > > Patch 01: MAINTAINERS entry > Patch 02: Driver documentation (Documentation/accel/qda/) > Patches 03-04: Core driver skeleton and compute bus > Patch 05: iommu: Register qda-compute-cb bus with IOMMU subsystem > Patches 06-07: CB device enumeration and memory manager > Patch 08: QUERY IOCTL and UAPI header > Patches 09-11: GEM buffer management and PRIME import > Patches 12-15: FastRPC protocol (invoke, session create/release, > map/unmap) > > Open Items > =========== > > 1. Device-Tree Compatible String > The QDA driver uses the same device-tree node structure and > properties as the existing fastrpc driver in drivers/misc/. A > mechanism is needed to allow the QDA driver to bind to its device > node independently of the fastrpc driver. > > The intended coexistence model is: platforms that require the > complete fastrpc feature set continue to use "qcom,fastrpc"; new > platforms where QDA's feature set is sufficient use a QDA-specific > compatible string. New feature development is directed toward QDA. > > The options under consideration are: > > a) Add a new "qcom,qda" compatible string to the existing > qcom,fastrpc.yaml binding, since the DT node structure and > properties are identical. No > > b) Introduce a separate qcom,qda.yaml binding that references or > inherits the fastrpc binding properties. No > > Seeking guidance from DT binding maintainers on the preferred > approach. Grow existing driver. You do not get new driver, you do not get new bindings. Best regards, Krzysztof
On 19-08-2026 00:43, Krzysztof Kozlowski wrote: > On 17/08/2026 06:47, Ekansh Gupta wrote: >> This patch series introduces the Qualcomm DSP Accelerator (QDA) driver, >> a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a >> standardized interface for offloading computational tasks to DSPs found >> on Qualcomm SoCs, supporting all DSP domains. >> >> The QDA driver implements the FastRPC protocol over the DRM accel >> subsystem. It uses the same device-tree node structure as the existing >> fastrpc driver in drivers/misc/. The approach for binding the QDA driver >> to device-tree nodes while coexisting with the fastrpc driver is an open >> item described below. > > No. Grow/replace/improve existing driver instead of coming with a duplicate. > > That's a standard upstream requirement, basically given on every > upstreaming guide. > > Please watch old talk from Greg - "I Don’t Want Your Code!". Posted discussion threads here[1]. Would seek comments from Dmitry, Srini as well. [1] https://lore.kernel.org/all/3476b5c3-7983-4994-a901-3d7d8bd75255@oss.qualcomm.com/ > >> >> v1: https://lore.kernel.org/all/20260519-qda-series-v1-0-b2d984c297f8@oss.qualcomm.com/ >> RFC: https://lore.kernel.org/dri-devel/20260224-qda-firstpost-v1-0-fe46a9c1a046@oss.qualcomm.com/T/ >> >> Changes since v1 >> ================ >> >> The v1 review raised two architectural objections and one correctness >> issue; all three are resolved in v2: >> >> * Christian König (dma-buf maintainer) pointed out that the imported- >> buffer path silently assumed the IOMMU maps every buffer as a single >> contiguous range, which is not guaranteed. v2 walks the scatterlist >> and cleanly rejects non-contiguous imports; contiguous imports (e.g. >> CMA DMA-buf heap) are accepted. (patch 11) >> >> * Dmitry Baryshkov objected to three different buffer-passing formats >> in the invoke IOCTL (DMA-BUF fd, direct/inline, DMA handle). v2 >> passes only GEM handles; userspace imports any fd to a GEM handle >> with DRM_IOCTL_PRIME_FD_TO_HANDLE before invoking. Packing and >> overlap handling are left to userspace. (patch 12) >> >> * The memory manager (patch 07) used a fixed 16-entry array without >> justification and leaked the device descriptor on teardown. v2 >> allocates the array from the DT context-bank count (as Dmitry >> suggested) and frees it correctly. >> >> User-space staging branch >> ========================= >> https://github.com/qualcomm/fastrpc/tree/accel/staging >> >> Key Features >> ============ >> >> * Standard DRM accelerator interface via /dev/accel/accelN >> * GEM-based buffer management with DMA-BUF import (PRIME) >> * IOMMU-based memory isolation using per-process context banks >> * FastRPC protocol implementation for DSP communication >> * RPMsg transport layer for reliable message passing >> * Support for all DSP domains (ADSP, CDSP, SDSP, GDSP) >> * DRM IOCTL interface for DSP session management, buffer allocation, >> and remote procedure invocation >> >> Architecture >> ============ >> >> 1. DRM Accelerator Framework Integration >> The driver registers as a DRM accel device, exposing a standard >> /dev/accel/accelN character device node. This provides established >> DRM infrastructure for device management, file operations, and >> IOCTL dispatch. >> >> 2. Memory Management >> Buffers are managed as GEM objects with PRIME support for DMA-BUF >> import. This enables buffer sharing with other DRM drivers (GPU, >> camera, video) using standard kernel mechanisms. Only contiguous >> imports are accepted; the driver verifies contiguity at import time >> rather than assuming it. >> >> 3. IOMMU Context Bank Management >> IOMMU context banks (CBs) are represented as proper struct device >> instances on a custom virtual bus (qda-compute-cb). Each CB device >> is registered with the IOMMU subsystem and receives its own IOMMU >> domain, enabling per-session address space isolation. The custom >> bus was introduced because IOMMU context banks are synthetic >> constructs — not real platform devices — and to ensure CB device >> lifetime is strictly subordinate to the parent QDA device. >> See also: https://lore.kernel.org/all/245d602f-3037-4ae3-9af9-d98f37258aae@oss.qualcomm.com/ >> >> 4. Memory Manager Architecture >> The memory manager maintains a registry of IOMMU devices in an >> array sized to the number of context banks described in the device >> tree, and coordinates per-process device assignment with reference- >> counted lifetime management. The DMA-coherent backend allocates >> buffers with SID-prefixed DMA addresses for DSP firmware >> compatibility. >> >> 5. Transport Layer >> RPMsg communication is handled in a dedicated transport layer >> (qda_rpmsg.c), separate from the core DRM driver logic. >> >> 6. Code Organization >> The driver is organized across multiple files (~4800 lines total): >> * qda_drv.c: Core driver and DRM integration >> * qda_rpmsg.c: RPMsg transport layer >> * qda_cb.c: Context bank device management >> * qda_compute_bus.c: Custom virtual bus for CB devices >> * qda_gem.c: GEM object management >> * qda_prime.c: DMA-BUF import (PRIME) >> * qda_memory_manager.c: IOMMU device registry and allocation >> * qda_memory_dma.c: DMA-coherent allocation backend >> * qda_fastrpc.c: FastRPC protocol implementation >> * qda_ioctl.c: IOCTL dispatch >> >> 7. UAPI Design >> The driver exposes DRM-style IOCTLs defined in >> include/uapi/drm/qda_accel.h, following DRM UAPI conventions >> (__u32/__u64 types, C++ guard, GPL-2.0-only WITH Linux-syscall-note). >> Buffer arguments are identified by GEM handles; the driver never >> accepts DMA-BUF fds directly in any IOCTL. >> >> Patch Series Organization >> ========================== >> >> Patch 01: MAINTAINERS entry >> Patch 02: Driver documentation (Documentation/accel/qda/) >> Patches 03-04: Core driver skeleton and compute bus >> Patch 05: iommu: Register qda-compute-cb bus with IOMMU subsystem >> Patches 06-07: CB device enumeration and memory manager >> Patch 08: QUERY IOCTL and UAPI header >> Patches 09-11: GEM buffer management and PRIME import >> Patches 12-15: FastRPC protocol (invoke, session create/release, >> map/unmap) >> >> Open Items >> =========== >> >> 1. Device-Tree Compatible String >> The QDA driver uses the same device-tree node structure and >> properties as the existing fastrpc driver in drivers/misc/. A >> mechanism is needed to allow the QDA driver to bind to its device >> node independently of the fastrpc driver. >> >> The intended coexistence model is: platforms that require the >> complete fastrpc feature set continue to use "qcom,fastrpc"; new >> platforms where QDA's feature set is sufficient use a QDA-specific >> compatible string. New feature development is directed toward QDA. >> >> The options under consideration are: >> >> a) Add a new "qcom,qda" compatible string to the existing >> qcom,fastrpc.yaml binding, since the DT node structure and >> properties are identical. > No > >> >> b) Introduce a separate qcom,qda.yaml binding that references or >> inherits the fastrpc binding properties. > > No > >> >> Seeking guidance from DT binding maintainers on the preferred >> approach. > > Grow existing driver. You do not get new driver, you do not get new > bindings. > > > Best regards, > Krzysztof
On 19/08/2026 15:26, Ekansh Gupta wrote: > On 19-08-2026 00:43, Krzysztof Kozlowski wrote: >> On 17/08/2026 06:47, Ekansh Gupta wrote: >>> This patch series introduces the Qualcomm DSP Accelerator (QDA) driver, >>> a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a >>> standardized interface for offloading computational tasks to DSPs found >>> on Qualcomm SoCs, supporting all DSP domains. >>> >>> The QDA driver implements the FastRPC protocol over the DRM accel >>> subsystem. It uses the same device-tree node structure as the existing >>> fastrpc driver in drivers/misc/. The approach for binding the QDA driver >>> to device-tree nodes while coexisting with the fastrpc driver is an open >>> item described below. >> >> No. Grow/replace/improve existing driver instead of coming with a duplicate. >> >> That's a standard upstream requirement, basically given on every >> upstreaming guide. >> >> Please watch old talk from Greg - "I Don’t Want Your Code!". > Posted discussion threads here[1]. Would seek comments from Dmitry, > Srini as well. > > [1] > https://lore.kernel.org/all/3476b5c3-7983-4994-a901-3d7d8bd75255@oss.qualcomm.com/ The rest of the comments is still valid even if you did not acknowledge them. Anyway, regarding above - again, watch the talk from Greg. You have ONE driver. Not two. Best regards, Krzysztof
On Wed, Aug 19, 2026 at 7:21 AM Krzysztof Kozlowski <krzk@kernel.org> wrote: > > On 19/08/2026 15:26, Ekansh Gupta wrote: > > On 19-08-2026 00:43, Krzysztof Kozlowski wrote: > >> On 17/08/2026 06:47, Ekansh Gupta wrote: > >>> This patch series introduces the Qualcomm DSP Accelerator (QDA) driver, > >>> a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a > >>> standardized interface for offloading computational tasks to DSPs found > >>> on Qualcomm SoCs, supporting all DSP domains. > >>> > >>> The QDA driver implements the FastRPC protocol over the DRM accel > >>> subsystem. It uses the same device-tree node structure as the existing > >>> fastrpc driver in drivers/misc/. The approach for binding the QDA driver > >>> to device-tree nodes while coexisting with the fastrpc driver is an open > >>> item described below. > >> > >> No. Grow/replace/improve existing driver instead of coming with a duplicate. > >> > >> That's a standard upstream requirement, basically given on every > >> upstreaming guide. > >> > >> Please watch old talk from Greg - "I Don’t Want Your Code!". > > Posted discussion threads here[1]. Would seek comments from Dmitry, > > Srini as well. > > > > [1] > > https://lore.kernel.org/all/3476b5c3-7983-4994-a901-3d7d8bd75255@oss.qualcomm.com/ > > The rest of the comments is still valid even if you did not acknowledge > them. > > Anyway, regarding above - again, watch the talk from Greg. > > You have ONE driver. Not two. Long term, moving to the common driver framework (which did not exist when fastrpc was first created) seems like a good thing. But does that not allow for some transition period? How can we get from here to there without otherwise breaking userspace? Is there some other precedent elsewhere in other driver subsystems? I suppose drm exposing legacy fbdev on top of drm drivers is _sort of_ a precedent if you squint a bit? I'm not really familiar enough to say if that would be reasonable/possible in this case. BR, -R > Best regards, > Krzysztof
On 19/08/2026 16:38, Rob Clark wrote: > On Wed, Aug 19, 2026 at 7:21 AM Krzysztof Kozlowski <krzk@kernel.org> wrote: >> >> On 19/08/2026 15:26, Ekansh Gupta wrote: >>> On 19-08-2026 00:43, Krzysztof Kozlowski wrote: >>>> On 17/08/2026 06:47, Ekansh Gupta wrote: >>>>> This patch series introduces the Qualcomm DSP Accelerator (QDA) driver, >>>>> a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a >>>>> standardized interface for offloading computational tasks to DSPs found >>>>> on Qualcomm SoCs, supporting all DSP domains. >>>>> >>>>> The QDA driver implements the FastRPC protocol over the DRM accel >>>>> subsystem. It uses the same device-tree node structure as the existing >>>>> fastrpc driver in drivers/misc/. The approach for binding the QDA driver >>>>> to device-tree nodes while coexisting with the fastrpc driver is an open >>>>> item described below. >>>> >>>> No. Grow/replace/improve existing driver instead of coming with a duplicate. >>>> >>>> That's a standard upstream requirement, basically given on every >>>> upstreaming guide. >>>> >>>> Please watch old talk from Greg - "I Don’t Want Your Code!". >>> Posted discussion threads here[1]. Would seek comments from Dmitry, >>> Srini as well. >>> >>> [1] >>> https://lore.kernel.org/all/3476b5c3-7983-4994-a901-3d7d8bd75255@oss.qualcomm.com/ >> >> The rest of the comments is still valid even if you did not acknowledge >> them. >> >> Anyway, regarding above - again, watch the talk from Greg. >> >> You have ONE driver. Not two. > > Long term, moving to the common driver framework (which did not exist > when fastrpc was first created) seems like a good thing. But does > that not allow for some transition period? How can we get from here > to there without otherwise breaking userspace? Is there some other > precedent elsewhere in other driver subsystems? Yes, Iris and Venus where we agreed for an exception (two drivers) as long as new driver supports old hardware / features. This is not the case here, right? > > I suppose drm exposing legacy fbdev on top of drm drivers is _sort of_ > a precedent if you squint a bit? I'm not really familiar enough to > say if that would be reasonable/possible in this case. Growing old driver does not look complicated itself. The only a bit tricky thing is to present somehow exclusive interface to user-space, like usage of one disables the second etc. Depending on actual differences in that interface. But having a duplicated driver is a clear no go and it is well known upstream requirement. Nothing new here. Best regards, Krzysztof
On Wed, Aug 19, 2026 at 7:43 AM Krzysztof Kozlowski <krzk@kernel.org> wrote: > > On 19/08/2026 16:38, Rob Clark wrote: > > On Wed, Aug 19, 2026 at 7:21 AM Krzysztof Kozlowski <krzk@kernel.org> wrote: > >> > >> On 19/08/2026 15:26, Ekansh Gupta wrote: > >>> On 19-08-2026 00:43, Krzysztof Kozlowski wrote: > >>>> On 17/08/2026 06:47, Ekansh Gupta wrote: > >>>>> This patch series introduces the Qualcomm DSP Accelerator (QDA) driver, > >>>>> a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a > >>>>> standardized interface for offloading computational tasks to DSPs found > >>>>> on Qualcomm SoCs, supporting all DSP domains. > >>>>> > >>>>> The QDA driver implements the FastRPC protocol over the DRM accel > >>>>> subsystem. It uses the same device-tree node structure as the existing > >>>>> fastrpc driver in drivers/misc/. The approach for binding the QDA driver > >>>>> to device-tree nodes while coexisting with the fastrpc driver is an open > >>>>> item described below. > >>>> > >>>> No. Grow/replace/improve existing driver instead of coming with a duplicate. > >>>> > >>>> That's a standard upstream requirement, basically given on every > >>>> upstreaming guide. > >>>> > >>>> Please watch old talk from Greg - "I Don’t Want Your Code!". > >>> Posted discussion threads here[1]. Would seek comments from Dmitry, > >>> Srini as well. > >>> > >>> [1] > >>> https://lore.kernel.org/all/3476b5c3-7983-4994-a901-3d7d8bd75255@oss.qualcomm.com/ > >> > >> The rest of the comments is still valid even if you did not acknowledge > >> them. > >> > >> Anyway, regarding above - again, watch the talk from Greg. > >> > >> You have ONE driver. Not two. > > > > Long term, moving to the common driver framework (which did not exist > > when fastrpc was first created) seems like a good thing. But does > > that not allow for some transition period? How can we get from here > > to there without otherwise breaking userspace? Is there some other > > precedent elsewhere in other driver subsystems? > > Yes, Iris and Venus where we agreed for an exception (two drivers) as > long as new driver supports old hardware / features. > > This is not the case here, right? > > > > > I suppose drm exposing legacy fbdev on top of drm drivers is _sort of_ > > a precedent if you squint a bit? I'm not really familiar enough to > > say if that would be reasonable/possible in this case. > > Growing old driver does not look complicated itself. The only a bit > tricky thing is to present somehow exclusive interface to user-space, > like usage of one disables the second etc. Depending on actual > differences in that interface. > > But having a duplicated driver is a clear no go and it is well known > upstream requirement. Nothing new here. Would it be acceptable as a first step to, by default (ie. when no cmdline override/etc) for the new driver to bind to new compatibles and the existing driver to old? Ie. have both drivers but only one binds on a given platform? BR, -R
On 19/08/2026 16:49, Rob Clark wrote: > On Wed, Aug 19, 2026 at 7:43 AM Krzysztof Kozlowski <krzk@kernel.org> wrote: >> >> On 19/08/2026 16:38, Rob Clark wrote: >>> On Wed, Aug 19, 2026 at 7:21 AM Krzysztof Kozlowski <krzk@kernel.org> wrote: >>>> >>>> On 19/08/2026 15:26, Ekansh Gupta wrote: >>>>> On 19-08-2026 00:43, Krzysztof Kozlowski wrote: >>>>>> On 17/08/2026 06:47, Ekansh Gupta wrote: >>>>>>> This patch series introduces the Qualcomm DSP Accelerator (QDA) driver, >>>>>>> a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a >>>>>>> standardized interface for offloading computational tasks to DSPs found >>>>>>> on Qualcomm SoCs, supporting all DSP domains. >>>>>>> >>>>>>> The QDA driver implements the FastRPC protocol over the DRM accel >>>>>>> subsystem. It uses the same device-tree node structure as the existing >>>>>>> fastrpc driver in drivers/misc/. The approach for binding the QDA driver >>>>>>> to device-tree nodes while coexisting with the fastrpc driver is an open >>>>>>> item described below. >>>>>> >>>>>> No. Grow/replace/improve existing driver instead of coming with a duplicate. >>>>>> >>>>>> That's a standard upstream requirement, basically given on every >>>>>> upstreaming guide. >>>>>> >>>>>> Please watch old talk from Greg - "I Don’t Want Your Code!". >>>>> Posted discussion threads here[1]. Would seek comments from Dmitry, >>>>> Srini as well. >>>>> >>>>> [1] >>>>> https://lore.kernel.org/all/3476b5c3-7983-4994-a901-3d7d8bd75255@oss.qualcomm.com/ >>>> >>>> The rest of the comments is still valid even if you did not acknowledge >>>> them. >>>> >>>> Anyway, regarding above - again, watch the talk from Greg. >>>> >>>> You have ONE driver. Not two. >>> >>> Long term, moving to the common driver framework (which did not exist >>> when fastrpc was first created) seems like a good thing. But does >>> that not allow for some transition period? How can we get from here >>> to there without otherwise breaking userspace? Is there some other >>> precedent elsewhere in other driver subsystems? >> >> Yes, Iris and Venus where we agreed for an exception (two drivers) as >> long as new driver supports old hardware / features. >> >> This is not the case here, right? >> >>> >>> I suppose drm exposing legacy fbdev on top of drm drivers is _sort of_ >>> a precedent if you squint a bit? I'm not really familiar enough to >>> say if that would be reasonable/possible in this case. >> >> Growing old driver does not look complicated itself. The only a bit >> tricky thing is to present somehow exclusive interface to user-space, >> like usage of one disables the second etc. Depending on actual >> differences in that interface. >> >> But having a duplicated driver is a clear no go and it is well known >> upstream requirement. Nothing new here. > > Would it be acceptable as a first step to, by default (ie. when no > cmdline override/etc) for the new driver to bind to new compatibles You have one compatible. > and the existing driver to old? Ie. have both drivers but only one > binds on a given platform? I don't see how is it possible to write such DTS, because - repeating my question - how many hardware blocks is there? I believe only one per given DSP, so how could you have two device nodes? Best regards, Krzysztof
On Wed, Aug 19, 2026 at 7:53 AM Krzysztof Kozlowski <krzk@kernel.org> wrote: > > On 19/08/2026 16:49, Rob Clark wrote: > > On Wed, Aug 19, 2026 at 7:43 AM Krzysztof Kozlowski <krzk@kernel.org> wrote: > >> > >> On 19/08/2026 16:38, Rob Clark wrote: > >>> On Wed, Aug 19, 2026 at 7:21 AM Krzysztof Kozlowski <krzk@kernel.org> wrote: > >>>> > >>>> On 19/08/2026 15:26, Ekansh Gupta wrote: > >>>>> On 19-08-2026 00:43, Krzysztof Kozlowski wrote: > >>>>>> On 17/08/2026 06:47, Ekansh Gupta wrote: > >>>>>>> This patch series introduces the Qualcomm DSP Accelerator (QDA) driver, > >>>>>>> a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a > >>>>>>> standardized interface for offloading computational tasks to DSPs found > >>>>>>> on Qualcomm SoCs, supporting all DSP domains. > >>>>>>> > >>>>>>> The QDA driver implements the FastRPC protocol over the DRM accel > >>>>>>> subsystem. It uses the same device-tree node structure as the existing > >>>>>>> fastrpc driver in drivers/misc/. The approach for binding the QDA driver > >>>>>>> to device-tree nodes while coexisting with the fastrpc driver is an open > >>>>>>> item described below. > >>>>>> > >>>>>> No. Grow/replace/improve existing driver instead of coming with a duplicate. > >>>>>> > >>>>>> That's a standard upstream requirement, basically given on every > >>>>>> upstreaming guide. > >>>>>> > >>>>>> Please watch old talk from Greg - "I Don’t Want Your Code!". > >>>>> Posted discussion threads here[1]. Would seek comments from Dmitry, > >>>>> Srini as well. > >>>>> > >>>>> [1] > >>>>> https://lore.kernel.org/all/3476b5c3-7983-4994-a901-3d7d8bd75255@oss.qualcomm.com/ > >>>> > >>>> The rest of the comments is still valid even if you did not acknowledge > >>>> them. > >>>> > >>>> Anyway, regarding above - again, watch the talk from Greg. > >>>> > >>>> You have ONE driver. Not two. > >>> > >>> Long term, moving to the common driver framework (which did not exist > >>> when fastrpc was first created) seems like a good thing. But does > >>> that not allow for some transition period? How can we get from here > >>> to there without otherwise breaking userspace? Is there some other > >>> precedent elsewhere in other driver subsystems? > >> > >> Yes, Iris and Venus where we agreed for an exception (two drivers) as > >> long as new driver supports old hardware / features. > >> > >> This is not the case here, right? > >> > >>> > >>> I suppose drm exposing legacy fbdev on top of drm drivers is _sort of_ > >>> a precedent if you squint a bit? I'm not really familiar enough to > >>> say if that would be reasonable/possible in this case. > >> > >> Growing old driver does not look complicated itself. The only a bit > >> tricky thing is to present somehow exclusive interface to user-space, > >> like usage of one disables the second etc. Depending on actual > >> differences in that interface. > >> > >> But having a duplicated driver is a clear no go and it is well known > >> upstream requirement. Nothing new here. > > > > Would it be acceptable as a first step to, by default (ie. when no > > cmdline override/etc) for the new driver to bind to new compatibles > > You have one compatible. > > > and the existing driver to old? Ie. have both drivers but only one > > binds on a given platform? > > I don't see how is it possible to write such DTS, because - repeating my > question - how many hardware blocks is there? I believe only one per > given DSP, so how could you have two device nodes? Maybe I'm misunderstanding something here.. I don't see any new bindings with this series so my assumption is that the bindings are the same. What I meant was something more like "qcom,glymur-fastrpc" would bind to new driver but "qcom,fastrpc" (which seems to be what is used on older platforms) would bind to the old driver. So single dts node, but different compatible strings picking which driver is used. BR, -R
On 19/08/2026 17:23, Rob Clark wrote: >>>> >>>> Growing old driver does not look complicated itself. The only a bit >>>> tricky thing is to present somehow exclusive interface to user-space, >>>> like usage of one disables the second etc. Depending on actual >>>> differences in that interface. >>>> >>>> But having a duplicated driver is a clear no go and it is well known >>>> upstream requirement. Nothing new here. >>> >>> Would it be acceptable as a first step to, by default (ie. when no >>> cmdline override/etc) for the new driver to bind to new compatibles >> >> You have one compatible. >> >>> and the existing driver to old? Ie. have both drivers but only one >>> binds on a given platform? >> >> I don't see how is it possible to write such DTS, because - repeating my >> question - how many hardware blocks is there? I believe only one per >> given DSP, so how could you have two device nodes? > > Maybe I'm misunderstanding something here.. I don't see any new > bindings with this series so my assumption is that the bindings are > the same. What I meant was something more like "qcom,glymur-fastrpc" > would bind to new driver but "qcom,fastrpc" (which seems to be what is > used on older platforms) would bind to the old driver. > > So single dts node, but different compatible strings picking which > driver is used. Glymur is already done, so imagining we talk about next/future SoC then we would be at point of duplicating drivers for the same hardware. So back to square one of my comments. The rule of usptream development is that we do not accept duplicated code, just because a vendor wants to write something new. This is basically the concept applied all over the drivers tree, where we pushed back against all sorts of duplications all over the vendors. What I miss in this thread is why would there be any exception here. We do not grant exceptions from standard practices on "I want" reasons. Best regards, Krzysztof
On Wed, Aug 19, 2026 at 8:27 AM Krzysztof Kozlowski <krzk@kernel.org> wrote: > > On 19/08/2026 17:23, Rob Clark wrote: > >>>> > >>>> Growing old driver does not look complicated itself. The only a bit > >>>> tricky thing is to present somehow exclusive interface to user-space, > >>>> like usage of one disables the second etc. Depending on actual > >>>> differences in that interface. > >>>> > >>>> But having a duplicated driver is a clear no go and it is well known > >>>> upstream requirement. Nothing new here. > >>> > >>> Would it be acceptable as a first step to, by default (ie. when no > >>> cmdline override/etc) for the new driver to bind to new compatibles > >> > >> You have one compatible. > >> > >>> and the existing driver to old? Ie. have both drivers but only one > >>> binds on a given platform? > >> > >> I don't see how is it possible to write such DTS, because - repeating my > >> question - how many hardware blocks is there? I believe only one per > >> given DSP, so how could you have two device nodes? > > > > Maybe I'm misunderstanding something here.. I don't see any new > > bindings with this series so my assumption is that the bindings are > > the same. What I meant was something more like "qcom,glymur-fastrpc" > > would bind to new driver but "qcom,fastrpc" (which seems to be what is > > used on older platforms) would bind to the old driver. > > > > So single dts node, but different compatible strings picking which > > driver is used. > > Glymur is already done, so imagining we talk about next/future SoC then > we would be at point of duplicating drivers for the same hardware. So > back to square one of my comments. I just picked glymur as an example, because I needed to pick _something_, and toplevel dts for devices available outside of qcom is only just starting to land. Maybe it wasn't the right example. Let's just say "qcom,supercoolfuturething-fastrpc" instead. I think my example otherwise still stands. (In either case, kernel cmdline or similar override to get the new driver instead of old would be a good idea.) > The rule of usptream development is that we do not accept duplicated > code, just because a vendor wants to write something new. This is > basically the concept applied all over the drivers tree, where we pushed > back against all sorts of duplications all over the vendors. > > What I miss in this thread is why would there be any exception here. We > do not grant exceptions from standard practices on "I want" reasons. I agree that we should not have duplicated drivers just for vendor lolz. But when it comes to adopting common frameworks and integrating better into the ecosystem, this doesn't seem like something we should actively discourage. I don't think this is a case of vendor lolz, but rather reacting to drm/accel emerging as the standard framework for this sort of driver. So how do we get from here to there? BR, -R
On 19/08/2026 17:48, Rob Clark wrote: > On Wed, Aug 19, 2026 at 8:27 AM Krzysztof Kozlowski <krzk@kernel.org> wrote: >> The rule of usptream development is that we do not accept duplicated >> code, just because a vendor wants to write something new. This is >> basically the concept applied all over the drivers tree, where we pushed >> back against all sorts of duplications all over the vendors. >> >> What I miss in this thread is why would there be any exception here. We >> do not grant exceptions from standard practices on "I want" reasons. > > I agree that we should not have duplicated drivers just for vendor > lolz. But when it comes to adopting common frameworks and integrating > better into the ecosystem, this doesn't seem like something we should > actively discourage. I don't think this is a case of vendor lolz, but No one discourages it. Following standard Linux kernel practices and requirements is not discouraging, do not twist the narrative here. Again, it is standard upstream review telling that we do not duplicate drivers. Ever, unless there is serious exception needed. I asked why there should be an exception granted? Is the reason for exception following: "We want to adopt common framework" ? > rather reacting to drm/accel emerging as the standard framework for > this sort of driver. > > So how do we get from here to there? What is wrong with my proposal? Best regards, Krzysztof
On Wed, Aug 19, 2026 at 11:15 PM Krzysztof Kozlowski <krzk@kernel.org> wrote: > > On 19/08/2026 17:48, Rob Clark wrote: > > On Wed, Aug 19, 2026 at 8:27 AM Krzysztof Kozlowski <krzk@kernel.org> wrote: > >> The rule of usptream development is that we do not accept duplicated > >> code, just because a vendor wants to write something new. This is > >> basically the concept applied all over the drivers tree, where we pushed > >> back against all sorts of duplications all over the vendors. > >> > >> What I miss in this thread is why would there be any exception here. We > >> do not grant exceptions from standard practices on "I want" reasons. > > > > I agree that we should not have duplicated drivers just for vendor > > lolz. But when it comes to adopting common frameworks and integrating > > better into the ecosystem, this doesn't seem like something we should > > actively discourage. I don't think this is a case of vendor lolz, but > > No one discourages it. Following standard Linux kernel practices and > requirements is not discouraging, do not twist the narrative here. > Again, it is standard upstream review telling that we do not duplicate > drivers. Ever, unless there is serious exception needed. I wasn't trying to twist the narrative, just trying to come up with a path forward that isn't "no" or "improve existing driver", since neither of those gets us towards a future using common frameworks. > I asked why there should be an exception granted? Is the reason for > exception following: > "We want to adopt common framework" > ? Possibly? But I don't think we want two drivers to be any sort of long term solution. (Ie. as long as venus/iris have co-exist.) > > > rather reacting to drm/accel emerging as the standard framework for > > this sort of driver. > > > > So how do we get from here to there? > > What is wrong with my proposal? Maybe I missed something, my understanding was your proposal was "Grow/replace/improve existing driver instead of coming with a duplicate".. grow or improve doesn't move us toward common frameworks. Maybe "replace" is a valid option. If there is something I missed, then I apologize. Options I can think of are: 1. Hardware cutoff.. new hw gets new driver, existing hw gets existing driver 2. Backwards compat chardev registered by new driver, providing existing UABI. I'm not 100% sure about the feasibility/drawbacks of this.. AFAIU the fastrpc folks where planning a backwards compat layer in userspace, so maybe it is possible. 3. exception? I'd like to know what the feasibility of #2 is, since at a high level that sounds like the best option. Possibly limit exposure of legacy UABI to existing hw so we don't get into a place of needing to extend the legacy UABI for new hw? But #1 sounds like a non-controversial place to start regardless. Possibly with #2 coming as followup and necessary step before eventual migration to new driver for existing hw? Even if we start with #2, how do we handle first-merge-window bugs/regressions without reverting addition of new driver and removal of old? It seems like we'd need a window of a couple release cycles where both drivers exist? Maybe others have other/better options in mind? BR, -R
On 20-08-2026 20:17, Rob Clark wrote: > On Wed, Aug 19, 2026 at 11:15 PM Krzysztof Kozlowski <krzk@kernel.org> wrote: >> >> On 19/08/2026 17:48, Rob Clark wrote: >>> On Wed, Aug 19, 2026 at 8:27 AM Krzysztof Kozlowski <krzk@kernel.org> wrote: >>>> The rule of usptream development is that we do not accept duplicated >>>> code, just because a vendor wants to write something new. This is >>>> basically the concept applied all over the drivers tree, where we pushed >>>> back against all sorts of duplications all over the vendors. >>>> >>>> What I miss in this thread is why would there be any exception here. We >>>> do not grant exceptions from standard practices on "I want" reasons. >>> >>> I agree that we should not have duplicated drivers just for vendor >>> lolz. But when it comes to adopting common frameworks and integrating >>> better into the ecosystem, this doesn't seem like something we should >>> actively discourage. I don't think this is a case of vendor lolz, but >> >> No one discourages it. Following standard Linux kernel practices and >> requirements is not discouraging, do not twist the narrative here. >> Again, it is standard upstream review telling that we do not duplicate >> drivers. Ever, unless there is serious exception needed. > > I wasn't trying to twist the narrative, just trying to come up with a > path forward that isn't "no" or "improve existing driver", since > neither of those gets us towards a future using common frameworks. > >> I asked why there should be an exception granted? Is the reason for >> exception following: >> "We want to adopt common framework" >> ? > > Possibly? But I don't think we want two drivers to be any sort of > long term solution. (Ie. as long as venus/iris have co-exist.) > >> >>> rather reacting to drm/accel emerging as the standard framework for >>> this sort of driver. >>> >>> So how do we get from here to there? >> >> What is wrong with my proposal? > > Maybe I missed something, my understanding was your proposal was > "Grow/replace/improve existing driver instead of coming with a > duplicate".. grow or improve doesn't move us toward common > frameworks. Maybe "replace" is a valid option. If there is something > I missed, then I apologize. > > Options I can think of are: > > 1. Hardware cutoff.. new hw gets new driver, existing hw gets existing > driver > 2. Backwards compat chardev registered by new driver, providing existing > UABI. I'm not 100% sure about the feasibility/drawbacks of this.. > AFAIU the fastrpc folks where planning a backwards compat layer in > userspace, so maybe it is possible. > 3. exception? > > I'd like to know what the feasibility of #2 is, since at a high level > that sounds like the best option. Possibly limit exposure of legacy > UABI to existing hw so we don't get into a place of needing to extend > the legacy UABI for new hw? > > But #1 sounds like a non-controversial place to start regardless. > Possibly with #2 coming as followup and necessary step before eventual > migration to new driver for existing hw? > > Even if we start with #2, how do we handle first-merge-window > bugs/regressions without reverting addition of new driver and removal > of old? It seems like we'd need a window of a couple release cycles > where both drivers exist? > > Maybe others have other/better options in mind? To all, I'm seeking on the approach I should follow to go ahead here. I can work on implementing #1(as per Rob's list) with hw specific compatible for v4 if it's acceptable. #2(compat driver) is something that we are still exploring as we couldn't find any standard way to achieve it. We might start a separate discussion for that once we have few possible designs with us. Happy to take any other suggestion also.> > BR, > -R
On Wed, Aug 26, 2026 at 06:37:00PM +0530, Ekansh Gupta wrote: > On 20-08-2026 20:17, Rob Clark wrote: > > On Wed, Aug 19, 2026 at 11:15 PM Krzysztof Kozlowski <krzk@kernel.org> wrote: > >> > >> On 19/08/2026 17:48, Rob Clark wrote: > >>> On Wed, Aug 19, 2026 at 8:27 AM Krzysztof Kozlowski <krzk@kernel.org> wrote: > >>>> The rule of usptream development is that we do not accept duplicated > >>>> code, just because a vendor wants to write something new. This is > >>>> basically the concept applied all over the drivers tree, where we pushed > >>>> back against all sorts of duplications all over the vendors. > >>>> > >>>> What I miss in this thread is why would there be any exception here. We > >>>> do not grant exceptions from standard practices on "I want" reasons. > >>> > >>> I agree that we should not have duplicated drivers just for vendor > >>> lolz. But when it comes to adopting common frameworks and integrating > >>> better into the ecosystem, this doesn't seem like something we should > >>> actively discourage. I don't think this is a case of vendor lolz, but > >> > >> No one discourages it. Following standard Linux kernel practices and > >> requirements is not discouraging, do not twist the narrative here. > >> Again, it is standard upstream review telling that we do not duplicate > >> drivers. Ever, unless there is serious exception needed. > > > > I wasn't trying to twist the narrative, just trying to come up with a > > path forward that isn't "no" or "improve existing driver", since > > neither of those gets us towards a future using common frameworks. > > > >> I asked why there should be an exception granted? Is the reason for > >> exception following: > >> "We want to adopt common framework" > >> ? > > > > Possibly? But I don't think we want two drivers to be any sort of > > long term solution. (Ie. as long as venus/iris have co-exist.) > > > >> > >>> rather reacting to drm/accel emerging as the standard framework for > >>> this sort of driver. > >>> > >>> So how do we get from here to there? > >> > >> What is wrong with my proposal? > > > > Maybe I missed something, my understanding was your proposal was > > "Grow/replace/improve existing driver instead of coming with a > > duplicate".. grow or improve doesn't move us toward common > > frameworks. Maybe "replace" is a valid option. If there is something > > I missed, then I apologize. > > > > Options I can think of are: > > > > 1. Hardware cutoff.. new hw gets new driver, existing hw gets existing > > driver > > 2. Backwards compat chardev registered by new driver, providing existing > > UABI. I'm not 100% sure about the feasibility/drawbacks of this.. > > AFAIU the fastrpc folks where planning a backwards compat layer in > > userspace, so maybe it is possible. > > 3. exception? > > > > I'd like to know what the feasibility of #2 is, since at a high level > > that sounds like the best option. Possibly limit exposure of legacy > > UABI to existing hw so we don't get into a place of needing to extend > > the legacy UABI for new hw? > > > > But #1 sounds like a non-controversial place to start regardless. > > Possibly with #2 coming as followup and necessary step before eventual > > migration to new driver for existing hw? > > > > Even if we start with #2, how do we handle first-merge-window > > bugs/regressions without reverting addition of new driver and removal > > of old? It seems like we'd need a window of a couple release cycles > > where both drivers exist? > > > > Maybe others have other/better options in mind? > To all, I'm seeking on the approach I should follow to go ahead here. I > can work on implementing #1(as per Rob's list) with hw specific > compatible for v4 if it's acceptable. > I don't see any reason for you to define a "hw specific compatible", because as you have shown in this series (and as Rob point out), there's no difference in the "hardware". The only reason for your "hw specific compatible" is to make a software selection in Linux - and that's not what DeviceTree is for. As such, I don't see that you have a DeviceTree problem at all, because this is a Linux-internal problem. > #2(compat driver) is something that we are still exploring as we > couldn't find any standard way to achieve it. We might start a separate > discussion for that once we have few possible designs with us. > This is the actual problem! We have existing user space that depends on the ioctl interface exposed by the current misc driver. You must not break these. Hardware cutoff is not a viable solution, because that's just a declaration that we'll let the old platforms rotten - or alternatively you commit to maintain two drivers to the very same feature and quality level. So the only reasonable solution is #2; from there it's a valid question if you reach that point my stepwise migrating the current misc driver that solution, or if you present a new driver with the fully backwards compatible interface, alongside the new ABI. But this does bring to a question which the cover letter should explain - but doesn't: what problem does this patch series actually solve? Regards, Bjorn > Happy to take any other suggestion also.> > > BR, > > -R > >
On Tue, Sep 08, 2026 at 05:44:10PM -0500, Bjorn Andersson wrote: > On Wed, Aug 26, 2026 at 06:37:00PM +0530, Ekansh Gupta wrote: > > On 20-08-2026 20:17, Rob Clark wrote: > > > On Wed, Aug 19, 2026 at 11:15 PM Krzysztof Kozlowski <krzk@kernel.org> wrote: > > >> > > >> On 19/08/2026 17:48, Rob Clark wrote: > > >>> On Wed, Aug 19, 2026 at 8:27 AM Krzysztof Kozlowski <krzk@kernel.org> wrote: > > >>>> The rule of usptream development is that we do not accept duplicated > > >>>> code, just because a vendor wants to write something new. This is > > >>>> basically the concept applied all over the drivers tree, where we pushed > > >>>> back against all sorts of duplications all over the vendors. > > >>>> > > >>>> What I miss in this thread is why would there be any exception here. We > > >>>> do not grant exceptions from standard practices on "I want" reasons. > > >>> > > >>> I agree that we should not have duplicated drivers just for vendor > > >>> lolz. But when it comes to adopting common frameworks and integrating > > >>> better into the ecosystem, this doesn't seem like something we should > > >>> actively discourage. I don't think this is a case of vendor lolz, but > > >> > > >> No one discourages it. Following standard Linux kernel practices and > > >> requirements is not discouraging, do not twist the narrative here. > > >> Again, it is standard upstream review telling that we do not duplicate > > >> drivers. Ever, unless there is serious exception needed. > > > > > > I wasn't trying to twist the narrative, just trying to come up with a > > > path forward that isn't "no" or "improve existing driver", since > > > neither of those gets us towards a future using common frameworks. > > > > > >> I asked why there should be an exception granted? Is the reason for > > >> exception following: > > >> "We want to adopt common framework" > > >> ? > > > > > > Possibly? But I don't think we want two drivers to be any sort of > > > long term solution. (Ie. as long as venus/iris have co-exist.) > > > > > >> > > >>> rather reacting to drm/accel emerging as the standard framework for > > >>> this sort of driver. > > >>> > > >>> So how do we get from here to there? > > >> > > >> What is wrong with my proposal? > > > > > > Maybe I missed something, my understanding was your proposal was > > > "Grow/replace/improve existing driver instead of coming with a > > > duplicate".. grow or improve doesn't move us toward common > > > frameworks. Maybe "replace" is a valid option. If there is something > > > I missed, then I apologize. > > > > > > Options I can think of are: > > > > > > 1. Hardware cutoff.. new hw gets new driver, existing hw gets existing > > > driver > > > 2. Backwards compat chardev registered by new driver, providing existing > > > UABI. I'm not 100% sure about the feasibility/drawbacks of this.. > > > AFAIU the fastrpc folks where planning a backwards compat layer in > > > userspace, so maybe it is possible. > > > 3. exception? > > > > > > I'd like to know what the feasibility of #2 is, since at a high level > > > that sounds like the best option. Possibly limit exposure of legacy > > > UABI to existing hw so we don't get into a place of needing to extend > > > the legacy UABI for new hw? > > > > > > But #1 sounds like a non-controversial place to start regardless. > > > Possibly with #2 coming as followup and necessary step before eventual > > > migration to new driver for existing hw? > > > > > > Even if we start with #2, how do we handle first-merge-window > > > bugs/regressions without reverting addition of new driver and removal > > > of old? It seems like we'd need a window of a couple release cycles > > > where both drivers exist? > > > > > > Maybe others have other/better options in mind? > > To all, I'm seeking on the approach I should follow to go ahead here. I > > can work on implementing #1(as per Rob's list) with hw specific > > compatible for v4 if it's acceptable. > > > > I don't see any reason for you to define a "hw specific compatible", > because as you have shown in this series (and as Rob point out), there's > no difference in the "hardware". > > The only reason for your "hw specific compatible" is to make a software > selection in Linux - and that's not what DeviceTree is for. That's not exactly true. There are protocol differences. For example, polling mode is supported only since a certain timeline in the history. Likewise other interface features are not supported on all the FastRPC devices. For the polling mode support we were already beaten by the lack of SoC-specific compats. > As such, I don't see that you have a DeviceTree problem at all, because > this is a Linux-internal problem. > > > #2(compat driver) is something that we are still exploring as we > > couldn't find any standard way to achieve it. We might start a separate > > discussion for that once we have few possible designs with us. > > > > This is the actual problem! > > We have existing user space that depends on the ioctl interface exposed > by the current misc driver. You must not break these. This is clear. > Hardware cutoff is not a viable solution, because that's just a > declaration that we'll let the old platforms rotten - or alternatively > you commit to maintain two drivers to the very same feature and quality > level. > > So the only reasonable solution is #2; from there it's a valid question > if you reach that point my stepwise migrating the current misc driver > that solution, or if you present a new driver with the fully backwards > compatible interface, alongside the new ABI. I think the general direction was #3 (or #2.1): implement a shim layer on top of the QDA driver as a separate module. Put all the historical over-complicated solutions into that shim module and let it die at some point. Current fastrpc driver lets userspace specify buffers in several different ways, forcing the kernel driver to perform a lot of work with buffer addresses. I don't think that this legacy code should be a part of the QDA. It can go to qda-fastrpc-shim, keeping it out of the kernel memory if it's not necessary. > But this does bring to a question which the cover letter should explain > - but doesn't: what problem does this patch series actually solve? I agree that it should be a part of the cover letter. As a person who triggered this work, I can propose my reasons: - Current driver has over-complicated memory manager (both on the userspace and on the kernel side). Correspondng uAPI is not really suitable for virtualization. Using GEM simplifies both the kernel driver and uAPI. Also using handle-offset-length to specify the buffers makes it easy to support virtual QDA devices. - Current driver predates the accel subsystem. Using common subsystem simplifies reviews. The QDA driver has gotten several comments about the usage of the DMA-BUFs. It'not unlikely that the same issues are present in the current FastRPC driver, just being unnoticed. - The ideas present in the current FastRPC driver also predate the current design practices. The uAPI was created in the ad-hoc way, just following the momentary needs. Driver code also shows the result of that, having enough of the spaghetti code. Given all of that, yes, it is possible to provide an evolution of the FastRPC driver into the accel+shim, improve the code quality meanwhile and end up with the good enough split. However I think that the path taken would be longer and the net result might be worse. With all of that in mind, my suggestion is to continue working on the QDA driver, get the core of it to integrate nicely with the accel and DMA-BUF usage requirements and implement the (minimal?) fastrpc shim on top of it. -- With best wishes Dmitry
On 09-09-2026 17:18, Dmitry Baryshkov wrote: > On Tue, Sep 08, 2026 at 05:44:10PM -0500, Bjorn Andersson wrote: >> On Wed, Aug 26, 2026 at 06:37:00PM +0530, Ekansh Gupta wrote: >>> On 20-08-2026 20:17, Rob Clark wrote: >>>> On Wed, Aug 19, 2026 at 11:15 PM Krzysztof Kozlowski <krzk@kernel.org> wrote: >>>>> >>>>> On 19/08/2026 17:48, Rob Clark wrote: >>>>>> On Wed, Aug 19, 2026 at 8:27 AM Krzysztof Kozlowski <krzk@kernel.org> wrote: >>>>>>> The rule of usptream development is that we do not accept duplicated >>>>>>> code, just because a vendor wants to write something new. This is >>>>>>> basically the concept applied all over the drivers tree, where we pushed >>>>>>> back against all sorts of duplications all over the vendors. >>>>>>> >>>>>>> What I miss in this thread is why would there be any exception here. We >>>>>>> do not grant exceptions from standard practices on "I want" reasons. >>>>>> >>>>>> I agree that we should not have duplicated drivers just for vendor >>>>>> lolz. But when it comes to adopting common frameworks and integrating >>>>>> better into the ecosystem, this doesn't seem like something we should >>>>>> actively discourage. I don't think this is a case of vendor lolz, but >>>>> >>>>> No one discourages it. Following standard Linux kernel practices and >>>>> requirements is not discouraging, do not twist the narrative here. >>>>> Again, it is standard upstream review telling that we do not duplicate >>>>> drivers. Ever, unless there is serious exception needed. >>>> >>>> I wasn't trying to twist the narrative, just trying to come up with a >>>> path forward that isn't "no" or "improve existing driver", since >>>> neither of those gets us towards a future using common frameworks. >>>> >>>>> I asked why there should be an exception granted? Is the reason for >>>>> exception following: >>>>> "We want to adopt common framework" >>>>> ? >>>> >>>> Possibly? But I don't think we want two drivers to be any sort of >>>> long term solution. (Ie. as long as venus/iris have co-exist.) >>>> >>>>> >>>>>> rather reacting to drm/accel emerging as the standard framework for >>>>>> this sort of driver. >>>>>> >>>>>> So how do we get from here to there? >>>>> >>>>> What is wrong with my proposal? >>>> >>>> Maybe I missed something, my understanding was your proposal was >>>> "Grow/replace/improve existing driver instead of coming with a >>>> duplicate".. grow or improve doesn't move us toward common >>>> frameworks. Maybe "replace" is a valid option. If there is something >>>> I missed, then I apologize. >>>> >>>> Options I can think of are: >>>> >>>> 1. Hardware cutoff.. new hw gets new driver, existing hw gets existing >>>> driver >>>> 2. Backwards compat chardev registered by new driver, providing existing >>>> UABI. I'm not 100% sure about the feasibility/drawbacks of this.. >>>> AFAIU the fastrpc folks where planning a backwards compat layer in >>>> userspace, so maybe it is possible. >>>> 3. exception? >>>> >>>> I'd like to know what the feasibility of #2 is, since at a high level >>>> that sounds like the best option. Possibly limit exposure of legacy >>>> UABI to existing hw so we don't get into a place of needing to extend >>>> the legacy UABI for new hw? >>>> >>>> But #1 sounds like a non-controversial place to start regardless. >>>> Possibly with #2 coming as followup and necessary step before eventual >>>> migration to new driver for existing hw? >>>> >>>> Even if we start with #2, how do we handle first-merge-window >>>> bugs/regressions without reverting addition of new driver and removal >>>> of old? It seems like we'd need a window of a couple release cycles >>>> where both drivers exist? >>>> >>>> Maybe others have other/better options in mind? >>> To all, I'm seeking on the approach I should follow to go ahead here. I >>> can work on implementing #1(as per Rob's list) with hw specific >>> compatible for v4 if it's acceptable. >>> >> >> I don't see any reason for you to define a "hw specific compatible", >> because as you have shown in this series (and as Rob point out), there's >> no difference in the "hardware". >> >> The only reason for your "hw specific compatible" is to make a software >> selection in Linux - and that's not what DeviceTree is for. > > That's not exactly true. There are protocol differences. For example, > polling mode is supported only since a certain timeline in the history. > Likewise other interface features are not supported on all the FastRPC > devices. For the polling mode support we were already beaten by the lack > of SoC-specific compats. > >> As such, I don't see that you have a DeviceTree problem at all, because >> this is a Linux-internal problem. >> >>> #2(compat driver) is something that we are still exploring as we >>> couldn't find any standard way to achieve it. We might start a separate >>> discussion for that once we have few possible designs with us. >>> >> >> This is the actual problem! >> >> We have existing user space that depends on the ioctl interface exposed >> by the current misc driver. You must not break these. > > This is clear. > >> Hardware cutoff is not a viable solution, because that's just a >> declaration that we'll let the old platforms rotten - or alternatively >> you commit to maintain two drivers to the very same feature and quality >> level. >> >> So the only reasonable solution is #2; from there it's a valid question >> if you reach that point my stepwise migrating the current misc driver >> that solution, or if you present a new driver with the fully backwards >> compatible interface, alongside the new ABI. > > I think the general direction was #3 (or #2.1): implement a shim layer > on top of the QDA driver as a separate module. Put all the historical > over-complicated solutions into that shim module and let it die at some > point. Current fastrpc driver lets userspace specify buffers in several > different ways, forcing the kernel driver to perform a lot of work > with buffer addresses. I don't think that this legacy code should be a > part of the QDA. It can go to qda-fastrpc-shim, keeping it out of the > kernel memory if it's not necessary. > >> But this does bring to a question which the cover letter should explain >> - but doesn't: what problem does this patch series actually solve? > > I agree that it should be a part of the cover letter. > > As a person who triggered this work, I can propose my reasons: > > - Current driver has over-complicated memory manager (both on the > userspace and on the kernel side). Correspondng uAPI is not really > suitable for virtualization. Using GEM simplifies both the kernel > driver and uAPI. Also using handle-offset-length to specify the > buffers makes it easy to support virtual QDA devices. > > - Current driver predates the accel subsystem. Using common subsystem > simplifies reviews. The QDA driver has gotten several comments about > the usage of the DMA-BUFs. It'not unlikely that the same issues are > present in the current FastRPC driver, just being unnoticed. > > - The ideas present in the current FastRPC driver also predate the > current design practices. The uAPI was created in the ad-hoc way, just > following the momentary needs. Driver code also shows the result of > that, having enough of the spaghetti code. > > Given all of that, yes, it is possible to provide an evolution of the > FastRPC driver into the accel+shim, improve the code quality meanwhile > and end up with the good enough split. However I think that the path > taken would be longer and the net result might be worse. > > With all of that in mind, my suggestion is to continue working on the > QDA driver, get the core of it to integrate nicely with the accel and > DMA-BUF usage requirements and implement the (minimal?) fastrpc shim on > top of it. > Agreed on all the above points. We discussed this internally and are exploring a shim module(something like qda_fastrpc_shim.ko) along these lines, sitting on top of the QDA driver rather than growing fastrpc.c in place. Still working through some of the details (e.g. who owns the rpmsg binding, and what "minimal" means precisely given feature parity is expected for the compat layer's behaviour). Will follow up with a design doc here once that's settled. To all, one separate question on daemon-attach security for feature parity: would running the daemon with CAP_SYS_ADMIN, and having the driver reject attach requests from any process without that capability, be an acceptable approach? The concern is preventing a malicious/unprivileged application from attaching to a privileged DSP PD and disrupting it. Thanks for the detailed feedback. //Ekansh
On Wed, Sep 09, 2026 at 02:48:17PM +0300, Dmitry Baryshkov wrote: > On Tue, Sep 08, 2026 at 05:44:10PM -0500, Bjorn Andersson wrote: > > On Wed, Aug 26, 2026 at 06:37:00PM +0530, Ekansh Gupta wrote: > > > On 20-08-2026 20:17, Rob Clark wrote: > > > > On Wed, Aug 19, 2026 at 11:15 PM Krzysztof Kozlowski <krzk@kernel.org> wrote: > > > >> > > > >> On 19/08/2026 17:48, Rob Clark wrote: > > > >>> On Wed, Aug 19, 2026 at 8:27 AM Krzysztof Kozlowski <krzk@kernel.org> wrote: > > > >>>> The rule of usptream development is that we do not accept duplicated > > > >>>> code, just because a vendor wants to write something new. This is > > > >>>> basically the concept applied all over the drivers tree, where we pushed > > > >>>> back against all sorts of duplications all over the vendors. > > > >>>> > > > >>>> What I miss in this thread is why would there be any exception here. We > > > >>>> do not grant exceptions from standard practices on "I want" reasons. > > > >>> > > > >>> I agree that we should not have duplicated drivers just for vendor > > > >>> lolz. But when it comes to adopting common frameworks and integrating > > > >>> better into the ecosystem, this doesn't seem like something we should > > > >>> actively discourage. I don't think this is a case of vendor lolz, but > > > >> > > > >> No one discourages it. Following standard Linux kernel practices and > > > >> requirements is not discouraging, do not twist the narrative here. > > > >> Again, it is standard upstream review telling that we do not duplicate > > > >> drivers. Ever, unless there is serious exception needed. > > > > > > > > I wasn't trying to twist the narrative, just trying to come up with a > > > > path forward that isn't "no" or "improve existing driver", since > > > > neither of those gets us towards a future using common frameworks. > > > > > > > >> I asked why there should be an exception granted? Is the reason for > > > >> exception following: > > > >> "We want to adopt common framework" > > > >> ? > > > > > > > > Possibly? But I don't think we want two drivers to be any sort of > > > > long term solution. (Ie. as long as venus/iris have co-exist.) > > > > > > > >> > > > >>> rather reacting to drm/accel emerging as the standard framework for > > > >>> this sort of driver. > > > >>> > > > >>> So how do we get from here to there? > > > >> > > > >> What is wrong with my proposal? > > > > > > > > Maybe I missed something, my understanding was your proposal was > > > > "Grow/replace/improve existing driver instead of coming with a > > > > duplicate".. grow or improve doesn't move us toward common > > > > frameworks. Maybe "replace" is a valid option. If there is something > > > > I missed, then I apologize. > > > > > > > > Options I can think of are: > > > > > > > > 1. Hardware cutoff.. new hw gets new driver, existing hw gets existing > > > > driver > > > > 2. Backwards compat chardev registered by new driver, providing existing > > > > UABI. I'm not 100% sure about the feasibility/drawbacks of this.. > > > > AFAIU the fastrpc folks where planning a backwards compat layer in > > > > userspace, so maybe it is possible. > > > > 3. exception? > > > > > > > > I'd like to know what the feasibility of #2 is, since at a high level > > > > that sounds like the best option. Possibly limit exposure of legacy > > > > UABI to existing hw so we don't get into a place of needing to extend > > > > the legacy UABI for new hw? > > > > > > > > But #1 sounds like a non-controversial place to start regardless. > > > > Possibly with #2 coming as followup and necessary step before eventual > > > > migration to new driver for existing hw? > > > > > > > > Even if we start with #2, how do we handle first-merge-window > > > > bugs/regressions without reverting addition of new driver and removal > > > > of old? It seems like we'd need a window of a couple release cycles > > > > where both drivers exist? > > > > > > > > Maybe others have other/better options in mind? > > > To all, I'm seeking on the approach I should follow to go ahead here. I > > > can work on implementing #1(as per Rob's list) with hw specific > > > compatible for v4 if it's acceptable. > > > > > > > I don't see any reason for you to define a "hw specific compatible", > > because as you have shown in this series (and as Rob point out), there's > > no difference in the "hardware". > > > > The only reason for your "hw specific compatible" is to make a software > > selection in Linux - and that's not what DeviceTree is for. > > That's not exactly true. There are protocol differences. For example, > polling mode is supported only since a certain timeline in the history. > Likewise other interface features are not supported on all the FastRPC > devices. For the polling mode support we were already beaten by the lack > of SoC-specific compats. > I can see the benefit of capturing some of the generational features in a compatible, like the changes related to address width. But for pure software features that doesn't have an actual bearing in the hardware, I'd prefer if we relied on dynamic discovery. But none of that applies to the question of "can I use compatible to select if we should use the new or old Linux driver". > > As such, I don't see that you have a DeviceTree problem at all, because > > this is a Linux-internal problem. > > > > > #2(compat driver) is something that we are still exploring as we > > > couldn't find any standard way to achieve it. We might start a separate > > > discussion for that once we have few possible designs with us. > > > > > > > This is the actual problem! > > > > We have existing user space that depends on the ioctl interface exposed > > by the current misc driver. You must not break these. > > This is clear. > > > Hardware cutoff is not a viable solution, because that's just a > > declaration that we'll let the old platforms rotten - or alternatively > > you commit to maintain two drivers to the very same feature and quality > > level. > > > > So the only reasonable solution is #2; from there it's a valid question > > if you reach that point my stepwise migrating the current misc driver > > that solution, or if you present a new driver with the fully backwards > > compatible interface, alongside the new ABI. > > I think the general direction was #3 (or #2.1): implement a shim layer > on top of the QDA driver as a separate module. Put all the historical > over-complicated solutions into that shim module and let it die at some > point. Current fastrpc driver lets userspace specify buffers in several > different ways, forcing the kernel driver to perform a lot of work > with buffer addresses. I don't think that this legacy code should be a > part of the QDA. It can go to qda-fastrpc-shim, keeping it out of the > kernel memory if it's not necessary. > > > But this does bring to a question which the cover letter should explain > > - but doesn't: what problem does this patch series actually solve? > > I agree that it should be a part of the cover letter. > > As a person who triggered this work, I can propose my reasons: > > - Current driver has over-complicated memory manager (both on the > userspace and on the kernel side). The userspace library is spaghetti, but when I wrote my own it turned out quite succinct. The ioctl structs certainly could have been cleaner, and documented, but it seems to me that a fair amount of the complexity comes from the different use cases - such as SMMU vs XPU, secure and unsecure buffers, remnants of now unsupported options. I'm presuming that QDA will need to adopt most of these, and that QDA support will be bolted into the spaghetti library. > Correspondng uAPI is not really > suitable for virtualization. Using GEM simplifies both the kernel > driver and uAPI. Also using handle-offset-length to specify the > buffers makes it easy to support virtual QDA devices. I'm looking forward to learn more about this! > > - Current driver predates the accel subsystem. Using common subsystem > simplifies reviews. The QDA driver has gotten several comments about > the usage of the DMA-BUFs. It'not unlikely that the same issues are > present in the current FastRPC driver, just being unnoticed. Yeah, this is unfortunate. It would certainly be nice to have a documented and clean ABI. > > - The ideas present in the current FastRPC driver also predate the > current design practices. The uAPI was created in the ad-hoc way, just > following the momentary needs. Driver code also shows the result of > that, having enough of the spaghetti code. > Yeah, again, this isn't desirable. > Given all of that, yes, it is possible to provide an evolution of the > FastRPC driver into the accel+shim, improve the code quality meanwhile > and end up with the good enough split. However I think that the path > taken would be longer and the net result might be worse. > > With all of that in mind, my suggestion is to continue working on the > QDA driver, get the core of it to integrate nicely with the accel and > DMA-BUF usage requirements and implement the (minimal?) fastrpc shim on > top of it. > If you believe this is the way to reach the proper design, then I won't object. My requirement is that my userspace continues to work when my distro suddenly switches FASTRPC=n/QDA=m. I have no problems with dropping the fastrpc driver once the QDA is drop-in-compatible. I'm also open to marking the compat layer deprecated and eventually drop it once we're certain that users have moved to a userspace that used the accel interface. But we can't merge the QDA driver as long as we believe that implementing a compat layer will be hard/impossible. Regards, Bjorn > > -- > With best wishes > Dmitry
On Thu, Sep 10, 2026 at 7:25 AM Bjorn Andersson <andersson@kernel.org> wrote: > > On Wed, Sep 09, 2026 at 02:48:17PM +0300, Dmitry Baryshkov wrote: > > On Tue, Sep 08, 2026 at 05:44:10PM -0500, Bjorn Andersson wrote: > > > On Wed, Aug 26, 2026 at 06:37:00PM +0530, Ekansh Gupta wrote: > > > > On 20-08-2026 20:17, Rob Clark wrote: > > > > > On Wed, Aug 19, 2026 at 11:15 PM Krzysztof Kozlowski <krzk@kernel.org> wrote: > > > > >> > > > > >> On 19/08/2026 17:48, Rob Clark wrote: > > > > >>> On Wed, Aug 19, 2026 at 8:27 AM Krzysztof Kozlowski <krzk@kernel.org> wrote: > > > > >>>> The rule of usptream development is that we do not accept duplicated > > > > >>>> code, just because a vendor wants to write something new. This is > > > > >>>> basically the concept applied all over the drivers tree, where we pushed > > > > >>>> back against all sorts of duplications all over the vendors. > > > > >>>> > > > > >>>> What I miss in this thread is why would there be any exception here. We > > > > >>>> do not grant exceptions from standard practices on "I want" reasons. > > > > >>> > > > > >>> I agree that we should not have duplicated drivers just for vendor > > > > >>> lolz. But when it comes to adopting common frameworks and integrating > > > > >>> better into the ecosystem, this doesn't seem like something we should > > > > >>> actively discourage. I don't think this is a case of vendor lolz, but > > > > >> > > > > >> No one discourages it. Following standard Linux kernel practices and > > > > >> requirements is not discouraging, do not twist the narrative here. > > > > >> Again, it is standard upstream review telling that we do not duplicate > > > > >> drivers. Ever, unless there is serious exception needed. > > > > > > > > > > I wasn't trying to twist the narrative, just trying to come up with a > > > > > path forward that isn't "no" or "improve existing driver", since > > > > > neither of those gets us towards a future using common frameworks. > > > > > > > > > >> I asked why there should be an exception granted? Is the reason for > > > > >> exception following: > > > > >> "We want to adopt common framework" > > > > >> ? > > > > > > > > > > Possibly? But I don't think we want two drivers to be any sort of > > > > > long term solution. (Ie. as long as venus/iris have co-exist.) > > > > > > > > > >> > > > > >>> rather reacting to drm/accel emerging as the standard framework for > > > > >>> this sort of driver. > > > > >>> > > > > >>> So how do we get from here to there? > > > > >> > > > > >> What is wrong with my proposal? > > > > > > > > > > Maybe I missed something, my understanding was your proposal was > > > > > "Grow/replace/improve existing driver instead of coming with a > > > > > duplicate".. grow or improve doesn't move us toward common > > > > > frameworks. Maybe "replace" is a valid option. If there is something > > > > > I missed, then I apologize. > > > > > > > > > > Options I can think of are: > > > > > > > > > > 1. Hardware cutoff.. new hw gets new driver, existing hw gets existing > > > > > driver > > > > > 2. Backwards compat chardev registered by new driver, providing existing > > > > > UABI. I'm not 100% sure about the feasibility/drawbacks of this.. > > > > > AFAIU the fastrpc folks where planning a backwards compat layer in > > > > > userspace, so maybe it is possible. > > > > > 3. exception? > > > > > > > > > > I'd like to know what the feasibility of #2 is, since at a high level > > > > > that sounds like the best option. Possibly limit exposure of legacy > > > > > UABI to existing hw so we don't get into a place of needing to extend > > > > > the legacy UABI for new hw? > > > > > > > > > > But #1 sounds like a non-controversial place to start regardless. > > > > > Possibly with #2 coming as followup and necessary step before eventual > > > > > migration to new driver for existing hw? > > > > > > > > > > Even if we start with #2, how do we handle first-merge-window > > > > > bugs/regressions without reverting addition of new driver and removal > > > > > of old? It seems like we'd need a window of a couple release cycles > > > > > where both drivers exist? > > > > > > > > > > Maybe others have other/better options in mind? > > > > To all, I'm seeking on the approach I should follow to go ahead here. I > > > > can work on implementing #1(as per Rob's list) with hw specific > > > > compatible for v4 if it's acceptable. > > > > > > > > > > I don't see any reason for you to define a "hw specific compatible", > > > because as you have shown in this series (and as Rob point out), there's > > > no difference in the "hardware". > > > > > > The only reason for your "hw specific compatible" is to make a software > > > selection in Linux - and that's not what DeviceTree is for. > > > > That's not exactly true. There are protocol differences. For example, > > polling mode is supported only since a certain timeline in the history. > > Likewise other interface features are not supported on all the FastRPC > > devices. For the polling mode support we were already beaten by the lack > > of SoC-specific compats. > > > > I can see the benefit of capturing some of the generational features in > a compatible, like the changes related to address width. But for pure > software features that doesn't have an actual bearing in the hardware, > I'd prefer if we relied on dynamic discovery. > > But none of that applies to the question of "can I use compatible to > select if we should use the new or old Linux driver". > Nit, driver could still have an allow/deny-list of machine compatibles. Might be something to keep in mind if a phased depreciation strategy was desirable.. BR, -R
On 09-09-2026 04:14, Bjorn Andersson wrote: > On Wed, Aug 26, 2026 at 06:37:00PM +0530, Ekansh Gupta wrote: >> On 20-08-2026 20:17, Rob Clark wrote: >>> On Wed, Aug 19, 2026 at 11:15 PM Krzysztof Kozlowski <krzk@kernel.org> wrote: >>>> >>>> On 19/08/2026 17:48, Rob Clark wrote: >>>>> On Wed, Aug 19, 2026 at 8:27 AM Krzysztof Kozlowski <krzk@kernel.org> wrote: >>>>>> The rule of usptream development is that we do not accept duplicated >>>>>> code, just because a vendor wants to write something new. This is >>>>>> basically the concept applied all over the drivers tree, where we pushed >>>>>> back against all sorts of duplications all over the vendors. >>>>>> >>>>>> What I miss in this thread is why would there be any exception here. We >>>>>> do not grant exceptions from standard practices on "I want" reasons. >>>>> >>>>> I agree that we should not have duplicated drivers just for vendor >>>>> lolz. But when it comes to adopting common frameworks and integrating >>>>> better into the ecosystem, this doesn't seem like something we should >>>>> actively discourage. I don't think this is a case of vendor lolz, but >>>> >>>> No one discourages it. Following standard Linux kernel practices and >>>> requirements is not discouraging, do not twist the narrative here. >>>> Again, it is standard upstream review telling that we do not duplicate >>>> drivers. Ever, unless there is serious exception needed. >>> >>> I wasn't trying to twist the narrative, just trying to come up with a >>> path forward that isn't "no" or "improve existing driver", since >>> neither of those gets us towards a future using common frameworks. >>> >>>> I asked why there should be an exception granted? Is the reason for >>>> exception following: >>>> "We want to adopt common framework" >>>> ? >>> >>> Possibly? But I don't think we want two drivers to be any sort of >>> long term solution. (Ie. as long as venus/iris have co-exist.) >>> >>>> >>>>> rather reacting to drm/accel emerging as the standard framework for >>>>> this sort of driver. >>>>> >>>>> So how do we get from here to there? >>>> >>>> What is wrong with my proposal? >>> >>> Maybe I missed something, my understanding was your proposal was >>> "Grow/replace/improve existing driver instead of coming with a >>> duplicate".. grow or improve doesn't move us toward common >>> frameworks. Maybe "replace" is a valid option. If there is something >>> I missed, then I apologize. >>> >>> Options I can think of are: >>> >>> 1. Hardware cutoff.. new hw gets new driver, existing hw gets existing >>> driver >>> 2. Backwards compat chardev registered by new driver, providing existing >>> UABI. I'm not 100% sure about the feasibility/drawbacks of this.. >>> AFAIU the fastrpc folks where planning a backwards compat layer in >>> userspace, so maybe it is possible. >>> 3. exception? >>> >>> I'd like to know what the feasibility of #2 is, since at a high level >>> that sounds like the best option. Possibly limit exposure of legacy >>> UABI to existing hw so we don't get into a place of needing to extend >>> the legacy UABI for new hw? >>> >>> But #1 sounds like a non-controversial place to start regardless. >>> Possibly with #2 coming as followup and necessary step before eventual >>> migration to new driver for existing hw? >>> >>> Even if we start with #2, how do we handle first-merge-window >>> bugs/regressions without reverting addition of new driver and removal >>> of old? It seems like we'd need a window of a couple release cycles >>> where both drivers exist? >>> >>> Maybe others have other/better options in mind? >> To all, I'm seeking on the approach I should follow to go ahead here. I >> can work on implementing #1(as per Rob's list) with hw specific >> compatible for v4 if it's acceptable. >> > > I don't see any reason for you to define a "hw specific compatible", > because as you have shown in this series (and as Rob point out), there's > no difference in the "hardware". > > The only reason for your "hw specific compatible" is to make a software > selection in Linux - and that's not what DeviceTree is for. > > > As such, I don't see that you have a DeviceTree problem at all, because > this is a Linux-internal problem. Agreed. > >> #2(compat driver) is something that we are still exploring as we >> couldn't find any standard way to achieve it. We might start a separate >> discussion for that once we have few possible designs with us. >> > > This is the actual problem! > > We have existing user space that depends on the ioctl interface exposed > by the current misc driver. You must not break these. > > Hardware cutoff is not a viable solution, because that's just a > declaration that we'll let the old platforms rotten - or alternatively > you commit to maintain two drivers to the very same feature and quality > level. > > So the only reasonable solution is #2; from there it's a valid question > if you reach that point my stepwise migrating the current misc driver > that solution, or if you present a new driver with the fully backwards > compatible interface, alongside the new ABI. > > > But this does bring to a question which the cover letter should explain > - but doesn't: what problem does this patch series actually solve? > > Regards, > Bjorn Agreed. I'll target #2: QDA implementing the existing fastrpc UABI alongside the new one, rather than a driver split by platform. On how to get there: the blocker we hit in v1 was that legacy fastrpc buffer semantics appeared to need a drm_file, and there's no exported way to construct one outside the DRM core. I want to re-examine that constraint rather than treat it as final, since the legacy interface's own buffer model (a dma_buf fd as the buffer identity, no GEM involved) is not inherently tied to drm_file, that's how the existing misc driver implements it today. I don't have a concrete design yet and would rather work through it here than commit to one prematurely. If anyone has thoughts on how the legacy UABI could be served without requiring a drm_file per session or if I can somehow bind drm_file with chardev by exposing some APIs from DRM core, I'd welcome them. On the cover letter: fair point, and I'll fix it. The problem this series solves is that a miscdevice interface requires us to hand-roll what the accel/DRM subsystem already provides as common infrastructure: GEM for buffer lifecycle and reference counting, PRIME for cross-driver import/export, per-file (per-open) context and handle-namespace isolation, and the existing debug and lifecycle tooling the DRM core already ships. Every accelerator driver added to drivers/accel (habanalabs, ivpu, qaic, rocket) has taken this path for the same reason, rather than each maintaining its own equivalent inside drivers/misc. Building QDA directly on this shared infrastructure, instead of extending fastrpc's own ad hoc buffer and session tracking to cover the same ground, avoids that duplication going forward. I'll put this plainly in the cover letter rather than leaving it implicit. /Ekansh > >> Happy to take any other suggestion also.> >>> BR, >>> -R >> >>
On 9/9/26 8:53 AM, Ekansh Gupta wrote: >> This is the actual problem! >> >> We have existing user space that depends on the ioctl interface exposed >> by the current misc driver. You must not break these. >> >> Hardware cutoff is not a viable solution, because that's just a >> declaration that we'll let the old platforms rotten - or alternatively >> you commit to maintain two drivers to the very same feature and quality >> level. >> >> So the only reasonable solution is #2; from there it's a valid question >> if you reach that point my stepwise migrating the current misc driver >> that solution, or if you present a new driver with the fully backwards >> compatible interface, alongside the new ABI. >> >> >> But this does bring to a question which the cover letter should explain >> - but doesn't: what problem does this patch series actually solve? >> >> Regards, >> Bjorn > Agreed. I'll target #2: QDA implementing the existing fastrpc UABI > alongside the new one, rather than a driver split by platform. > > On how to get there: the blocker we hit in v1 was that legacy fastrpc > buffer semantics appeared to need a drm_file, and there's no exported > way to construct one outside the DRM core. I want to re-examine that > constraint rather than treat it as final, since the legacy interface's > own buffer model (a dma_buf fd as the buffer identity, no GEM > involved) is not inherently tied to drm_file, that's how the existing > misc driver implements it today. I don't have a concrete design yet > and would rather work through it here than commit to one prematurely. > If anyone has thoughts on how the legacy UABI could be served without > requiring a drm_file per session or if I can somehow bind drm_file with > chardev by exposing some APIs from DRM core, I'd welcome them. > > On the cover letter: fair point, and I'll fix it. The problem this > series solves is that a miscdevice interface requires us to hand-roll > what the accel/DRM subsystem already provides as common > infrastructure: GEM for buffer lifecycle and reference counting, > PRIME for cross-driver import/export, per-file (per-open) context and > handle-namespace isolation, and the existing debug and lifecycle > tooling the DRM core already ships. Every accelerator driver added to > drivers/accel (habanalabs, ivpu, qaic, rocket) has taken this path for > the same reason, rather than each maintaining its own equivalent > inside drivers/misc. Building QDA directly on this shared > infrastructure, instead of extending fastrpc's own ad hoc buffer and > session tracking to cover the same ground, avoids that duplication Am sure you must have already tried this, but Can you not migrate existing fastrpc driver to this shared infrastructure under the hood? Can you elaborate on what are the blockers you hit in doing so? This will ensure that UAPI is retained and still get benefit of QDA. --srini > going forward.
On 09-09-2026 13:32, Srinivas Kandagatla wrote: > On 9/9/26 8:53 AM, Ekansh Gupta wrote: >>> This is the actual problem! >>> >>> We have existing user space that depends on the ioctl interface exposed >>> by the current misc driver. You must not break these. >>> >>> Hardware cutoff is not a viable solution, because that's just a >>> declaration that we'll let the old platforms rotten - or alternatively >>> you commit to maintain two drivers to the very same feature and quality >>> level. >>> >>> So the only reasonable solution is #2; from there it's a valid question >>> if you reach that point my stepwise migrating the current misc driver >>> that solution, or if you present a new driver with the fully backwards >>> compatible interface, alongside the new ABI. >>> >>> >>> But this does bring to a question which the cover letter should explain >>> - but doesn't: what problem does this patch series actually solve? >>> >>> Regards, >>> Bjorn >> Agreed. I'll target #2: QDA implementing the existing fastrpc UABI >> alongside the new one, rather than a driver split by platform. >> >> On how to get there: the blocker we hit in v1 was that legacy fastrpc >> buffer semantics appeared to need a drm_file, and there's no exported >> way to construct one outside the DRM core. I want to re-examine that >> constraint rather than treat it as final, since the legacy interface's >> own buffer model (a dma_buf fd as the buffer identity, no GEM >> involved) is not inherently tied to drm_file, that's how the existing >> misc driver implements it today. I don't have a concrete design yet >> and would rather work through it here than commit to one prematurely. >> If anyone has thoughts on how the legacy UABI could be served without >> requiring a drm_file per session or if I can somehow bind drm_file with >> chardev by exposing some APIs from DRM core, I'd welcome them. >> >> On the cover letter: fair point, and I'll fix it. The problem this >> series solves is that a miscdevice interface requires us to hand-roll >> what the accel/DRM subsystem already provides as common >> infrastructure: GEM for buffer lifecycle and reference counting, >> PRIME for cross-driver import/export, per-file (per-open) context and >> handle-namespace isolation, and the existing debug and lifecycle >> tooling the DRM core already ships. Every accelerator driver added to >> drivers/accel (habanalabs, ivpu, qaic, rocket) has taken this path for >> the same reason, rather than each maintaining its own equivalent >> inside drivers/misc. Building QDA directly on this shared >> infrastructure, instead of extending fastrpc's own ad hoc buffer and >> session tracking to cover the same ground, avoids that duplication > Am sure you must have already tried this, but Can you not migrate > existing fastrpc driver to this shared infrastructure under the hood? > > Can you elaborate on what are the blockers you hit in doing so? > > This will ensure that UAPI is retained and still get benefit of QDA. > QDA's session model is built around drm_file, and fastrpc's chardev has no drm_file associated with it. I'm still checking whether that dependency is unavoidable for fastrpc's case, or whether there's a way to use the underlying buffer-management infrastructure without it. Please correct me if you are asking/suggesting a different approach? > --srini >> going forward. >
On 8/19/26 4:38 PM, Rob Clark wrote: > On Wed, Aug 19, 2026 at 7:21 AM Krzysztof Kozlowski <krzk@kernel.org> wrote: >> >> On 19/08/2026 15:26, Ekansh Gupta wrote: >>> On 19-08-2026 00:43, Krzysztof Kozlowski wrote: >>>> On 17/08/2026 06:47, Ekansh Gupta wrote: >>>>> This patch series introduces the Qualcomm DSP Accelerator (QDA) driver, >>>>> a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a >>>>> standardized interface for offloading computational tasks to DSPs found >>>>> on Qualcomm SoCs, supporting all DSP domains. >>>>> >>>>> The QDA driver implements the FastRPC protocol over the DRM accel >>>>> subsystem. It uses the same device-tree node structure as the existing >>>>> fastrpc driver in drivers/misc/. The approach for binding the QDA driver >>>>> to device-tree nodes while coexisting with the fastrpc driver is an open >>>>> item described below. >>>> >>>> No. Grow/replace/improve existing driver instead of coming with a duplicate. >>>> >>>> That's a standard upstream requirement, basically given on every >>>> upstreaming guide. >>>> >>>> Please watch old talk from Greg - "I Don’t Want Your Code!". >>> Posted discussion threads here[1]. Would seek comments from Dmitry, >>> Srini as well. >>> >>> [1] >>> https://lore.kernel.org/all/3476b5c3-7983-4994-a901-3d7d8bd75255@oss.qualcomm.com/ >> >> The rest of the comments is still valid even if you did not acknowledge >> them. >> >> Anyway, regarding above - again, watch the talk from Greg. >> >> You have ONE driver. Not two. > > Long term, moving to the common driver framework (which did not exist > when fastrpc was first created) seems like a good thing. But does > that not allow for some transition period? How can we get from here > to there without otherwise breaking userspace? Is there some other > precedent elsewhere in other driver subsystems? > > I suppose drm exposing legacy fbdev on top of drm drivers is _sort of_ > a precedent if you squint a bit? I'm not really familiar enough to > say if that would be reasonable/possible in this case. Also, venus/iris, mdp5/dpu1 Konrad
On 18/08/2026 21:13, Krzysztof Kozlowski wrote: > On 17/08/2026 06:47, Ekansh Gupta wrote: >> This patch series introduces the Qualcomm DSP Accelerator (QDA) driver, >> a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a >> standardized interface for offloading computational tasks to DSPs found >> on Qualcomm SoCs, supporting all DSP domains. >> >> The QDA driver implements the FastRPC protocol over the DRM accel >> subsystem. It uses the same device-tree node structure as the existing >> fastrpc driver in drivers/misc/. The approach for binding the QDA driver >> to device-tree nodes while coexisting with the fastrpc driver is an open >> item described below. > > No. Grow/replace/improve existing driver instead of coming with a duplicate. > > That's a standard upstream requirement, basically given on every > upstreaming guide. > > Please watch old talk from Greg - "I Don’t Want Your Code!". > >> >> v1: https://lore.kernel.org/all/20260519-qda-series-v1-0-b2d984c297f8@oss.qualcomm.com/ >> RFC: https://lore.kernel.org/dri-devel/20260224-qda-firstpost-v1-0-fe46a9c1a046@oss.qualcomm.com/T/ >> >> Changes since v1 >> ================ >> >> The v1 review raised two architectural objections and one correctness >> issue; all three are resolved in v2: >> >> * Christian König (dma-buf maintainer) pointed out that the imported- >> buffer path silently assumed the IOMMU maps every buffer as a single >> contiguous range, which is not guaranteed. v2 walks the scatterlist >> and cleanly rejects non-contiguous imports; contiguous imports (e.g. >> CMA DMA-buf heap) are accepted. (patch 11) >> >> * Dmitry Baryshkov objected to three different buffer-passing formats >> in the invoke IOCTL (DMA-BUF fd, direct/inline, DMA handle). v2 >> passes only GEM handles; userspace imports any fd to a GEM handle >> with DRM_IOCTL_PRIME_FD_TO_HANDLE before invoking. Packing and >> overlap handling are left to userspace. (patch 12) >> >> * The memory manager (patch 07) used a fixed 16-entry array without >> justification and leaked the device descriptor on teardown. v2 >> allocates the array from the DT context-bank count (as Dmitry >> suggested) and frees it correctly. >> >> User-space staging branch >> ========================= >> https://github.com/qualcomm/fastrpc/tree/accel/staging >> >> Key Features >> ============ >> >> * Standard DRM accelerator interface via /dev/accel/accelN >> * GEM-based buffer management with DMA-BUF import (PRIME) >> * IOMMU-based memory isolation using per-process context banks >> * FastRPC protocol implementation for DSP communication >> * RPMsg transport layer for reliable message passing >> * Support for all DSP domains (ADSP, CDSP, SDSP, GDSP) >> * DRM IOCTL interface for DSP session management, buffer allocation, >> and remote procedure invocation >> >> Architecture >> ============ >> >> 1. DRM Accelerator Framework Integration >> The driver registers as a DRM accel device, exposing a standard >> /dev/accel/accelN character device node. This provides established >> DRM infrastructure for device management, file operations, and >> IOCTL dispatch. >> >> 2. Memory Management >> Buffers are managed as GEM objects with PRIME support for DMA-BUF >> import. This enables buffer sharing with other DRM drivers (GPU, >> camera, video) using standard kernel mechanisms. Only contiguous >> imports are accepted; the driver verifies contiguity at import time >> rather than assuming it. >> >> 3. IOMMU Context Bank Management >> IOMMU context banks (CBs) are represented as proper struct device >> instances on a custom virtual bus (qda-compute-cb). Each CB device >> is registered with the IOMMU subsystem and receives its own IOMMU >> domain, enabling per-session address space isolation. The custom >> bus was introduced because IOMMU context banks are synthetic >> constructs — not real platform devices — and to ensure CB device >> lifetime is strictly subordinate to the parent QDA device. >> See also: https://lore.kernel.org/all/245d602f-3037-4ae3-9af9-d98f37258aae@oss.qualcomm.com/ >> >> 4. Memory Manager Architecture >> The memory manager maintains a registry of IOMMU devices in an >> array sized to the number of context banks described in the device >> tree, and coordinates per-process device assignment with reference- >> counted lifetime management. The DMA-coherent backend allocates >> buffers with SID-prefixed DMA addresses for DSP firmware >> compatibility. >> >> 5. Transport Layer >> RPMsg communication is handled in a dedicated transport layer >> (qda_rpmsg.c), separate from the core DRM driver logic. >> >> 6. Code Organization >> The driver is organized across multiple files (~4800 lines total): >> * qda_drv.c: Core driver and DRM integration >> * qda_rpmsg.c: RPMsg transport layer >> * qda_cb.c: Context bank device management >> * qda_compute_bus.c: Custom virtual bus for CB devices >> * qda_gem.c: GEM object management >> * qda_prime.c: DMA-BUF import (PRIME) >> * qda_memory_manager.c: IOMMU device registry and allocation >> * qda_memory_dma.c: DMA-coherent allocation backend >> * qda_fastrpc.c: FastRPC protocol implementation >> * qda_ioctl.c: IOCTL dispatch >> >> 7. UAPI Design >> The driver exposes DRM-style IOCTLs defined in >> include/uapi/drm/qda_accel.h, following DRM UAPI conventions >> (__u32/__u64 types, C++ guard, GPL-2.0-only WITH Linux-syscall-note). >> Buffer arguments are identified by GEM handles; the driver never >> accepts DMA-BUF fds directly in any IOCTL. >> >> Patch Series Organization >> ========================== >> >> Patch 01: MAINTAINERS entry >> Patch 02: Driver documentation (Documentation/accel/qda/) >> Patches 03-04: Core driver skeleton and compute bus >> Patch 05: iommu: Register qda-compute-cb bus with IOMMU subsystem >> Patches 06-07: CB device enumeration and memory manager >> Patch 08: QUERY IOCTL and UAPI header >> Patches 09-11: GEM buffer management and PRIME import >> Patches 12-15: FastRPC protocol (invoke, session create/release, >> map/unmap) >> >> Open Items >> =========== >> >> 1. Device-Tree Compatible String >> The QDA driver uses the same device-tree node structure and >> properties as the existing fastrpc driver in drivers/misc/. A >> mechanism is needed to allow the QDA driver to bind to its device >> node independently of the fastrpc driver. >> >> The intended coexistence model is: platforms that require the >> complete fastrpc feature set continue to use "qcom,fastrpc"; new >> platforms where QDA's feature set is sufficient use a QDA-specific >> compatible string. New feature development is directed toward QDA. >> >> The options under consideration are: >> >> a) Add a new "qcom,qda" compatible string to the existing >> qcom,fastrpc.yaml binding, since the DT node structure and >> properties are identical. > No > >> >> b) Introduce a separate qcom,qda.yaml binding that references or >> inherits the fastrpc binding properties. > > No > >> >> Seeking guidance from DT binding maintainers on the preferred >> approach. > > Grow existing driver. You do not get new driver, you do not get new > bindings. > And this was already questioned at v1 (the true v1, not v1+1) but you ignored the comment. Great, so here goes away trust. NAK Best regards, Krzysztof
On 19-08-2026 00:51, Krzysztof Kozlowski wrote: > On 18/08/2026 21:13, Krzysztof Kozlowski wrote: >> On 17/08/2026 06:47, Ekansh Gupta wrote: >>> This patch series introduces the Qualcomm DSP Accelerator (QDA) driver, >>> a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a >>> standardized interface for offloading computational tasks to DSPs found >>> on Qualcomm SoCs, supporting all DSP domains. >>> >>> The QDA driver implements the FastRPC protocol over the DRM accel >>> subsystem. It uses the same device-tree node structure as the existing >>> fastrpc driver in drivers/misc/. The approach for binding the QDA driver >>> to device-tree nodes while coexisting with the fastrpc driver is an open >>> item described below. >> >> No. Grow/replace/improve existing driver instead of coming with a duplicate. >> >> That's a standard upstream requirement, basically given on every >> upstreaming guide. >> >> Please watch old talk from Greg - "I Don’t Want Your Code!". >> >>> >>> v1: https://lore.kernel.org/all/20260519-qda-series-v1-0-b2d984c297f8@oss.qualcomm.com/ >>> RFC: https://lore.kernel.org/dri-devel/20260224-qda-firstpost-v1-0-fe46a9c1a046@oss.qualcomm.com/T/ >>> >>> Changes since v1 >>> ================ >>> >>> The v1 review raised two architectural objections and one correctness >>> issue; all three are resolved in v2: >>> >>> * Christian König (dma-buf maintainer) pointed out that the imported- >>> buffer path silently assumed the IOMMU maps every buffer as a single >>> contiguous range, which is not guaranteed. v2 walks the scatterlist >>> and cleanly rejects non-contiguous imports; contiguous imports (e.g. >>> CMA DMA-buf heap) are accepted. (patch 11) >>> >>> * Dmitry Baryshkov objected to three different buffer-passing formats >>> in the invoke IOCTL (DMA-BUF fd, direct/inline, DMA handle). v2 >>> passes only GEM handles; userspace imports any fd to a GEM handle >>> with DRM_IOCTL_PRIME_FD_TO_HANDLE before invoking. Packing and >>> overlap handling are left to userspace. (patch 12) >>> >>> * The memory manager (patch 07) used a fixed 16-entry array without >>> justification and leaked the device descriptor on teardown. v2 >>> allocates the array from the DT context-bank count (as Dmitry >>> suggested) and frees it correctly. >>> >>> User-space staging branch >>> ========================= >>> https://github.com/qualcomm/fastrpc/tree/accel/staging >>> >>> Key Features >>> ============ >>> >>> * Standard DRM accelerator interface via /dev/accel/accelN >>> * GEM-based buffer management with DMA-BUF import (PRIME) >>> * IOMMU-based memory isolation using per-process context banks >>> * FastRPC protocol implementation for DSP communication >>> * RPMsg transport layer for reliable message passing >>> * Support for all DSP domains (ADSP, CDSP, SDSP, GDSP) >>> * DRM IOCTL interface for DSP session management, buffer allocation, >>> and remote procedure invocation >>> >>> Architecture >>> ============ >>> >>> 1. DRM Accelerator Framework Integration >>> The driver registers as a DRM accel device, exposing a standard >>> /dev/accel/accelN character device node. This provides established >>> DRM infrastructure for device management, file operations, and >>> IOCTL dispatch. >>> >>> 2. Memory Management >>> Buffers are managed as GEM objects with PRIME support for DMA-BUF >>> import. This enables buffer sharing with other DRM drivers (GPU, >>> camera, video) using standard kernel mechanisms. Only contiguous >>> imports are accepted; the driver verifies contiguity at import time >>> rather than assuming it. >>> >>> 3. IOMMU Context Bank Management >>> IOMMU context banks (CBs) are represented as proper struct device >>> instances on a custom virtual bus (qda-compute-cb). Each CB device >>> is registered with the IOMMU subsystem and receives its own IOMMU >>> domain, enabling per-session address space isolation. The custom >>> bus was introduced because IOMMU context banks are synthetic >>> constructs — not real platform devices — and to ensure CB device >>> lifetime is strictly subordinate to the parent QDA device. >>> See also: https://lore.kernel.org/all/245d602f-3037-4ae3-9af9-d98f37258aae@oss.qualcomm.com/ >>> >>> 4. Memory Manager Architecture >>> The memory manager maintains a registry of IOMMU devices in an >>> array sized to the number of context banks described in the device >>> tree, and coordinates per-process device assignment with reference- >>> counted lifetime management. The DMA-coherent backend allocates >>> buffers with SID-prefixed DMA addresses for DSP firmware >>> compatibility. >>> >>> 5. Transport Layer >>> RPMsg communication is handled in a dedicated transport layer >>> (qda_rpmsg.c), separate from the core DRM driver logic. >>> >>> 6. Code Organization >>> The driver is organized across multiple files (~4800 lines total): >>> * qda_drv.c: Core driver and DRM integration >>> * qda_rpmsg.c: RPMsg transport layer >>> * qda_cb.c: Context bank device management >>> * qda_compute_bus.c: Custom virtual bus for CB devices >>> * qda_gem.c: GEM object management >>> * qda_prime.c: DMA-BUF import (PRIME) >>> * qda_memory_manager.c: IOMMU device registry and allocation >>> * qda_memory_dma.c: DMA-coherent allocation backend >>> * qda_fastrpc.c: FastRPC protocol implementation >>> * qda_ioctl.c: IOCTL dispatch >>> >>> 7. UAPI Design >>> The driver exposes DRM-style IOCTLs defined in >>> include/uapi/drm/qda_accel.h, following DRM UAPI conventions >>> (__u32/__u64 types, C++ guard, GPL-2.0-only WITH Linux-syscall-note). >>> Buffer arguments are identified by GEM handles; the driver never >>> accepts DMA-BUF fds directly in any IOCTL. >>> >>> Patch Series Organization >>> ========================== >>> >>> Patch 01: MAINTAINERS entry >>> Patch 02: Driver documentation (Documentation/accel/qda/) >>> Patches 03-04: Core driver skeleton and compute bus >>> Patch 05: iommu: Register qda-compute-cb bus with IOMMU subsystem >>> Patches 06-07: CB device enumeration and memory manager >>> Patch 08: QUERY IOCTL and UAPI header >>> Patches 09-11: GEM buffer management and PRIME import >>> Patches 12-15: FastRPC protocol (invoke, session create/release, >>> map/unmap) >>> >>> Open Items >>> =========== >>> >>> 1. Device-Tree Compatible String >>> The QDA driver uses the same device-tree node structure and >>> properties as the existing fastrpc driver in drivers/misc/. A >>> mechanism is needed to allow the QDA driver to bind to its device >>> node independently of the fastrpc driver. >>> >>> The intended coexistence model is: platforms that require the >>> complete fastrpc feature set continue to use "qcom,fastrpc"; new >>> platforms where QDA's feature set is sufficient use a QDA-specific >>> compatible string. New feature development is directed toward QDA. >>> >>> The options under consideration are: >>> >>> a) Add a new "qcom,qda" compatible string to the existing >>> qcom,fastrpc.yaml binding, since the DT node structure and >>> properties are identical. >> No >> >>> >>> b) Introduce a separate qcom,qda.yaml binding that references or >>> inherits the fastrpc binding properties. >> >> No >> >>> >>> Seeking guidance from DT binding maintainers on the preferred >>> approach. >> >> Grow existing driver. You do not get new driver, you do not get new >> bindings. >> > > And this was already questioned at v1 (the true v1, not v1+1) but you > ignored the comment. The discussion were around compat layers in v1 patch (which is not yet concluded) and on whether this driver is going to be an alternative or a replacement. Please excuse be of overlooking any NAK on the true v1 series, but I can't still find it. Thanks for your review, Krzysztof. //Ekansh> > Great, so here goes away trust. > > NAK > > Best regards, > Krzysztof
On 19/08/2026 15:32, Ekansh Gupta wrote: > On 19-08-2026 00:51, Krzysztof Kozlowski wrote: >> On 18/08/2026 21:13, Krzysztof Kozlowski wrote: >>> On 17/08/2026 06:47, Ekansh Gupta wrote: >>>> This patch series introduces the Qualcomm DSP Accelerator (QDA) driver, >>>> a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a >>>> standardized interface for offloading computational tasks to DSPs found >>>> on Qualcomm SoCs, supporting all DSP domains. >>>> >>>> The QDA driver implements the FastRPC protocol over the DRM accel >>>> subsystem. It uses the same device-tree node structure as the existing >>>> fastrpc driver in drivers/misc/. The approach for binding the QDA driver >>>> to device-tree nodes while coexisting with the fastrpc driver is an open >>>> item described below. >>> >>> No. Grow/replace/improve existing driver instead of coming with a duplicate. >>> >>> That's a standard upstream requirement, basically given on every >>> upstreaming guide. >>> >>> Please watch old talk from Greg - "I Don’t Want Your Code!". >>> >>>> >>>> v1: https://lore.kernel.org/all/20260519-qda-series-v1-0-b2d984c297f8@oss.qualcomm.com/ >>>> RFC: https://lore.kernel.org/dri-devel/20260224-qda-firstpost-v1-0-fe46a9c1a046@oss.qualcomm.com/T/ >>>> >>>> Changes since v1 >>>> ================ >>>> >>>> The v1 review raised two architectural objections and one correctness >>>> issue; all three are resolved in v2: >>>> >>>> * Christian König (dma-buf maintainer) pointed out that the imported- >>>> buffer path silently assumed the IOMMU maps every buffer as a single >>>> contiguous range, which is not guaranteed. v2 walks the scatterlist >>>> and cleanly rejects non-contiguous imports; contiguous imports (e.g. >>>> CMA DMA-buf heap) are accepted. (patch 11) >>>> >>>> * Dmitry Baryshkov objected to three different buffer-passing formats >>>> in the invoke IOCTL (DMA-BUF fd, direct/inline, DMA handle). v2 >>>> passes only GEM handles; userspace imports any fd to a GEM handle >>>> with DRM_IOCTL_PRIME_FD_TO_HANDLE before invoking. Packing and >>>> overlap handling are left to userspace. (patch 12) >>>> >>>> * The memory manager (patch 07) used a fixed 16-entry array without >>>> justification and leaked the device descriptor on teardown. v2 >>>> allocates the array from the DT context-bank count (as Dmitry >>>> suggested) and frees it correctly. >>>> >>>> User-space staging branch >>>> ========================= >>>> https://github.com/qualcomm/fastrpc/tree/accel/staging >>>> >>>> Key Features >>>> ============ >>>> >>>> * Standard DRM accelerator interface via /dev/accel/accelN >>>> * GEM-based buffer management with DMA-BUF import (PRIME) >>>> * IOMMU-based memory isolation using per-process context banks >>>> * FastRPC protocol implementation for DSP communication >>>> * RPMsg transport layer for reliable message passing >>>> * Support for all DSP domains (ADSP, CDSP, SDSP, GDSP) >>>> * DRM IOCTL interface for DSP session management, buffer allocation, >>>> and remote procedure invocation >>>> >>>> Architecture >>>> ============ >>>> >>>> 1. DRM Accelerator Framework Integration >>>> The driver registers as a DRM accel device, exposing a standard >>>> /dev/accel/accelN character device node. This provides established >>>> DRM infrastructure for device management, file operations, and >>>> IOCTL dispatch. >>>> >>>> 2. Memory Management >>>> Buffers are managed as GEM objects with PRIME support for DMA-BUF >>>> import. This enables buffer sharing with other DRM drivers (GPU, >>>> camera, video) using standard kernel mechanisms. Only contiguous >>>> imports are accepted; the driver verifies contiguity at import time >>>> rather than assuming it. >>>> >>>> 3. IOMMU Context Bank Management >>>> IOMMU context banks (CBs) are represented as proper struct device >>>> instances on a custom virtual bus (qda-compute-cb). Each CB device >>>> is registered with the IOMMU subsystem and receives its own IOMMU >>>> domain, enabling per-session address space isolation. The custom >>>> bus was introduced because IOMMU context banks are synthetic >>>> constructs — not real platform devices — and to ensure CB device >>>> lifetime is strictly subordinate to the parent QDA device. >>>> See also: https://lore.kernel.org/all/245d602f-3037-4ae3-9af9-d98f37258aae@oss.qualcomm.com/ >>>> >>>> 4. Memory Manager Architecture >>>> The memory manager maintains a registry of IOMMU devices in an >>>> array sized to the number of context banks described in the device >>>> tree, and coordinates per-process device assignment with reference- >>>> counted lifetime management. The DMA-coherent backend allocates >>>> buffers with SID-prefixed DMA addresses for DSP firmware >>>> compatibility. >>>> >>>> 5. Transport Layer >>>> RPMsg communication is handled in a dedicated transport layer >>>> (qda_rpmsg.c), separate from the core DRM driver logic. >>>> >>>> 6. Code Organization >>>> The driver is organized across multiple files (~4800 lines total): >>>> * qda_drv.c: Core driver and DRM integration >>>> * qda_rpmsg.c: RPMsg transport layer >>>> * qda_cb.c: Context bank device management >>>> * qda_compute_bus.c: Custom virtual bus for CB devices >>>> * qda_gem.c: GEM object management >>>> * qda_prime.c: DMA-BUF import (PRIME) >>>> * qda_memory_manager.c: IOMMU device registry and allocation >>>> * qda_memory_dma.c: DMA-coherent allocation backend >>>> * qda_fastrpc.c: FastRPC protocol implementation >>>> * qda_ioctl.c: IOCTL dispatch >>>> >>>> 7. UAPI Design >>>> The driver exposes DRM-style IOCTLs defined in >>>> include/uapi/drm/qda_accel.h, following DRM UAPI conventions >>>> (__u32/__u64 types, C++ guard, GPL-2.0-only WITH Linux-syscall-note). >>>> Buffer arguments are identified by GEM handles; the driver never >>>> accepts DMA-BUF fds directly in any IOCTL. >>>> >>>> Patch Series Organization >>>> ========================== >>>> >>>> Patch 01: MAINTAINERS entry >>>> Patch 02: Driver documentation (Documentation/accel/qda/) >>>> Patches 03-04: Core driver skeleton and compute bus >>>> Patch 05: iommu: Register qda-compute-cb bus with IOMMU subsystem >>>> Patches 06-07: CB device enumeration and memory manager >>>> Patch 08: QUERY IOCTL and UAPI header >>>> Patches 09-11: GEM buffer management and PRIME import >>>> Patches 12-15: FastRPC protocol (invoke, session create/release, >>>> map/unmap) >>>> >>>> Open Items >>>> =========== >>>> >>>> 1. Device-Tree Compatible String >>>> The QDA driver uses the same device-tree node structure and >>>> properties as the existing fastrpc driver in drivers/misc/. A >>>> mechanism is needed to allow the QDA driver to bind to its device >>>> node independently of the fastrpc driver. >>>> >>>> The intended coexistence model is: platforms that require the >>>> complete fastrpc feature set continue to use "qcom,fastrpc"; new >>>> platforms where QDA's feature set is sufficient use a QDA-specific >>>> compatible string. New feature development is directed toward QDA. >>>> >>>> The options under consideration are: >>>> >>>> a) Add a new "qcom,qda" compatible string to the existing >>>> qcom,fastrpc.yaml binding, since the DT node structure and >>>> properties are identical. >>> No >>> >>>> >>>> b) Introduce a separate qcom,qda.yaml binding that references or >>>> inherits the fastrpc binding properties. >>> >>> No >>> >>>> >>>> Seeking guidance from DT binding maintainers on the preferred >>>> approach. >>> >>> Grow existing driver. You do not get new driver, you do not get new >>> bindings. >>> >> >> And this was already questioned at v1 (the true v1, not v1+1) but you >> ignored the comment. > The discussion were around compat layers in v1 patch (which is not yet > concluded) and on whether this driver is going to be an alternative or a No, you got comment, from Trilok I think, asking what is the plan in respect of existing fastrpc driver. Best regards, Krzysztof
On 17/08/2026 06:47, Ekansh Gupta wrote: > This patch series introduces the Qualcomm DSP Accelerator (QDA) driver, > a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a > standardized interface for offloading computational tasks to DSPs found > on Qualcomm SoCs, supporting all DSP domains. > > The QDA driver implements the FastRPC protocol over the DRM accel > subsystem. It uses the same device-tree node structure as the existing > fastrpc driver in drivers/misc/. The approach for binding the QDA driver > to device-tree nodes while coexisting with the fastrpc driver is an open > item described below. > > v1: https://lore.kernel.org/all/20260519-qda-series-v1-0-b2d984c297f8@oss.qualcomm.com/ > RFC: https://lore.kernel.org/dri-devel/20260224-qda-firstpost-v1-0-fe46a9c1a046@oss.qualcomm.com/T/ So this is a v3, not v2. Please start using b4 correctly, so versioning will be kept instead of faking the numbers. Best regards, Krzysztof
On 19-08-2026 00:48, Krzysztof Kozlowski wrote: > On 17/08/2026 06:47, Ekansh Gupta wrote: >> This patch series introduces the Qualcomm DSP Accelerator (QDA) driver, >> a DRM-based accelerator driver for Qualcomm DSPs. The driver provides a >> standardized interface for offloading computational tasks to DSPs found >> on Qualcomm SoCs, supporting all DSP domains. >> >> The QDA driver implements the FastRPC protocol over the DRM accel >> subsystem. It uses the same device-tree node structure as the existing >> fastrpc driver in drivers/misc/. The approach for binding the QDA driver >> to device-tree nodes while coexisting with the fastrpc driver is an open >> item described below. >> >> v1: https://lore.kernel.org/all/20260519-qda-series-v1-0-b2d984c297f8@oss.qualcomm.com/ >> RFC: https://lore.kernel.org/dri-devel/20260224-qda-firstpost-v1-0-fe46a9c1a046@oss.qualcomm.com/T/ > > So this is a v3, not v2. Please start using b4 correctly, so versioning > will be kept instead of faking the numbers. I'll correct the versioning in the next post. Thank you for pointing this. > > Best regards, > Krzysztof
© 2016 - 2026 Red Hat, Inc.