.../display/tegra/nvidia,tegra124-vic.yaml | 8 + .../bindings/display/tegra/nvidia,tegra186-dc.yaml | 10 + .../bindings/display/tegra/nvidia,tegra20-dc.yaml | 10 +- .../display/tegra/nvidia,tegra20-host1x.yaml | 7 + .../bindings/gpu/host1x/nvidia,tegra234-nvdec.yaml | 8 + .../nvidia,tegra-video-protection-region.yaml | 75 ++ arch/arm64/boot/dts/nvidia/tegra234.dtsi | 45 + arch/arm64/boot/dts/nvidia/tegra264.dtsi | 33 + arch/arm64/mm/pageattr.c | 2 + drivers/dma-buf/dma-heap.c | 52 + drivers/dma-buf/heaps/Kconfig | 12 + drivers/dma-buf/heaps/Makefile | 10 + drivers/dma-buf/heaps/tegra-vpr-init.c | 133 +++ drivers/dma-buf/heaps/tegra-vpr.c | 1210 ++++++++++++++++++++ drivers/dma-buf/heaps/tegra-vpr.h | 73 ++ drivers/of/of_numa.c | 1 + include/linux/bitmap.h | 25 +- include/linux/cma.h | 4 + include/linux/dma-heap.h | 2 + include/trace/events/cma.h | 77 +- include/trace/events/tegra_vpr.h | 57 + mm/cma.c | 88 +- 22 files changed, 1913 insertions(+), 29 deletions(-)
This series adds support for the video protection region (VPR) used on
Tegra SoC devices. It's a special region of memory that is protected
from accesses by the CPU and used to store DRM protected content (both
decrypted stream data as well as decoded video frames).
Patches 1 through 3 add DT binding documentation for the VPR and add the
VPR to the list of memory-region items for display, host1x and NVDEC.
The set_direct_map_*_noflush() functions that will be used later in this
series are exported in patch 4 so that the drivers that use them can be
built as a module.
Patch 5 adds bitmap_allocate(), which is like bitmap_allocate_region()
but works on sizes that are not a power of two.
The of_node_to_nid() function is exported in patch 6 because it is used
in a later patch adding a driver that can be built as a module.
Patch 7 introduces new APIs needed by the Tegra VPR implementation that
allow memory to be allocated at a fixed offset within a CMA area. Tegra
VPR needs this in order to implement its own allocator on top of CMA to
meet the strict hardware requirements. This replaces the dynamic CMA
area creation patch from earlier versions.
Patch 8 adds some infrastructure for DMA heap implementations to provide
information through debugfs.
The Tegra VPR implementation is added in patch 9. See its commit message
for more details about the specifics of this implementation.
Finally, patches 10-12 add the VPR placeholder node on Tegra234 and
Tegra264 and hook it up to the host1x node so that it can make use of
this region.
Changes in v6:
- refactor cma_alloc_at() to maximize code reuse
- allow building Tegra VPR support as a module
- requires two new patches to export symbols
- Link to v5: https://patch.msgid.link/20260814-tegra-vpr-v5-0-71832b5d0246@nvidia.com
Changes in v5:
- use a single CMA area in combination with the new cma_alloc_at() API
- drop dynamic CMA area allocation patch
- various cleanups
- Link to v4: https://patch.msgid.link/20260807-tegra-vpr-v4-0-5510d16af89e@nvidia.com
Changes in v4:
- Link to v3: https://patch.msgid.link/20260701-tegra-vpr-v3-0-d80f7b871bb4@nvidia.com
- fully remove from linear map while chunks are active
- address checkpatch.pl and Sashiko comments
- improve error handling
- remove freezer support
Changes in v3:
- Link to v2: https://patch.msgid.link/20260122161009.3865888-1-thierry.reding@kernel.org
- introduce set_memory_device() and set_memory_normal()
- rename VPR nodes to "protected"
- add Tegra264 placeholder nodes
Changes in v2:
- Link to v1: https://patch.msgid.link/20250902154630.4032984-1-thierry.reding@gmail.com
- Tegra VPR implementation is now more optimized to reduce the number of
(very slow) resize operations, and allows cross-chunk allocations
- dynamic CMA areas are now trackd separately from static ones, but the
global number of CMA pages accounts for all areas
Thierry
Signed-off-by: Thierry Reding <treding@nvidia.com>
---
Thierry Reding (12):
dt-bindings: reserved-memory: Document Tegra VPR
dt-bindings: display: tegra: Document memory regions
dt-bindings: gpu: host1x: Document memory-regions for NVDEC
arm64/mm: Export set_direct_map_*_noflush() APIs
bitmap: Add bitmap_allocate() function
of: Export of_node_to_nid()
mm/cma: Introduce cma_alloc_at() API
dma-buf: heaps: Add debugfs support
dma-buf: heaps: Add support for Tegra VPR
arm64: tegra: Add VPR placeholder node on Tegra234
arm64: tegra: Hook up VPR to host1x
arm64: tegra: Add VPR placeholder node on Tegra264
.../display/tegra/nvidia,tegra124-vic.yaml | 8 +
.../bindings/display/tegra/nvidia,tegra186-dc.yaml | 10 +
.../bindings/display/tegra/nvidia,tegra20-dc.yaml | 10 +-
.../display/tegra/nvidia,tegra20-host1x.yaml | 7 +
.../bindings/gpu/host1x/nvidia,tegra234-nvdec.yaml | 8 +
.../nvidia,tegra-video-protection-region.yaml | 75 ++
arch/arm64/boot/dts/nvidia/tegra234.dtsi | 45 +
arch/arm64/boot/dts/nvidia/tegra264.dtsi | 33 +
arch/arm64/mm/pageattr.c | 2 +
drivers/dma-buf/dma-heap.c | 52 +
drivers/dma-buf/heaps/Kconfig | 12 +
drivers/dma-buf/heaps/Makefile | 10 +
drivers/dma-buf/heaps/tegra-vpr-init.c | 133 +++
drivers/dma-buf/heaps/tegra-vpr.c | 1210 ++++++++++++++++++++
drivers/dma-buf/heaps/tegra-vpr.h | 73 ++
drivers/of/of_numa.c | 1 +
include/linux/bitmap.h | 25 +-
include/linux/cma.h | 4 +
include/linux/dma-heap.h | 2 +
include/trace/events/cma.h | 77 +-
include/trace/events/tegra_vpr.h | 57 +
mm/cma.c | 88 +-
22 files changed, 1913 insertions(+), 29 deletions(-)
---
base-commit: 850de4c714df026eb4c869fdb01dd722003e1ae3
change-id: 20260507-tegra-vpr-cd4bc2509c4c
Best regards,
--
Thierry Reding <treding@nvidia.com>
Hi Thierry, On Fri, Sep 04, 2026 at 12:44:51PM +0200, Thierry Reding wrote: > This series adds support for the video protection region (VPR) used on > Tegra SoC devices. It's a special region of memory that is protected > from accesses by the CPU and used to store DRM protected content (both > decrypted stream data as well as decoded video frames). > > Patches 1 through 3 add DT binding documentation for the VPR and add the > VPR to the list of memory-region items for display, host1x and NVDEC. > > The set_direct_map_*_noflush() functions that will be used later in this > series are exported in patch 4 so that the drivers that use them can be > built as a module. > > Patch 5 adds bitmap_allocate(), which is like bitmap_allocate_region() > but works on sizes that are not a power of two. > > The of_node_to_nid() function is exported in patch 6 because it is used > in a later patch adding a driver that can be built as a module. > > Patch 7 introduces new APIs needed by the Tegra VPR implementation that > allow memory to be allocated at a fixed offset within a CMA area. Tegra > VPR needs this in order to implement its own allocator on top of CMA to > meet the strict hardware requirements. This replaces the dynamic CMA > area creation patch from earlier versions. Did you get a chance to see how this could work with Vincent's series: https://lore.kernel.org/r/20260902104712.2399797-1-vdonnefort@google.com ? I think that should remove your reliance on can_set_direct_map() and mean that you can retain block mappings for most of the linear mapping. Looks like you forgot to cc him, so I added him here. Cheers, Will
On Fri, Sep 04, 2026 at 12:41:05PM +0100, Will Deacon wrote: > Hi Thierry, > > On Fri, Sep 04, 2026 at 12:44:51PM +0200, Thierry Reding wrote: > > This series adds support for the video protection region (VPR) used on > > Tegra SoC devices. It's a special region of memory that is protected > > from accesses by the CPU and used to store DRM protected content (both > > decrypted stream data as well as decoded video frames). > > > > Patches 1 through 3 add DT binding documentation for the VPR and add the > > VPR to the list of memory-region items for display, host1x and NVDEC. > > > > The set_direct_map_*_noflush() functions that will be used later in this > > series are exported in patch 4 so that the drivers that use them can be > > built as a module. > > > > Patch 5 adds bitmap_allocate(), which is like bitmap_allocate_region() > > but works on sizes that are not a power of two. > > > > The of_node_to_nid() function is exported in patch 6 because it is used > > in a later patch adding a driver that can be built as a module. > > > > Patch 7 introduces new APIs needed by the Tegra VPR implementation that > > allow memory to be allocated at a fixed offset within a CMA area. Tegra > > VPR needs this in order to implement its own allocator on top of CMA to > > meet the strict hardware requirements. This replaces the dynamic CMA > > area creation patch from earlier versions. > > Did you get a chance to see how this could work with Vincent's series: > > https://lore.kernel.org/r/20260902104712.2399797-1-vdonnefort@google.com > > ? I think that should remove your reliance on can_set_direct_map() and > mean that you can retain block mappings for most of the linear mapping. I'm not sure if it would help all that much. Yes, if we mark the VPR region as LLMAP (or PTE_MAP, whichever it ends up being), it should make the checks for can_set_direct_map() redundant. However, from what I can tell, Vincent's series still forces page-granularity on these regions, so it won't retain block mappings at all for them. The block mappings can be retained for the non-VPR memory, so that's nice. It also reduces the amount of external prerequisites, but I had kind of hoped that we could go one step further and keep block mappings even for the VPR memory if the region happened to be a multiple of the block size. The recent addition of page count to the set_direct_map_*() functions helps reduce the amount of checks that need to be run, so maybe there's not too much to be gained from removing whole block mappings at once from the linear map. Thierry
On Tue, Sep 08, 2026 at 10:39:17AM +0200, Thierry Reding wrote:
> On Fri, Sep 04, 2026 at 12:41:05PM +0100, Will Deacon wrote:
> > Hi Thierry,
> >
> > On Fri, Sep 04, 2026 at 12:44:51PM +0200, Thierry Reding wrote:
> > > This series adds support for the video protection region (VPR) used on
> > > Tegra SoC devices. It's a special region of memory that is protected
> > > from accesses by the CPU and used to store DRM protected content (both
> > > decrypted stream data as well as decoded video frames).
> > >
> > > Patches 1 through 3 add DT binding documentation for the VPR and add the
> > > VPR to the list of memory-region items for display, host1x and NVDEC.
> > >
> > > The set_direct_map_*_noflush() functions that will be used later in this
> > > series are exported in patch 4 so that the drivers that use them can be
> > > built as a module.
> > >
> > > Patch 5 adds bitmap_allocate(), which is like bitmap_allocate_region()
> > > but works on sizes that are not a power of two.
> > >
> > > The of_node_to_nid() function is exported in patch 6 because it is used
> > > in a later patch adding a driver that can be built as a module.
> > >
> > > Patch 7 introduces new APIs needed by the Tegra VPR implementation that
> > > allow memory to be allocated at a fixed offset within a CMA area. Tegra
> > > VPR needs this in order to implement its own allocator on top of CMA to
> > > meet the strict hardware requirements. This replaces the dynamic CMA
> > > area creation patch from earlier versions.
> >
> > Did you get a chance to see how this could work with Vincent's series:
> >
> > https://lore.kernel.org/r/20260902104712.2399797-1-vdonnefort@google.com
> >
> > ? I think that should remove your reliance on can_set_direct_map() and
> > mean that you can retain block mappings for most of the linear mapping.
>
> I'm not sure if it would help all that much. Yes, if we mark the VPR
> region as LLMAP (or PTE_MAP, whichever it ends up being), it should make
> the checks for can_set_direct_map() redundant. However, from what I can
> tell, Vincent's series still forces page-granularity on these regions,
> so it won't retain block mappings at all for them.
>
> The block mappings can be retained for the non-VPR memory, so that's
> nice. It also reduces the amount of external prerequisites, but I had
> kind of hoped that we could go one step further and keep block mappings
> even for the VPR memory if the region happened to be a multiple of the
> block size.
>
> The recent addition of page count to the set_direct_map_*() functions
> helps reduce the amount of checks that need to be run, so maybe there's
> not too much to be gained from removing whole block mappings at once
> from the linear map.
I don't think everyone received Sashiko's review, so let me discuss this
here. Sashiko rightly pointed out that set_direct_map_invalid_noflush()
and set_direct_map_default_noflush() return 0 when can_set_direct_map()
fails and that code will then simply continue to work as if the pages
had been removed (or added back) even though they weren't.
arch/arm64/mm/pageattr.c:set_direct_map_invalid_noflush() {
...
if (!can_set_direct_map())
return 0;
...
}
Looking into this a bit, it looks like this is maybe a remnant from the
early days when the check was simpler ("if (!rodata_full)", though I'm
not sure the 0 return value made sense even then), but it seems wrong
indeed for this to result in success when clearly the operation was
skipped.
None of the other architectures seem to have similar checks, except for
clear cases of no-ops (like the address being outside the linear
mapping, the number of pages being 0 or there not being any actual
changes). All of the three callers seem to be prepared to deal with
failure, so I think we should just make these fail instead of returning
0.
Any thought?
Thierry
On Tue, Sep 08, 2026 at 10:39:17AM +0200, Thierry Reding wrote: > On Fri, Sep 04, 2026 at 12:41:05PM +0100, Will Deacon wrote: > > Hi Thierry, > > > > On Fri, Sep 04, 2026 at 12:44:51PM +0200, Thierry Reding wrote: > > > This series adds support for the video protection region (VPR) used on > > > Tegra SoC devices. It's a special region of memory that is protected > > > from accesses by the CPU and used to store DRM protected content (both > > > decrypted stream data as well as decoded video frames). > > > > > > Patches 1 through 3 add DT binding documentation for the VPR and add the > > > VPR to the list of memory-region items for display, host1x and NVDEC. > > > > > > The set_direct_map_*_noflush() functions that will be used later in this > > > series are exported in patch 4 so that the drivers that use them can be > > > built as a module. > > > > > > Patch 5 adds bitmap_allocate(), which is like bitmap_allocate_region() > > > but works on sizes that are not a power of two. > > > > > > The of_node_to_nid() function is exported in patch 6 because it is used > > > in a later patch adding a driver that can be built as a module. > > > > > > Patch 7 introduces new APIs needed by the Tegra VPR implementation that > > > allow memory to be allocated at a fixed offset within a CMA area. Tegra > > > VPR needs this in order to implement its own allocator on top of CMA to > > > meet the strict hardware requirements. This replaces the dynamic CMA > > > area creation patch from earlier versions. > > > > Did you get a chance to see how this could work with Vincent's series: > > > > https://lore.kernel.org/r/20260902104712.2399797-1-vdonnefort@google.com > > > > ? I think that should remove your reliance on can_set_direct_map() and > > mean that you can retain block mappings for most of the linear mapping. > > I'm not sure if it would help all that much. Yes, if we mark the VPR > region as LLMAP (or PTE_MAP, whichever it ends up being), it should make > the checks for can_set_direct_map() redundant. However, from what I can > tell, Vincent's series still forces page-granularity on these regions, > so it won't retain block mappings at all for them. > > The block mappings can be retained for the non-VPR memory, so that's > nice. It also reduces the amount of external prerequisites, but I had > kind of hoped that we could go one step further and keep block mappings > even for the VPR memory if the region happened to be a multiple of the > block size. > > The recent addition of page count to the set_direct_map_*() functions > helps reduce the amount of checks that need to be run, so maybe there's > not too much to be gained from removing whole block mappings at once > from the linear map. > > Thierry I should be able to add PMD_SIZE mapping support to the series. That was actually my original idea as we can easily force the CMA allocation granule to be PMD_SIZE too. I didn't implement it as I thought there were not much interest in the end (and also as contiguous.c is always using PAGE_SIZE granularity). But now as I have implemented a specific pool (and do not use contiguous.c as originally planned), if you believe it is important for the VPR driver, let me see if I can extend the support in a V2. -- Vincent
On Tue, Sep 08, 2026 at 09:57:57AM +0100, Vincent Donnefort wrote: > On Tue, Sep 08, 2026 at 10:39:17AM +0200, Thierry Reding wrote: > > On Fri, Sep 04, 2026 at 12:41:05PM +0100, Will Deacon wrote: > > > Hi Thierry, > > > > > > On Fri, Sep 04, 2026 at 12:44:51PM +0200, Thierry Reding wrote: > > > > This series adds support for the video protection region (VPR) used on > > > > Tegra SoC devices. It's a special region of memory that is protected > > > > from accesses by the CPU and used to store DRM protected content (both > > > > decrypted stream data as well as decoded video frames). > > > > > > > > Patches 1 through 3 add DT binding documentation for the VPR and add the > > > > VPR to the list of memory-region items for display, host1x and NVDEC. > > > > > > > > The set_direct_map_*_noflush() functions that will be used later in this > > > > series are exported in patch 4 so that the drivers that use them can be > > > > built as a module. > > > > > > > > Patch 5 adds bitmap_allocate(), which is like bitmap_allocate_region() > > > > but works on sizes that are not a power of two. > > > > > > > > The of_node_to_nid() function is exported in patch 6 because it is used > > > > in a later patch adding a driver that can be built as a module. > > > > > > > > Patch 7 introduces new APIs needed by the Tegra VPR implementation that > > > > allow memory to be allocated at a fixed offset within a CMA area. Tegra > > > > VPR needs this in order to implement its own allocator on top of CMA to > > > > meet the strict hardware requirements. This replaces the dynamic CMA > > > > area creation patch from earlier versions. > > > > > > Did you get a chance to see how this could work with Vincent's series: > > > > > > https://lore.kernel.org/r/20260902104712.2399797-1-vdonnefort@google.com > > > > > > ? I think that should remove your reliance on can_set_direct_map() and > > > mean that you can retain block mappings for most of the linear mapping. > > > > I'm not sure if it would help all that much. Yes, if we mark the VPR > > region as LLMAP (or PTE_MAP, whichever it ends up being), it should make > > the checks for can_set_direct_map() redundant. However, from what I can > > tell, Vincent's series still forces page-granularity on these regions, > > so it won't retain block mappings at all for them. > > > > The block mappings can be retained for the non-VPR memory, so that's > > nice. It also reduces the amount of external prerequisites, but I had > > kind of hoped that we could go one step further and keep block mappings > > even for the VPR memory if the region happened to be a multiple of the > > block size. > > > > The recent addition of page count to the set_direct_map_*() functions > > helps reduce the amount of checks that need to be run, so maybe there's > > not too much to be gained from removing whole block mappings at once > > from the linear map. > > > > Thierry > > I should be able to add PMD_SIZE mapping support to the series. That was > actually my original idea as we can easily force the CMA allocation granule to > be PMD_SIZE too. > > I didn't implement it as I thought there were not much interest in the end (and > also as contiguous.c is always using PAGE_SIZE granularity). > > But now as I have implemented a specific pool (and do not use contiguous.c as > originally planned), if you believe it is important for the VPR driver, let me > see if I can extend the support in a V2. I don't think it needs to be part of a v2 and can be a follow-up. It should be transparent from an API point of view and merely be an optimisation for that specific case. Eventually it'd be nice to have, though it might also be worth checking what the actually gains are. Thierry
On Tue, Sep 08, 2026 at 11:08:26AM +0200, Thierry Reding wrote: > On Tue, Sep 08, 2026 at 09:57:57AM +0100, Vincent Donnefort wrote: > > On Tue, Sep 08, 2026 at 10:39:17AM +0200, Thierry Reding wrote: > > > On Fri, Sep 04, 2026 at 12:41:05PM +0100, Will Deacon wrote: > > > > Hi Thierry, > > > > > > > > On Fri, Sep 04, 2026 at 12:44:51PM +0200, Thierry Reding wrote: > > > > > This series adds support for the video protection region (VPR) used on > > > > > Tegra SoC devices. It's a special region of memory that is protected > > > > > from accesses by the CPU and used to store DRM protected content (both > > > > > decrypted stream data as well as decoded video frames). > > > > > > > > > > Patches 1 through 3 add DT binding documentation for the VPR and add the > > > > > VPR to the list of memory-region items for display, host1x and NVDEC. > > > > > > > > > > The set_direct_map_*_noflush() functions that will be used later in this > > > > > series are exported in patch 4 so that the drivers that use them can be > > > > > built as a module. > > > > > > > > > > Patch 5 adds bitmap_allocate(), which is like bitmap_allocate_region() > > > > > but works on sizes that are not a power of two. > > > > > > > > > > The of_node_to_nid() function is exported in patch 6 because it is used > > > > > in a later patch adding a driver that can be built as a module. > > > > > > > > > > Patch 7 introduces new APIs needed by the Tegra VPR implementation that > > > > > allow memory to be allocated at a fixed offset within a CMA area. Tegra > > > > > VPR needs this in order to implement its own allocator on top of CMA to > > > > > meet the strict hardware requirements. This replaces the dynamic CMA > > > > > area creation patch from earlier versions. > > > > > > > > Did you get a chance to see how this could work with Vincent's series: > > > > > > > > https://lore.kernel.org/r/20260902104712.2399797-1-vdonnefort@google.com > > > > > > > > ? I think that should remove your reliance on can_set_direct_map() and > > > > mean that you can retain block mappings for most of the linear mapping. > > > > > > I'm not sure if it would help all that much. Yes, if we mark the VPR > > > region as LLMAP (or PTE_MAP, whichever it ends up being), it should make > > > the checks for can_set_direct_map() redundant. However, from what I can > > > tell, Vincent's series still forces page-granularity on these regions, > > > so it won't retain block mappings at all for them. > > > > > > The block mappings can be retained for the non-VPR memory, so that's > > > nice. It also reduces the amount of external prerequisites, but I had > > > kind of hoped that we could go one step further and keep block mappings > > > even for the VPR memory if the region happened to be a multiple of the > > > block size. > > > > > > The recent addition of page count to the set_direct_map_*() functions > > > helps reduce the amount of checks that need to be run, so maybe there's > > > not too much to be gained from removing whole block mappings at once > > > from the linear map. > > > > > > Thierry > > > > I should be able to add PMD_SIZE mapping support to the series. That was > > actually my original idea as we can easily force the CMA allocation granule to > > be PMD_SIZE too. > > > > I didn't implement it as I thought there were not much interest in the end (and > > also as contiguous.c is always using PAGE_SIZE granularity). > > > > But now as I have implemented a specific pool (and do not use contiguous.c as > > originally planned), if you believe it is important for the VPR driver, let me > > see if I can extend the support in a V2. > > I don't think it needs to be part of a v2 and can be a follow-up. It > should be transparent from an API point of view and merely be an > optimisation for that specific case. > > Eventually it'd be nice to have, though it might also be worth checking > what the actually gains are. > > Thierry Ack. I'll keep it as is then and we will see later how to extend it, if it is necessary. -- Vincent
© 2016 - 2026 Red Hat, Inc.