.../bindings/perf/brcm,bcm2835-axiperf.yaml | 77 + drivers/perf/rpi_axi_pmu.c | 2903 +++++++++++++++++ 2 files changed, 2980 insertions(+) create mode 100644 Documentation/devicetree/bindings/perf/brcm,bcm2835-axiperf.yaml create mode 100644 drivers/perf/rpi_axi_pmu.c
Motivation & Background
==============================
Currently, Linux lacks a standard perf-API compatible driver for the
Broadcom AXI performance counter blocks on Raspberry Pi platforms. Prior
out-of-tree vendor solutions relied on custom debugfs nodes and ad-hoc
kthreads, preventing integration with standard Linux perf tooling (`perf stat`,
`perf list`, etc.).
This driver implements standard `struct pmu` hardware uncore callbacks
under `drivers/perf/`, exposing human-readable sysfs event aliases, unit
scaling (`Bytes`), and bus filtering directly to user space.
Key Architectural Improvements & Features
==============================
1. Standard Linux Perf Integration:
- Exposes uncore AXI interconnect events via `/sys/bus/event_source/devices/rpi_axi_pmu/`.
- Supports event sampling and hardware counter accumulation (`local64_add`),
automatically managing 31-bit hardware counter wraparound across high-bandwidth
interconnect transfers.
2. CPU Hotplug Support (`cpuhp`):
- Registers dynamic CPU hotplug notifiers (`CPUHP_AP_ONLINE_DYN`).
- Automatically migrates PMU context (`perf_pmu_migrate_context`) to an online
CPU core when a designated CPU goes offline, safely guarding memory teardown boundaries.
3. Hybrid Memory-Mapped & Mailbox Work Queue Architecture:
- System Monitor (MMIO): Performs fast atomic-safe memory reads (~15ns)
directly mapped over ARM physical memory space (`MON_SYSTEM`).
- VPU Monitor (Mailbox IPC): For Broadcom BCM2835-BCM2711 platforms (RPi 1-4),
VideoCore VPU monitor IPC calls are offloaded to process context via a dedicated
workqueue (`vpu_work`) and serialized under `vpu_mutex`. This avoids atomic
sleeps or blocking in timer/interrupt context.
4. SoC Generation Support:
- Patch 1 documents device bindings for Broadcom.
- Patch 2 adds core driver support for Broadcom BCM2835-BCM2711 (RPi 1-4).
- Patch 3 expands support for Broadcom BCM2712 (Raspberry Pi 5), adding PCIe RP1
Southbridge links, HEVC decoder, HVS display engine, and Cortex-A76 DSU L3
interconnect monitoring dynamically via native MMIO readouts.
Hardware Validation
==============================
The driver has been rigorously validated on real hardware:
- Raspberry Pi 400 (BCM2711): Validated System L2, ARM CPU, and VideoCore
VPU firmware mailbox IPC performance counters organically spanning
multi-slice multiplexing boundaries.
- Raspberry Pi 5 (BCM2712): Validated live byte throughput across HVS
display refresh cycles, Cortex-A76 DSU L3 interconnect memory traffic,
and PCIe RP1 Southbridge transfers independently configuring `bcm2712_`
filters.
Changes since v12
==============================
- Schema Enforcement: Added an `allOf` / `if` block to
`brcm,bcm2835-axiperf.yaml` to mandate the `firmware` property when
compiling dtbs against older silicon (`brcm,bcm2835-axiperf` and
`brcm,bcm2711-axiperf`). Validated the documentation example to
securely meet this constraint.
- Sysfs Arrays: Rehydrated the missing dynamic PMU Sysfs string
definitions (`PMU_EVENT_ATTR_STRING`) and memory pointers for the
BCM2712 CPU L2 and Master ID filtered events, ensuring they are
cleanly exported to Userspace.
- Whitespace formatting: Purged excess consecutive vertical blank
lines inadvertently dropped into the sysfs definitions across
boundaries in Patch 2 and Patch 3.
- Documentation nits: Rectified and clarified a misleading comment
regarding synchronous reads for the BCM2712 VPU explicitly isolated
from native runtime monitoring.
Ian Rogers (3):
dt-bindings: perf: Add Broadcom Raspberry Pi AXI PMU definition
perf: Add Raspberry Pi BCM2835 AXI PMU driver
perf: Add Raspberry Pi 5 (BCM2712) AXI PMU support
.../bindings/perf/brcm,bcm2835-axiperf.yaml | 77 +
drivers/perf/rpi_axi_pmu.c | 2903 +++++++++++++++++
2 files changed, 2980 insertions(+)
create mode 100644 Documentation/devicetree/bindings/perf/brcm,bcm2835-axiperf.yaml
create mode 100644 drivers/perf/rpi_axi_pmu.c
--
2.55.0.691.gc56d675ccc-goog
On 15/08/2026 19:27, Ian Rogers wrote: > Motivation & Background > ============================== > Currently, Linux lacks a standard perf-API compatible driver for the > Broadcom AXI performance counter blocks on Raspberry Pi platforms. Prior > out-of-tree vendor solutions relied on custom debugfs nodes and ad-hoc > kthreads, preventing integration with standard Linux perf tooling (`perf stat`, > `perf list`, etc.). > Slow down! One version per 24h in normal cycle, not 13 patchsets in two days! (and even rarer during the merge window) Best regards, Krzysztof
On Sun, Aug 16, 2026 at 11:20 PM Krzysztof Kozlowski <krzk@kernel.org> wrote: > > On 15/08/2026 19:27, Ian Rogers wrote: > > Motivation & Background > > ============================== > > Currently, Linux lacks a standard perf-API compatible driver for the > > Broadcom AXI performance counter blocks on Raspberry Pi platforms. Prior > > out-of-tree vendor solutions relied on custom debugfs nodes and ad-hoc > > kthreads, preventing integration with standard Linux perf tooling (`perf stat`, > > `perf list`, etc.). > > > > Slow down! > > One version per 24h in normal cycle, not 13 patchsets in two days! (and > even rarer during the merge window) I'm familiar with perf-tools which was set up to follow BPF's example, both have -next repos. The -next repos merge into linux-next for testing. I agree that cherry-picking shouldn't be done straight into a pull request to Linus. Is there a principle against using -next and linux-next? 13 patch sets resulted from Sashiko reviews and subsequent addressing of those comments. Yes, Sashiko was run internally before mailing the changes. No, waiting a week between versions (a proposal on this thread) doesn't make sense as addressing the feedback would take 3+ months. As already mentioned, the current churn simply reflects where we are with the tooling. I apologize; I did what I could to reduce it and you can see in the changes per version in the cover letter that this addresses more than just the mailing list Sashiko review. Fwiw, if you're not familiar with Sashiko's reviews it seems you should get it enabled for your mailing list - Linus is positive on it and it generally improves reviews and reduces reviewer/maintainer burden. b4 handles reply-to git send-emails with `b4 am`, each reply-to gets a fresh message-id and that's all the tools care about. I forgot to run get-maintainers.pl for the Documentation change, sorry. Since more important drivers (in particular Apple M1) are currently unaddressed in drivers/perf, I don't hold out much hope of this driver landing or receiving maintainer attention. I sent the patches merely to help the ecosystem and encourage vendors to interface with perf. So far the only constructive human review feedback I've addressed is from Florian. Thanks, Ian > Best regards, > Krzysztof
On 17/08/2026 08:57, Ian Rogers wrote: > On Sun, Aug 16, 2026 at 11:20 PM Krzysztof Kozlowski <krzk@kernel.org> wrote: >> >> On 15/08/2026 19:27, Ian Rogers wrote: >>> Motivation & Background >>> ============================== >>> Currently, Linux lacks a standard perf-API compatible driver for the >>> Broadcom AXI performance counter blocks on Raspberry Pi platforms. Prior >>> out-of-tree vendor solutions relied on custom debugfs nodes and ad-hoc >>> kthreads, preventing integration with standard Linux perf tooling (`perf stat`, >>> `perf list`, etc.). >>> >> >> Slow down! >> >> One version per 24h in normal cycle, not 13 patchsets in two days! (and >> even rarer during the merge window) > > I'm familiar with perf-tools which was set up to follow BPF's example, > both have -next repos. The -next repos merge into linux-next for > testing. I agree that cherry-picking shouldn't be done straight into a > pull request to Linus. Is there a principle against using -next and > linux-next? How is this relevant? I don't think you read my comment. Best regards, Krzysztof
On Mon, Aug 17, 2026 at 12:26 AM Krzysztof Kozlowski <krzk@kernel.org> wrote: > > On 17/08/2026 08:57, Ian Rogers wrote: > > On Sun, Aug 16, 2026 at 11:20 PM Krzysztof Kozlowski <krzk@kernel.org> wrote: > >> > >> On 15/08/2026 19:27, Ian Rogers wrote: > >>> Motivation & Background > >>> ============================== > >>> Currently, Linux lacks a standard perf-API compatible driver for the > >>> Broadcom AXI performance counter blocks on Raspberry Pi platforms. Prior > >>> out-of-tree vendor solutions relied on custom debugfs nodes and ad-hoc > >>> kthreads, preventing integration with standard Linux perf tooling (`perf stat`, > >>> `perf list`, etc.). > >>> > >> > >> Slow down! > >> > >> One version per 24h in normal cycle, not 13 patchsets in two days! (and > >> even rarer during the merge window) > > > > I'm familiar with perf-tools which was set up to follow BPF's example, > > both have -next repos. The -next repos merge into linux-next for > > testing. I agree that cherry-picking shouldn't be done straight into a > > pull request to Linus. Is there a principle against using -next and > > linux-next? > > How is this relevant? I don't think you read my comment. I think I addressed all the points in your email. The -next development is relevant because it directly addresses your point about not sending emails during a merge window. Except when the -next branch is becoming the main branch, merging patches at will is acceptable and you also benefit from the linux-next testing. Not sending patches 5 times a year during the two-week merge window, it is nice to think that this timetable fits everybody. Thanks, Ian > Best regards, > Krzysztof
On 17/08/2026 09:46, Ian Rogers wrote: > On Mon, Aug 17, 2026 at 12:26 AM Krzysztof Kozlowski <krzk@kernel.org> wrote: >> >> On 17/08/2026 08:57, Ian Rogers wrote: >>> On Sun, Aug 16, 2026 at 11:20 PM Krzysztof Kozlowski <krzk@kernel.org> wrote: >>>> >>>> On 15/08/2026 19:27, Ian Rogers wrote: >>>>> Motivation & Background >>>>> ============================== >>>>> Currently, Linux lacks a standard perf-API compatible driver for the >>>>> Broadcom AXI performance counter blocks on Raspberry Pi platforms. Prior >>>>> out-of-tree vendor solutions relied on custom debugfs nodes and ad-hoc >>>>> kthreads, preventing integration with standard Linux perf tooling (`perf stat`, >>>>> `perf list`, etc.). >>>>> >>>> >>>> Slow down! >>>> >>>> One version per 24h in normal cycle, not 13 patchsets in two days! (and >>>> even rarer during the merge window) >>> >>> I'm familiar with perf-tools which was set up to follow BPF's example, >>> both have -next repos. The -next repos merge into linux-next for >>> testing. I agree that cherry-picking shouldn't be done straight into a >>> pull request to Linus. Is there a principle against using -next and >>> linux-next? >> >> How is this relevant? I don't think you read my comment. > > I think I addressed all the points in your email. The -next > development is relevant because it directly addresses your point about > not sending emails during a merge window. Except when the -next branch > is becoming the main branch, merging patches at will is acceptable and > you also benefit from the linux-next testing. Not sending patches 5 > times a year during the two-week merge window, it is nice to think > that this timetable fits everybody. No, next is not relevant, because I am not even speaking about applying patches. I explained you the rule of one posting per 24h. Not 5! Best regards, Krzysztof
On Mon, Aug 17, 2026 at 12:54 AM Krzysztof Kozlowski <krzk@kernel.org> wrote: > > On 17/08/2026 09:46, Ian Rogers wrote: > > On Mon, Aug 17, 2026 at 12:26 AM Krzysztof Kozlowski <krzk@kernel.org> wrote: > >> > >> On 17/08/2026 08:57, Ian Rogers wrote: > >>> On Sun, Aug 16, 2026 at 11:20 PM Krzysztof Kozlowski <krzk@kernel.org> wrote: > >>>> > >>>> On 15/08/2026 19:27, Ian Rogers wrote: > >>>>> Motivation & Background > >>>>> ============================== > >>>>> Currently, Linux lacks a standard perf-API compatible driver for the > >>>>> Broadcom AXI performance counter blocks on Raspberry Pi platforms. Prior > >>>>> out-of-tree vendor solutions relied on custom debugfs nodes and ad-hoc > >>>>> kthreads, preventing integration with standard Linux perf tooling (`perf stat`, > >>>>> `perf list`, etc.). > >>>>> > >>>> > >>>> Slow down! > >>>> > >>>> One version per 24h in normal cycle, not 13 patchsets in two days! (and > >>>> even rarer during the merge window) > >>> > >>> I'm familiar with perf-tools which was set up to follow BPF's example, > >>> both have -next repos. The -next repos merge into linux-next for > >>> testing. I agree that cherry-picking shouldn't be done straight into a > >>> pull request to Linus. Is there a principle against using -next and > >>> linux-next? > >> > >> How is this relevant? I don't think you read my comment. > > > > I think I addressed all the points in your email. The -next > > development is relevant because it directly addresses your point about > > not sending emails during a merge window. Except when the -next branch > > is becoming the main branch, merging patches at will is acceptable and > > you also benefit from the linux-next testing. Not sending patches 5 > > times a year during the two-week merge window, it is nice to think > > that this timetable fits everybody. > > No, next is not relevant, because I am not even speaking about applying > patches. I explained you the rule of one posting per 24h. Not 5! So, when a patch series receives feedback from Sashiko indicating issues, it's natural for the author to address those issues and assume a human won't review further given the existing problems with the series. For example, this series with 9 versions posted in 2 day: https://lore.kernel.org/linux-perf-users/?q=%22perf+build%3A+Add+target+to+install+devel+packages+needed+to+build+perf%22 The author then requested human review once Sashiko was happy: https://lore.kernel.org/linux-perf-users/anyrSMrYWJz-TpBE@x1/ Sashiko has an option to embargo/delay its review for a specified number of hours and you can see the small number of mailing lists using it: https://github.com/sashiko-dev/sashiko/blob/main/sashiko.dev/email_policy.toml#L79 I sent the versions to the mailing lists suggested by get-maintainers.pl plus RPi, the Sashiko reviews were only coming from linux-perf-users as the others haven't set Sashiko up (their choice but seems not desirable imo). As already shown, it is normal on linux-perf-users to send >1 version per day, and I think this is preferable than having mailing list people waiting for the next patch version that fixes the already spotted issues. Waiting >24h per version wasn't something I could do with this series, I didn't have 2 weeks to spread the emails over, so appologies for the inconvenience receiving an email has caused you. If you have constructive comments, particularly on the device tree, I'm keen to incorporate them. At the moment I don't see anything actionable and the feedback has been similar in quality to feedback I've already chosen to stop paying attention to. Thanks, Ian > Best regards, > Krzysztof
On Sat, Aug 15, 2026 at 10:27 AM Ian Rogers <irogers@google.com> wrote:
>
> Motivation & Background
> ==============================
> Currently, Linux lacks a standard perf-API compatible driver for the
> Broadcom AXI performance counter blocks on Raspberry Pi platforms.
The patches now have no Sashiko reported issues except for the
deliberate backward compatibility #ifdef to support earlier than Linux
6.13 builds (I believe Raspberry Pi OS is currently using the LTS
v6.12 kernel and when they switch to v6.18 this can be removed):
https://sashiko.dev/#/patchset/20260815172712.50119-1-irogers%40google.com
If you are testing the patches the easiest way (imo) is to build them
outside the Linux tree doing something like:
# Remove the non-perf driver
$ sudo rmmod raspberrypi_axi_monitor
In a directory have rpi_axi_pmu.c (from the patches) and a Makefile with:
```
obj-m += rpi_axi_pmu.o
KDIR ?= /lib/modules/$(shell uname -r)/build
PWD := $(shell pwd)
default:
$(MAKE) -C $(KDIR) M=$(PWD) modules
clean:
$(MAKE) -C $(KDIR) M=$(PWD) clean
```
then run make and `sudo insmod rpi_axi_pmu.ko`.
On a Raspberry Pi 4 and earlier you should be able to test counters like:
```
$ sudo perf stat -e rpi_axi_pmu/h264_rtrans/,rpi_axi_pmu/h264_wtrans/
-a -- ffmpeg -vcodec h264_v4l2m2m -i test.mp4 -f null -
ffmpeg version 5.1.9-0+deb12u1+rpt1 Copyright (c) 2000-2026 the FFmpeg
developers
built with gcc 12 (Debian 12.2.0-14+deb12u1)
configuration: --prefix=/usr --extra-version=0+deb12u1+rpt1
--toolchain=hardened --incdir=/usr/include/aarch64-linux-gnu
--enable-gpl --disable-stripping --disable-mmal --enable-gnutls
--enable-ladspa --enable-libaom --enable-libass --enable-libbluray
--enable-libbs2b --enable-libcaca --enable-libcdio --enable-libcodec2
--enable-libdav1d --enable-libflite --enable-libfontconfig
--enable-libfreetype --enable-libfribidi --enable-libglslang
--enable-libgme --enable-libgsm --enable-libjack --enable-libmp3lame
--enable-libmysofa --enable-libopenjpeg --enable-libopenmpt
--enable-libopus --enable-libpulse --enable-librabbitmq
--enable-librist --enable-librubberband --enable-libshine
--enable-libsnappy --enable-libsoxr --enable-libspeex --enable-libsrt
--enable-libssh --enable-libsvtav1 --enable-libtheora
--enable-libtwolame --enable-libvidstab --enable-libvorbis
--enable-libvpx --enable-libwebp --enable-libx265 --enable-libxml2
--enable-libxvid --enable-libzimg --enable-libzmq --enable-libzvbi
--enable-lv2 --enable-omx --enable-openal --enable-opencl
--enable-opengl --enable-sand --enable-sdl2 --disable-sndio
--enable-libjxl --enable-neon --enable-v4l2-request --enable-libudev
--enable-epoxy --libdir=/usr/lib/aarch64-linux-gnu --arch=arm64
--enable-pocketsphinx --enable-librsvg --enable-libdc1394
--enable-libdrm --enable-vout-drm --enable-libiec61883
--enable-chromaprint --enable-frei0r --enable-libx264
--enable-libplacebo --enable-librav1e --enable-shared
libavutil 57. 28.100 / 57. 28.100
libavcodec 59. 37.100 / 59. 37.100
libavformat 59. 27.100 / 59. 27.100
libavdevice 59. 7.100 / 59. 7.100
libavfilter 8. 44.100 / 8. 44.100
libswscale 6. 7.100 / 6. 7.100
libswresample 4. 7.100 / 4. 7.100
libpostproc 56. 6.100 / 56. 6.100
Input #0, mov,mp4,m4a,3gp,3g2,mj2, from 'test.mp4':
Metadata:
major_brand : isom
minor_version : 512
compatible_brands: isomiso2avc1mp41
encoder : Lavf59.27.100
Duration: 00:00:05.00, start: 0.000000, bitrate: 204 kb/s
Stream #0:0[0x1](und): Video: h264 (High) (avc1 / 0x31637661),
yuv420p(tv, smpte170m, progressive), 640x480, 202 kb/s, SAR 1:1 DAR
4:3, 30 fps, 30 tbr, 15360 tbn (default)
Metadata:
handler_name : VideoHandler
vendor_id : [0][0][0][0]
encoder : Lavc59.37.100 h264_v4l2m2m
[h264_v4l2m2m @ 0x559bdbefb0] Using device /dev/video10
[h264_v4l2m2m @ 0x559bdbefb0] driver 'bcm2835-codec' on card
'bcm2835-codec-decode' in mplane mode
[h264_v4l2m2m @ 0x559bdbefb0] requesting formats: output=H264 capture=YU12
Stream mapping:
Stream #0:0 -> #0:0 (h264 (h264_v4l2m2m) -> wrapped_avframe (native))
Press [q] to stop, [?] for help
Output #0, null, to 'pipe:':
Metadata:
major_brand : isom
minor_version : 512
compatible_brands: isomiso2avc1mp41
encoder : Lavf59.27.100
Stream #0:0(und): Video: wrapped_avframe, yuv420p(tv,
smpte170m/bt470m/bt709, progressive), 640x480 [SAR 1:1 DAR 4:3],
q=2-31, 200 kb/s, 30 fps, 30 tbn (default)
Metadata:
handler_name : VideoHandler
vendor_id : [0][0][0][0]
encoder : Lavc59.37.100 wrapped_avframe
frame= 150 fps=0.0 q=-0.0 Lsize=N/A time=00:00:05.00 bitrate=N/A
speed=12.9x
video:69kB audio:0kB subtitle:0kB other streams:0kB global headers:0kB
muxing overhead: unknown
Performance counter stats for 'system wide':
429,721,472 Bytes rpi_axi_pmu/h264_rtrans/
163,431,488 Bytes rpi_axi_pmu/h264_wtrans/
2.152513625 seconds time elapsed
```
and on a Raspberry Pi 5 like:
```
$ sudo perf stat -e
rpi_axi_pmu/a76_dsu_l3_rtrans/,rpi_axi_pmu/a76_dsu_l3_wtrans/,rpi_axi_pmu/dma_l2_wtrans/
-a -- sh -c 'dd if=/dev/zero of=/dev/shm/test_ram bs=1M count=1000 &&
rm /dev/shm/test_ram'
1000+0 records in
1000+0 records out
1048576000 bytes (1.0 GB, 1000 MiB) copied, 0.255655 s, 4.1 GB/s
Performance counter stats for 'system wide':
2,656 Bytes rpi_axi_pmu/a76_dsu_l3_rtrans/
1,152 Bytes rpi_axi_pmu/a76_dsu_l3_wtrans/
576 Bytes rpi_axi_pmu/dma_l2_wtrans/
0.279860667 seconds time elapsed
```
Thanks,
Ian
© 2016 - 2026 Red Hat, Inc.