[PATCH v12 00/25] Allow AET to use PMT as loadable module

Tony Luck posted 25 patches 1 week, 1 day ago
Documentation/filesystems/resctrl.rst      |  18 ++-
include/linux/arm_mpam.h                   |   3 -
include/linux/intel_vsec.h                 |  14 ++
include/linux/resctrl.h                    |  82 ++++++++++-
arch/x86/include/asm/processor.h           |   4 -
arch/x86/include/asm/resctrl.h             |  24 +---
arch/x86/kernel/cpu/resctrl/internal.h     |  23 +---
arch/x86/kernel/cpu/amd.c                  |   3 -
arch/x86/kernel/cpu/cpuid-deps.c           |   1 +
arch/x86/kernel/cpu/hygon.c                |   3 -
arch/x86/kernel/cpu/intel.c                |   7 -
arch/x86/kernel/cpu/resctrl/core.c         | 150 ++++++++++++---------
arch/x86/kernel/cpu/resctrl/intel_aet.c    | 114 ++++++++++++++--
arch/x86/kernel/cpu/resctrl/monitor.c      |  71 ++++++----
drivers/platform/x86/intel/pmt/telemetry.c |  47 ++++++-
drivers/resctrl/mpam_resctrl.c             |  44 +++---
fs/resctrl/monitor.c                       | 119 +++++++++++-----
fs/resctrl/pseudo_lock.c                   |   6 +-
fs/resctrl/rdtgroup.c                      |  79 +++++++----
arch/x86/Kconfig                           |  15 +--
arch/x86/kernel/cpu/resctrl/Makefile       |   3 +-
21 files changed, 562 insertions(+), 268 deletions(-)
[PATCH v12 00/25] Allow AET to use PMT as loadable module
Posted by Tony Luck 1 week, 1 day ago
Requiring INTEL_PMT_TELEMETRY=y to enable AET is a functional workaround
to enable enumeration of Application Energy Telemetry (AET) events, but
unacceptable to many users. It results in increased configuration complexity,
increased kernel memory footprint and inability to patch problems by unloading
a module and loading an updated version.

Add a registration function to the AET code that can be used by
INTEL_PMT_TELEMETRY to provide the enumeration functions.

INTEL_PMT_TELEMETRY can be loaded/unloaded independently of
resctrl file system mount/unmount. Perform enumeration on
every mount and cleanup on every unmount.

Patch series based on v7.3-rc3

Signed-off-by: Tony Luck <tony.luck@intel.com>

Changes since v11:
Link: https://lore.kernel.org/all/20260831174421.13921-1-tony.luck@intel.com/

Patch 12 "Handle change in number of RMIDs on each mount"
	split into three parts (12, 13, 14 in this series)
Patch 14 "Enforce system RMID limit on AET event groups"
	massively simplified, and moved later (now patch 20).

See individual patches for changes to each part.

Tony Luck (25):
  x86/cpufeatures: Add missing CQM feature dependency
  x86/resctrl: Check if monitoring features are supported
  x86/resctrl: Enumerate monitor features in rdt_get_l3_mon_config()
  x86/resctrl: Apply Intel MBM quirk from rdt_get_l3_mon_config()
  x86/resctrl: Delete resctrl_cpu_detect()
  arm,x86,fs/resctrl: Replace architecture
    resctrl_arch_{alloc,mon}_capable()
  x86/resctrl: Update special case for Intel Haswell enumeration
  x86/resctrl: Delete rdt_alloc_capable and rdt_mon_capable
  fs/resctrl: Remove redundant calls to resctrl_{alloc,mon}_capable()
  x86/resctrl: Honor rdt=perf option to force enable AET perf events
  fs/resctrl: Add interface to disable a monitor event
  arm,x86,fs/resctrl: Allocate maximum needed rmid_ptrs[]
  arm,x86,fs/resctrl: Allocate right size for L3 monitor arrays
  fs/resctrl: Rebuild free RMID list on each mount
  x86,fs/resctrl: Handle systems where AET is the only resource
  x86/resctrl: Add PMT registration API for AET enumeration callbacks
  platform/x86/intel/pmt: Register enumeration functions with resctrl
  x86/resctrl: Use registered function pointers for AET enumeration
  arm,x86,fs/resctrl: Enumerate AET on every resctrl mount
  x86/resctrl: Enforce system RMID limit on AET
  x86/resctrl: Export interface to report telemetry unbind/remove
  platform/x86/intel/pmt: Inform resctrl when MMIO maps are being
    removed
  x86/resctrl: Require 64-bit x86 for resctrl support
  x86/resctrl: Simplify Kconfig options for resctrl
  x86,fs/resctrl: Document telemetry mount timing caveat

 Documentation/filesystems/resctrl.rst      |  18 ++-
 include/linux/arm_mpam.h                   |   3 -
 include/linux/intel_vsec.h                 |  14 ++
 include/linux/resctrl.h                    |  82 ++++++++++-
 arch/x86/include/asm/processor.h           |   4 -
 arch/x86/include/asm/resctrl.h             |  24 +---
 arch/x86/kernel/cpu/resctrl/internal.h     |  23 +---
 arch/x86/kernel/cpu/amd.c                  |   3 -
 arch/x86/kernel/cpu/cpuid-deps.c           |   1 +
 arch/x86/kernel/cpu/hygon.c                |   3 -
 arch/x86/kernel/cpu/intel.c                |   7 -
 arch/x86/kernel/cpu/resctrl/core.c         | 150 ++++++++++++---------
 arch/x86/kernel/cpu/resctrl/intel_aet.c    | 114 ++++++++++++++--
 arch/x86/kernel/cpu/resctrl/monitor.c      |  71 ++++++----
 drivers/platform/x86/intel/pmt/telemetry.c |  47 ++++++-
 drivers/resctrl/mpam_resctrl.c             |  44 +++---
 fs/resctrl/monitor.c                       | 119 +++++++++++-----
 fs/resctrl/pseudo_lock.c                   |   6 +-
 fs/resctrl/rdtgroup.c                      |  79 +++++++----
 arch/x86/Kconfig                           |  15 +--
 arch/x86/kernel/cpu/resctrl/Makefile       |   3 +-
 21 files changed, 562 insertions(+), 268 deletions(-)


base-commit: fd73f4a6659897191fa0d40695fe370925dd3780
-- 
2.55.0
Re: [PATCH v12 00/25] Allow AET to use PMT as loadable module
Posted by Luck, Tony 1 week ago
On Wed, Sep 16, 2026 at 04:12:55PM -0700, Tony Luck wrote:
> Requiring INTEL_PMT_TELEMETRY=y to enable AET is a functional workaround
> to enable enumeration of Application Energy Telemetry (AET) events, but
> unacceptable to many users. It results in increased configuration complexity,
> increased kernel memory footprint and inability to patch problems by unloading
> a module and loading an updated version.
> 
> Add a registration function to the AET code that can be used by
> INTEL_PMT_TELEMETRY to provide the enumeration functions.
> 
> INTEL_PMT_TELEMETRY can be loaded/unloaded independently of
> resctrl file system mount/unmount. Perform enumeration on
> every mount and cleanup on every unmount.

Sashiko report here:
	https://sashiko.dev/#/patchset/20260916231320.14502-1-tony.luck%40intel.com

Only issues in parts 11, 20, 21

Patch 11: [PATCH v12 11/25] fs/resctrl: Add interface to disable a monitor event

    This isn't a bug, but the kerneldoc for resctrl_disable_mon_event() appears
    to contradict the core safety invariant described in the commit message.
    The commit message states the architecture is responsible for calling this
    interface "only while resctrl is unmounted", but this documentation says not
    to disable an event that may be accessed while "unmounted".
    Could this lead to confusion for callers reading the header file? Should this
    say "while the file system is mounted" instead?

The kerneldoc comment is the better description here (supplied by
Reinette in the review of the v11 version of this series).

  https://lore.kernel.org/all/f9f3cb40-bc98-449d-a801-6af836900e76@intel.com/

With the intent of reminding developers that resctrl code may not be
idle just because the file system is not mounted. The limbo timer code
will continue to run until LLC cache occupancy counters reduce to the
threshold value to stop tracking.

Commit message could be updated to match if we need a new series.


Patch 20: [PATCH v12 20/25] x86/resctrl: Enforce system RMID limit on AET

    Does this code successfully enforce the system RMID limit on systems with SNC
    enabled as stated in the commit message?
    When SNC is enabled, the true maximum usable RMID limit is scaled down and
    available via resctrl_arch_system_max_rmid_idx(). By capping AET's num_rmid
    against pqr_assoc_num_rmid (the unscaled physical limit), the resulting limit
    could remain incorrectly large, continuing to display an unachievable value to
    users in info/PERF_PKG_MON/num_rmids.

This code is doing what I intend. Making sure that the value reported in
info/PERF_PKG_MON/num_rmids shows how many RMIDs can be supported by AET.

Perhaps the commit message could better explain this intent.


Patch 21: [PATCH v12 21/25] x86/resctrl: Export interface to report telemetry unbind/remove

    Can this result in an invalid cast for non-PCI devices?
    The PMT subsystem allows non-PCI devices (such as ACPI platform devices from
    pwrm_telemetry.c) to register endpoints. Using to_pci_dev() blindly here
    without verifying dev_is_pci() generates a bogus pointer for non-PCI devices.
    ...
    When this bogus pointer is passed into intel_vsec_get_mapping() and
    eventually to pci_match_id(), will it cause out-of-bounds memory reads or
    KASAN panics when dereferencing pdev->vendor and pdev->device?

The AET endpoints are always PCIe (enumeration uses the VSEC feature).

    Does dropping ep_lock here create a race condition?
    While ep_lock is dropped, stale endpoints still remain in the global
    telem_array list. A concurrent resctrl mount could invoke
    intel_pmt_get_regions_by_feature(), acquire the lock, and cache pointers to
    the MMIO resources of the devices currently being removed.
    When pmt_telem_remove() resumes and re-acquires the lock, it unmaps those
    regions. Won't the concurrent reader be left with validly cached but unmapped
    memory pointers, leading to a kernel panic when dereferenced by AET?

This is an existing issue in the pmt_telemetry driver. Scenario is a
race between a resctrl mount and an unbind of a device. The unbind gets
to pmt_telem_remove() but loses the race to acquire ep_lock to the mount
code calling intel_pmt_get_regions_by_feature(). All devices report
valid MMIO addresses and ep_lock is released then pmt_telem_remove()
invalidates the MMIO mappings for the device being unbound/removed.

Perhaps the telemetry driver should prevent removal of devices for the
interval from intel_pmt_get_regions_by_feature() to intel_pmt_put_feature_group()?

Can it do that?

-Tony
Re: [PATCH v12 00/25] Allow AET to use PMT as loadable module
Posted by Luck, Tony 1 week ago
On Thu, Sep 17, 2026 at 09:32:58AM -0700, Luck, Tony wrote:
>     Does dropping ep_lock here create a race condition?
>     While ep_lock is dropped, stale endpoints still remain in the global
>     telem_array list. A concurrent resctrl mount could invoke
>     intel_pmt_get_regions_by_feature(), acquire the lock, and cache pointers to
>     the MMIO resources of the devices currently being removed.
>     When pmt_telem_remove() resumes and re-acquires the lock, it unmaps those
>     regions. Won't the concurrent reader be left with validly cached but unmapped
>     memory pointers, leading to a kernel panic when dereferenced by AET?
> 
> This is an existing issue in the pmt_telemetry driver. Scenario is a
> race between a resctrl mount and an unbind of a device. The unbind gets
> to pmt_telem_remove() but loses the race to acquire ep_lock to the mount
> code calling intel_pmt_get_regions_by_feature(). All devices report
> valid MMIO addresses and ep_lock is released then pmt_telem_remove()
> invalidates the MMIO mappings for the device being unbound/removed.
> 
> Perhaps the telemetry driver should prevent removal of devices for the
> interval from intel_pmt_get_regions_by_feature() to intel_pmt_put_feature_group()?
> 
> Can it do that?

It looks like I can avoid making this worse if I add a new log "aet_mmio"lock"
and use that to protect against invalidation of the MMIO virtual
pointers. Then I don't need to drop and reacquire ep_lock.

-Tony