From nobody Fri Sep 25 22:19:33 2026 Received: from mail-pj1-f41.google.com (mail-pj1-f41.google.com [209.85.216.41]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1F2B0336884 for ; Tue, 8 Sep 2026 02:03:52 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.41 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788833036; cv=none; b=sEAoSGO3KrzdeQT5Oq1L2IS4mjlPo50s7oegcEtgVmaxm6MQQGNRY0YUYJIAhK+BfGHS6UVtgQuWLT9yRRZzc/NM/tfuwGZnzsh4uZvvugoQdjcVMEmLAbsY1C0lW2fEpFVj3wJM0xtLmPQSZvJ3gekbX6AxzpULvV2J9jEda/A= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788833036; c=relaxed/simple; bh=Y+fZFT4WfTgdxtJ4tDxLL2Ax0Z8zt4omfywrKp/y+8g=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=cHwm/3kmatz7gW+Q74L55RdNUuzu+Yqy4NBB4XACaOlB+dZqYn6so6Y0/HCCvuny7dnETbJ/fVpjw+fMbjf5tLK/dkScolNNdgJcNz3x5wqJYYWF61PnDv9ZACrorZEhga4JS2EMH8ocQ2TZR0902bNwVJCzKn+mjMEF1w3Hxpk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=sifive.com; spf=pass smtp.mailfrom=sifive.com; dkim=pass (2048-bit key) header.d=sifive.com header.i=@sifive.com header.b=fNJb0mN2; arc=none smtp.client-ip=209.85.216.41 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=sifive.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=sifive.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=sifive.com header.i=@sifive.com header.b="fNJb0mN2" Received: by mail-pj1-f41.google.com with SMTP id 98e67ed59e1d1-39647aa9d52so3896873a91.0 for ; Mon, 07 Sep 2026 19:03:52 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=sifive.com; s=google; t=1788833032; x=1789437832; darn=vger.kernel.org; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:from:to:cc:subject :date:message-id:reply-to:content-type; bh=4RC30WVkLLWYCjQiap4Av0KDiZuNXH0eLkJ0+mv6m/Y=; b=fNJb0mN203h9vIrHU3oEITsJV0/4DhWe08vbp7yY2dcyNckLP4ZqCfE2ybyW1ZcPnd qGqwacJDYtgiy0zILrFbghLW5bK6TxlMsYBjzx5qnX9CYMKDf1iRaqU2aYKEkhja4Cum 2doStopowyJIR8rJsZnbKQ2xl+UsVbV8p9DOtPGU4fmgIfRuks1o+ndS7NP++IR0sr0T LMB/VUnnJ0GreIH6c53M8hnEJASW0p1xDRLmqKTP/5AocO3FB1khHtVOQ3o6DEgNDuUx +lTWoaeNjICYGns88Lv5ZwnCQn0S3woVTMWp6BhYBKQRDuUnh0zMQ30pUrwGkTk3IxUY El8A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788833032; x=1789437832; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=4RC30WVkLLWYCjQiap4Av0KDiZuNXH0eLkJ0+mv6m/Y=; b=dMWSpNpujWzKbIS7gYv/drb9Ttfdw5vuU8DGum3ESWDp0BOjAtFph7E4iqXb11zMJ8 axctygrkyoQAEz2FX4Rttdq8dJJb2bipm2Z0KjzQhXjTLrTx+Vo9zod6mSYwhz3vW8kH LKR/0xSxuSMbRc54YS7iGzowLmRuMFr6qQaeKL03ODrHTZdDMHcdel7q48Now0m/pWy2 vShG7OIF4RYPfluOL/lI901bLjDbDVgMAPl4vy4Ag38njvElJZbslh836VAs30+CN05P Ws2Yb73rTvch5SmVJMg0bUtfGf5TtyOeWvl32CGT8uXgd1Dd41llYGDLFLMXzSyTmwRF 9Nkg== X-Forwarded-Encrypted: i=1; AKwUvBwWmttOqp84ODdVt5VDXHEwEOcItNdq3nesM5RbgBqk/PR6LVClQyNUFgCKBx4crsnIRvujVjMIRj076BY=@vger.kernel.org X-Gm-Message-State: AFuF++n4p/H22WLSBuL45SIFqo4H4EexySOZM2TYXHShC71RXWohoAzZ 3NkKc0PngLDRAtN8+K5yZJXiHUg61AP0kn4Ec8PZTzyXVneVve+R9siBnBidex6aJ+M= X-Gm-Gg: AYBFou1hGYoLsesQvwEFwGEL9rx5wkvbpsIeFvDFTvVcz+Y+ov5vOZCi2EPhGrL1p0n EQxj9lNF9RMPOSj3mYNbvI+TwVvsgOmmF0KIlae10ExLiUulpgcj6KtbxK84sNpyPYW0k631wS5 u48IgAJ6VApYHN8P0RF3CfkkhvgU77j2HRJYfQ+m83GtxGQF0TltOxttw85zLayrVopcWiumPUA YQBgATPUbAPRj2F0Mz5DAMvbC3aax5Y0ymyt4d24oO0VeRb+9dQbbWnWySIXeeDAkJe3sH4Uzf8 HWgSgVnSuOoge/eFqt2VcPu6mfp2ejHM7cSxNDfjFhLgBawCC3obpjQp3ayLUwKaBNfQvgSyO1L 9zDoWIWSnCIFdHCO/O3y/WPXMwC0z90devV1deWJRMfZHNQfAnCj6NNPtfSreCT29QyNF+80M0L /kdEaywr0E1osrDMVCQpOgHiG74oOvMLx0ULRImbzSKuCAw5M5Uo9m+rGWRAsdluEGutAiDn6j X-Received: by 2002:a17:90a:ec84:b0:36b:b903:994 with SMTP id 98e67ed59e1d1-39b07f44dbdmr36625558a91.4.1788833032068; Mon, 07 Sep 2026 19:03:52 -0700 (PDT) Received: from sw04.internal.sifive.com ([4.53.31.132]) by smtp.gmail.com with ESMTPSA id a92af1059eb24-1434c745a09sm893671c88.7.2026.09.07.19.03.50 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 07 Sep 2026 19:03:51 -0700 (PDT) From: Zong Li To: tomasz.jeznach@linux.dev, joro@8bytes.org, will@kernel.org, robin.murphy@arm.com, pjw@kernel.org, palmer@dabbelt.com, aou@eecs.berkeley.edu, alex@ghiti.fr, mark.rutland@arm.com, andrew.jones@oss.qualcomm.com, guoren@kernel.org, david.laight.linux@gmail.com, zhangzhanpeng.jasper@bytedance.com, yang.yicong@picoheart.com, nutty.liu@hotmail.com, iommu@lists.linux.dev, linux-riscv@lists.infradead.org, linux-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org Cc: Zong Li , Chen Pei , Fangyu Yu Subject: [PATCH v9 1/2] drivers/perf: riscv-iommu: add risc-v iommu pmu driver Date: Mon, 7 Sep 2026 19:03:44 -0700 Message-ID: <20260908020347.1836653-2-zong.li@sifive.com> X-Mailer: git-send-email @GIT_VERSION@ In-Reply-To: <20260908020347.1836653-1-zong.li@sifive.com> References: <20260908020347.1836653-1-zong.li@sifive.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Add a new driver to support the RISC-V IOMMU PMU. This is an auxiliary device driver created by the parent RISC-V IOMMU driver. The performance monitor provides counters with filtering support to collect events for specific device ID/process ID, or GSCID/PSCID. The RISC-V IOMMU PMU separates the cycle counter from the event counters. The cycle counter is not associated with iohpmevt0, so a software-defined cycle event is required for the perf subsystem. The number and width of the counters are hardware-implemented and must be detected at runtime. Leave out all the dead cleanup code (i.e. .remove() operation) if the PMU driver is tied to the IOMMU driver and can never realistically be removed. PMU-related definitions are moved into the perf driver, where they are used exclusively. According to RISC-V IOMMU specification Chapter 6: Whether an 8 byte access to an IOMMU register is single-copy atomic is UNSPECIFIED. Use two separate 4 byte accesses for hardware compatibility. Tested-by: Chen Pei Tested-by: Fangyu Yu Reviewed-by: Nutty Liu Reviewed-by: Guo Ren (Alibaba DAMO Academy) Reviewed-by: Yicong Yang Suggested-by: David Laight Suggested-by: Guo Ren Link: https://lore.kernel.org/linux-riscv/20260618143634.7f3dd6c5@pumpkin/ Signed-off-by: Zong Li --- drivers/iommu/riscv/iommu-bits.h | 61 -- drivers/perf/Kconfig | 12 + drivers/perf/Makefile | 1 + drivers/perf/riscv_iommu_pmu.c | 1020 ++++++++++++++++++++++++++++++ 4 files changed, 1033 insertions(+), 61 deletions(-) create mode 100644 drivers/perf/riscv_iommu_pmu.c diff --git a/drivers/iommu/riscv/iommu-bits.h b/drivers/iommu/riscv/iommu-b= its.h index f2ef9bd3cde9..6b5de913a032 100644 --- a/drivers/iommu/riscv/iommu-bits.h +++ b/drivers/iommu/riscv/iommu-bits.h @@ -192,67 +192,6 @@ enum riscv_iommu_ddtp_modes { #define RISCV_IOMMU_IPSR_PMIP BIT(RISCV_IOMMU_INTR_PM) #define RISCV_IOMMU_IPSR_PIP BIT(RISCV_IOMMU_INTR_PQ) =20 -/* 5.19 Performance monitoring counter overflow status (32bits) */ -#define RISCV_IOMMU_REG_IOCOUNTOVF 0x0058 -#define RISCV_IOMMU_IOCOUNTOVF_CY BIT(0) -#define RISCV_IOMMU_IOCOUNTOVF_HPM GENMASK_ULL(31, 1) - -/* 5.20 Performance monitoring counter inhibits (32bits) */ -#define RISCV_IOMMU_REG_IOCOUNTINH 0x005C -#define RISCV_IOMMU_IOCOUNTINH_CY BIT(0) -#define RISCV_IOMMU_IOCOUNTINH_HPM GENMASK(31, 1) - -/* 5.21 Performance monitoring cycles counter (64bits) */ -#define RISCV_IOMMU_REG_IOHPMCYCLES 0x0060 -#define RISCV_IOMMU_IOHPMCYCLES_COUNTER GENMASK_ULL(62, 0) -#define RISCV_IOMMU_IOHPMCYCLES_OF BIT_ULL(63) - -/* 5.22 Performance monitoring event counters (31 * 64bits) */ -#define RISCV_IOMMU_REG_IOHPMCTR_BASE 0x0068 -#define RISCV_IOMMU_REG_IOHPMCTR(_n) (RISCV_IOMMU_REG_IOHPMCTR_BASE + ((_n= ) * 0x8)) - -/* 5.23 Performance monitoring event selectors (31 * 64bits) */ -#define RISCV_IOMMU_REG_IOHPMEVT_BASE 0x0160 -#define RISCV_IOMMU_REG_IOHPMEVT(_n) (RISCV_IOMMU_REG_IOHPMEVT_BASE + ((_n= ) * 0x8)) -#define RISCV_IOMMU_IOHPMEVT_EVENTID GENMASK_ULL(14, 0) -#define RISCV_IOMMU_IOHPMEVT_DMASK BIT_ULL(15) -#define RISCV_IOMMU_IOHPMEVT_PID_PSCID GENMASK_ULL(35, 16) -#define RISCV_IOMMU_IOHPMEVT_DID_GSCID GENMASK_ULL(59, 36) -#define RISCV_IOMMU_IOHPMEVT_PV_PSCV BIT_ULL(60) -#define RISCV_IOMMU_IOHPMEVT_DV_GSCV BIT_ULL(61) -#define RISCV_IOMMU_IOHPMEVT_IDT BIT_ULL(62) -#define RISCV_IOMMU_IOHPMEVT_OF BIT_ULL(63) - -/* Number of defined performance-monitoring event selectors */ -#define RISCV_IOMMU_IOHPMEVT_CNT 31 - -/** - * enum riscv_iommu_hpmevent_id - Performance-monitoring event identifier - * - * @RISCV_IOMMU_HPMEVENT_INVALID: Invalid event, do not count - * @RISCV_IOMMU_HPMEVENT_URQ: Untranslated requests - * @RISCV_IOMMU_HPMEVENT_TRQ: Translated requests - * @RISCV_IOMMU_HPMEVENT_ATS_RQ: ATS translation requests - * @RISCV_IOMMU_HPMEVENT_TLB_MISS: TLB misses - * @RISCV_IOMMU_HPMEVENT_DD_WALK: Device directory walks - * @RISCV_IOMMU_HPMEVENT_PD_WALK: Process directory walks - * @RISCV_IOMMU_HPMEVENT_S_VS_WALKS: First-stage page table walks - * @RISCV_IOMMU_HPMEVENT_G_WALKS: Second-stage page table walks - * @RISCV_IOMMU_HPMEVENT_MAX: Value to denote maximum Event IDs - */ -enum riscv_iommu_hpmevent_id { - RISCV_IOMMU_HPMEVENT_INVALID =3D 0, - RISCV_IOMMU_HPMEVENT_URQ =3D 1, - RISCV_IOMMU_HPMEVENT_TRQ =3D 2, - RISCV_IOMMU_HPMEVENT_ATS_RQ =3D 3, - RISCV_IOMMU_HPMEVENT_TLB_MISS =3D 4, - RISCV_IOMMU_HPMEVENT_DD_WALK =3D 5, - RISCV_IOMMU_HPMEVENT_PD_WALK =3D 6, - RISCV_IOMMU_HPMEVENT_S_VS_WALKS =3D 7, - RISCV_IOMMU_HPMEVENT_G_WALKS =3D 8, - RISCV_IOMMU_HPMEVENT_MAX =3D 9 -}; - /* 5.24 Translation request IOVA (64bits) */ #define RISCV_IOMMU_REG_TR_REQ_IOVA 0x0258 #define RISCV_IOMMU_TR_REQ_IOVA_VPN GENMASK_ULL(63, 12) diff --git a/drivers/perf/Kconfig b/drivers/perf/Kconfig index 245e7bb763b9..8cce6c2ea626 100644 --- a/drivers/perf/Kconfig +++ b/drivers/perf/Kconfig @@ -105,6 +105,18 @@ config RISCV_PMU_SBI full perf feature support i.e. counter overflow, privilege mode filtering, counter configuration. =20 +config RISCV_IOMMU_PMU + depends on RISCV || COMPILE_TEST + depends on RISCV_IOMMU + bool "RISC-V IOMMU Hardware Performance Monitor" + default y + help + Say Y if you want to use the RISC-V IOMMU performance monitor + implementation. The performance monitor is an optional hardware + feature, and whether it is actually enabled depends on IOMMU + hardware support. If the underlying hardware does not implement + the PMU, this option will have no effect. + config STARFIVE_STARLINK_PMU depends on ARCH_STARFIVE || COMPILE_TEST depends on 64BIT diff --git a/drivers/perf/Makefile b/drivers/perf/Makefile index eb8a022dad9a..90c75f3c0ac1 100644 --- a/drivers/perf/Makefile +++ b/drivers/perf/Makefile @@ -20,6 +20,7 @@ obj-$(CONFIG_QCOM_L3_PMU) +=3D qcom_l3_pmu.o obj-$(CONFIG_RISCV_PMU) +=3D riscv_pmu.o obj-$(CONFIG_RISCV_PMU_LEGACY) +=3D riscv_pmu_legacy.o obj-$(CONFIG_RISCV_PMU_SBI) +=3D riscv_pmu_sbi.o +obj-$(CONFIG_RISCV_IOMMU_PMU) +=3D riscv_iommu_pmu.o obj-$(CONFIG_STARFIVE_STARLINK_PMU) +=3D starfive_starlink_pmu.o obj-$(CONFIG_THUNDERX2_PMU) +=3D thunderx2_pmu.o obj-$(CONFIG_XGENE_PMU) +=3D xgene_pmu.o diff --git a/drivers/perf/riscv_iommu_pmu.c b/drivers/perf/riscv_iommu_pmu.c new file mode 100644 index 000000000000..3a1cf366a79f --- /dev/null +++ b/drivers/perf/riscv_iommu_pmu.c @@ -0,0 +1,1020 @@ +// SPDX-License-Identifier: GPL-2.0-only +/* + * Copyright (C) 2026 SiFive + * + * Authors + * Zong Li + */ + +#include +#include +#include +#include +#include +#include + +#include "../iommu/riscv/iommu.h" + +/* 5.19 Performance monitoring counter overflow status (32bits) */ +#define RISCV_IOMMU_REG_IOCOUNTOVF 0x0058 +#define RISCV_IOMMU_IOCOUNTOVF_CY BIT(0) +#define RISCV_IOMMU_IOCOUNTOVF_HPM GENMASK_ULL(31, 1) + +/* 5.20 Performance monitoring counter inhibits (32bits) */ +#define RISCV_IOMMU_REG_IOCOUNTINH 0x005C +#define RISCV_IOMMU_IOCOUNTINH_CY BIT(0) +#define RISCV_IOMMU_IOCOUNTINH_HPM GENMASK(31, 0) + +/* 5.21 Performance monitoring cycles counter (64bits) */ +#define RISCV_IOMMU_REG_IOHPMCYCLES 0x0060 +#define RISCV_IOMMU_IOHPMCYCLES_COUNTER GENMASK_ULL(62, 0) +#define RISCV_IOMMU_IOHPMCYCLES_OF BIT_ULL(63) +#define RISCV_IOMMU_REG_IOHPMCTR(_n) (RISCV_IOMMU_REG_IOHPMCYCLES + ((_n) = * 0x8)) + +/* 5.22 Performance monitoring event counters (31 * 64bits) */ +#define RISCV_IOMMU_REG_IOHPMCTR_BASE 0x0068 +#define RISCV_IOMMU_IOHPMCTR_COUNTER GENMASK_ULL(63, 0) + +/* 5.23 Performance monitoring event selectors (31 * 64bits) */ +#define RISCV_IOMMU_REG_IOHPMEVT_BASE 0x0160 +#define RISCV_IOMMU_REG_IOHPMEVT(_n) (RISCV_IOMMU_REG_IOHPMEVT_BASE + ((_n= ) * 0x8)) +#define RISCV_IOMMU_IOHPMEVT_EVENTID GENMASK_ULL(14, 0) +#define RISCV_IOMMU_IOHPMEVT_DMASK BIT_ULL(15) +#define RISCV_IOMMU_IOHPMEVT_PID_PSCID GENMASK_ULL(35, 16) +#define RISCV_IOMMU_IOHPMEVT_DID_GSCID GENMASK_ULL(59, 36) +#define RISCV_IOMMU_IOHPMEVT_PV_PSCV BIT_ULL(60) +#define RISCV_IOMMU_IOHPMEVT_DV_GSCV BIT_ULL(61) +#define RISCV_IOMMU_IOHPMEVT_IDT BIT_ULL(62) +#define RISCV_IOMMU_IOHPMEVT_OF BIT_ULL(63) +#define RISCV_IOMMU_IOHPMEVT_EVENT GENMASK_ULL(62, 0) + +/* The total number of counters is 31 event counters plus 1 cycle counter = */ +#define RISCV_IOMMU_HPM_COUNTER_NUM 32 + +/* Counter index 0 is the cycle counter, the event counters start at index= 1 */ +#define RISCV_IOMMU_HPM_CYCLE_IDX 0 + +static int cpuhp_state; + +/** + * enum riscv_iommu_hpmevent_id - Performance-monitoring event identifier + * + * @RISCV_IOMMU_HPMEVENT_CYCLE: Clock cycle counter + * @RISCV_IOMMU_HPMEVENT_URQ: Untranslated requests + * @RISCV_IOMMU_HPMEVENT_TRQ: Translated requests + * @RISCV_IOMMU_HPMEVENT_ATS_RQ: ATS translation requests + * @RISCV_IOMMU_HPMEVENT_TLB_MISS: TLB misses + * @RISCV_IOMMU_HPMEVENT_DD_WALK: Device directory walks + * @RISCV_IOMMU_HPMEVENT_PD_WALK: Process directory walks + * @RISCV_IOMMU_HPMEVENT_S_VS_WALKS: First-stage page table walks + * @RISCV_IOMMU_HPMEVENT_G_WALKS: Second-stage page table walks + * @RISCV_IOMMU_HPMEVENT_MAX: Value to denote maximum Event IDs + * + * The specification does not define an event ID for counting the + * number of clock cycles, meaning there is no associated 'iohpmevt0'. + * Event ID 0 is an invalid event and does not overlap with any valid + * event ID. Let's repurpose ID 0 as the cycle for perf, the cycle + * event is not actually written into any register, it serves solely + * as an identifier. + */ +enum riscv_iommu_hpmevent_id { + RISCV_IOMMU_HPMEVENT_CYCLE =3D 0, + RISCV_IOMMU_HPMEVENT_URQ =3D 1, + RISCV_IOMMU_HPMEVENT_TRQ =3D 2, + RISCV_IOMMU_HPMEVENT_ATS_RQ =3D 3, + RISCV_IOMMU_HPMEVENT_TLB_MISS =3D 4, + RISCV_IOMMU_HPMEVENT_DD_WALK =3D 5, + RISCV_IOMMU_HPMEVENT_PD_WALK =3D 6, + RISCV_IOMMU_HPMEVENT_S_VS_WALKS =3D 7, + RISCV_IOMMU_HPMEVENT_G_WALKS =3D 8, + RISCV_IOMMU_HPMEVENT_MAX =3D 9 +}; + +struct riscv_iommu_pmu { + struct pmu pmu; + struct hlist_node node; + void __iomem *reg; + int on_cpu; + unsigned int irq; + int numa_node; + unsigned int num_counters; + u64 cycle_cntr_mask; + u64 event_cntr_mask; + struct perf_event *events[RISCV_IOMMU_HPM_COUNTER_NUM]; + DECLARE_BITMAP(used_counters, RISCV_IOMMU_HPM_COUNTER_NUM); + /* Defers overflow processing to on_cpu when the interrupt lands elsewher= e */ + struct irq_work work; +}; + +#define to_riscv_iommu_pmu(p) (container_of(p, struct riscv_iommu_pmu, pmu= )) + +#define RISCV_IOMMU_PMU_ATTR_EXTRACTOR(_name, _mask) \ + static inline u32 get_##_name(struct perf_event *event) \ + { \ + return FIELD_GET(_mask, event->attr.config); \ + } \ + +RISCV_IOMMU_PMU_ATTR_EXTRACTOR(event, RISCV_IOMMU_IOHPMEVT_EVENTID); +RISCV_IOMMU_PMU_ATTR_EXTRACTOR(partial_matching, RISCV_IOMMU_IOHPMEVT_DMAS= K); +RISCV_IOMMU_PMU_ATTR_EXTRACTOR(pid_pscid, RISCV_IOMMU_IOHPMEVT_PID_PSCID); +RISCV_IOMMU_PMU_ATTR_EXTRACTOR(did_gscid, RISCV_IOMMU_IOHPMEVT_DID_GSCID); +RISCV_IOMMU_PMU_ATTR_EXTRACTOR(filter_pid_pscid, RISCV_IOMMU_IOHPMEVT_PV_P= SCV); +RISCV_IOMMU_PMU_ATTR_EXTRACTOR(filter_did_gscid, RISCV_IOMMU_IOHPMEVT_DV_G= SCV); +RISCV_IOMMU_PMU_ATTR_EXTRACTOR(filter_id_type, RISCV_IOMMU_IOHPMEVT_IDT); + +/* Formats */ +PMU_FORMAT_ATTR(event, "config:0-14"); +PMU_FORMAT_ATTR(partial_matching, "config:15"); +PMU_FORMAT_ATTR(pid_pscid, "config:16-35"); +PMU_FORMAT_ATTR(did_gscid, "config:36-59"); +PMU_FORMAT_ATTR(filter_pid_pscid, "config:60"); +PMU_FORMAT_ATTR(filter_did_gscid, "config:61"); +PMU_FORMAT_ATTR(filter_id_type, "config:62"); + +static struct attribute *riscv_iommu_pmu_formats[] =3D { + &format_attr_event.attr, + &format_attr_partial_matching.attr, + &format_attr_pid_pscid.attr, + &format_attr_did_gscid.attr, + &format_attr_filter_pid_pscid.attr, + &format_attr_filter_did_gscid.attr, + &format_attr_filter_id_type.attr, + NULL, +}; + +static const struct attribute_group riscv_iommu_pmu_format_group =3D { + .name =3D "format", + .attrs =3D riscv_iommu_pmu_formats, +}; + +/* Events */ +static ssize_t riscv_iommu_pmu_event_show(struct device *dev, + struct device_attribute *attr, + char *page) +{ + struct perf_pmu_events_attr *pmu_attr; + + pmu_attr =3D container_of(attr, struct perf_pmu_events_attr, attr); + + return sysfs_emit(page, "event=3D0x%02llx\n", pmu_attr->id); +} + +#define RISCV_IOMMU_PMU_EVENT_ATTR(name, id) \ + PMU_EVENT_ATTR_ID(name, riscv_iommu_pmu_event_show, id) + +static struct attribute *riscv_iommu_pmu_events[] =3D { + RISCV_IOMMU_PMU_EVENT_ATTR(cycle, RISCV_IOMMU_HPMEVENT_CYCLE), + RISCV_IOMMU_PMU_EVENT_ATTR(untranslated_req, RISCV_IOMMU_HPMEVENT_URQ), + RISCV_IOMMU_PMU_EVENT_ATTR(translated_req, RISCV_IOMMU_HPMEVENT_TRQ), + RISCV_IOMMU_PMU_EVENT_ATTR(ats_trans_req, RISCV_IOMMU_HPMEVENT_ATS_RQ), + RISCV_IOMMU_PMU_EVENT_ATTR(tlb_miss, RISCV_IOMMU_HPMEVENT_TLB_MISS), + RISCV_IOMMU_PMU_EVENT_ATTR(ddt_walks, RISCV_IOMMU_HPMEVENT_DD_WALK), + RISCV_IOMMU_PMU_EVENT_ATTR(pdt_walks, RISCV_IOMMU_HPMEVENT_PD_WALK), + RISCV_IOMMU_PMU_EVENT_ATTR(s_vs_pt_walks, RISCV_IOMMU_HPMEVENT_S_VS_WALKS= ), + RISCV_IOMMU_PMU_EVENT_ATTR(g_pt_walks, RISCV_IOMMU_HPMEVENT_G_WALKS), + NULL, +}; + +static const struct attribute_group riscv_iommu_pmu_events_group =3D { + .name =3D "events", + .attrs =3D riscv_iommu_pmu_events, +}; + +/* cpumask */ +static ssize_t riscv_iommu_cpumask_show(struct device *dev, + struct device_attribute *attr, + char *buf) +{ + struct riscv_iommu_pmu *pmu =3D to_riscv_iommu_pmu(dev_get_drvdata(dev)); + int on_cpu =3D pmu->on_cpu; + + /* + * riscv_iommu_pmu_offline_cpu() leaves on_cpu at -1 when it cannot + * find another online CPU to migrate to. Report an empty mask rather + * than feeding -1 to cpumask_of(), which indexes out of bounds. + */ + if (on_cpu < 0) + return sysfs_emit(buf, "%*pbl\n", cpumask_pr_args(cpu_none_mask)); + + return sysfs_emit(buf, "%*pbl\n", cpumask_pr_args(cpumask_of(on_cpu))); +} + +static struct device_attribute riscv_iommu_cpumask_attr =3D + __ATTR(cpumask, 0444, riscv_iommu_cpumask_show, NULL); + +static struct attribute *riscv_iommu_cpumask_attrs[] =3D { + &riscv_iommu_cpumask_attr.attr, + NULL +}; + +static const struct attribute_group riscv_iommu_pmu_cpumask_group =3D { + .attrs =3D riscv_iommu_cpumask_attrs, +}; + +static const struct attribute_group *riscv_iommu_pmu_attr_grps[] =3D { + &riscv_iommu_pmu_cpumask_group, + &riscv_iommu_pmu_format_group, + &riscv_iommu_pmu_events_group, + NULL, +}; + +/* + * Register access wrapper + * + * According to RISC-V IOMMU specification Chapter 6: + * A 4 byte access to an IOMMU register must be single-copy atomic. + * Whether an 8 byte access to an IOMMU register is single-copy atomic is = UNSPECIFIED + * + * Use two separate 4 byte accesses for hardware compatibility + */ +static u64 riscv_iommu_pmu_readq(void __iomem *addr) +{ + return hi_lo_readq(addr); +} + +static void riscv_iommu_pmu_writeq(u64 value, void __iomem *addr) +{ + hi_lo_writeq(value, addr); +} + +/* PMU Operations */ +static void riscv_iommu_pmu_set_counter(struct riscv_iommu_pmu *pmu, u32 i= dx, + u64 value) +{ + u64 counter_mask =3D idx ? pmu->event_cntr_mask : pmu->cycle_cntr_mask; + + riscv_iommu_pmu_writeq(value & counter_mask, pmu->reg + RISCV_IOMMU_REG_I= OHPMCTR(idx)); +} + +/* + * As stated in the RISC-V IOMMU Specification, Chapter 6: + * Whether an 8 byte access to an IOMMU register is single-copy atomic + * is UNSPECIFIED, and such an access may appear, internally to the + * IOMMU, as if two separate 4 byte accesses -=E2=80=89first to the high h= alf + * and second to the low half=E2=80=89-=E2=80=89were performed + * + * To make sure the driver works correctly on different hardware, + * the software will always use two 4-byte access for the counter. + * + * This function implements the hi-lo-hi pattern to detect and handle + * wraparound during the read operation: + * 1. Read high half (hi) + * 2. Read low half (lo) + * 3. Read high half again (hi_again) + * + * If both reads of the high half agree, then the low half did not carry + * into the high half in between, so the two halves belong together. If + * they differ, the low half wrapped during the read and is re-read to + * pair it with the high half observed after the carry. A second carry + * cannot follow within these few register accesses, as that would + * require the counter to advance by another 2^32 in the meantime. + * + * Note that the comparison must be made between the two reads of the + * high half within this call. Comparing against a value cached from an + * earlier call cannot work: a carry is invisible to such a check + * whenever the cached low half happens to be smaller than the low half + * observed after the wrap. + */ +static u64 riscv_iommu_pmu_get_counter(struct riscv_iommu_pmu *pmu, u32 id= x) +{ + void __iomem *addr =3D pmu->reg + RISCV_IOMMU_REG_IOHPMCTR(idx); + u64 value, counter_mask =3D idx ? pmu->event_cntr_mask : pmu->cycle_cntr_= mask; + u32 hi, lo, hi_again; + + hi =3D readl(addr + 4); + lo =3D readl(addr); + hi_again =3D readl(addr + 4); + + if (hi_again !=3D hi) { + hi =3D hi_again; + lo =3D readl(addr); + } + + value =3D (((u64)hi << 32) | lo) & counter_mask; + + /* The bit 63 of cycle counter (i.e., idx =3D=3D 0) is OF bit */ + return idx ? value : (value & ~RISCV_IOMMU_IOHPMCYCLES_OF); +} + +static bool is_cycle_event(u64 event) +{ + return FIELD_GET(RISCV_IOMMU_IOHPMEVT_EVENTID, event) =3D=3D + RISCV_IOMMU_HPMEVENT_CYCLE; +} + +static void riscv_iommu_pmu_set_event(struct riscv_iommu_pmu *pmu, u32 idx, + u64 value) +{ + /* There is no associated IOHPMEVT0 for IOHPMCYCLES */ + if (is_cycle_event(value)) + return; + + /* Event counter start from idx 1 */ + riscv_iommu_pmu_writeq(FIELD_GET(RISCV_IOMMU_IOHPMEVT_EVENT, value), + pmu->reg + RISCV_IOMMU_REG_IOHPMEVT(idx - 1)); +} + +static void riscv_iommu_pmu_enable_counter(struct riscv_iommu_pmu *pmu, u3= 2 idx) +{ + void __iomem *addr =3D pmu->reg + RISCV_IOMMU_REG_IOCOUNTINH; + u32 value =3D readl(addr); + + writel(value & ~BIT(idx), addr); +} + +static void riscv_iommu_pmu_disable_counter(struct riscv_iommu_pmu *pmu, u= 32 idx) +{ + void __iomem *addr =3D pmu->reg + RISCV_IOMMU_REG_IOCOUNTINH; + u32 value =3D readl(addr); + + writel(value | BIT(idx), addr); +} + +static void riscv_iommu_pmu_clear_ovf(struct riscv_iommu_pmu *pmu, u32 idx) +{ + u64 value; + + /* Counter is disabled here, making it safe to read and write registers */ + if (idx =3D=3D RISCV_IOMMU_HPM_CYCLE_IDX) { + value =3D riscv_iommu_pmu_readq(pmu->reg + RISCV_IOMMU_REG_IOHPMCYCLES) & + ~RISCV_IOMMU_IOHPMCYCLES_OF; + riscv_iommu_pmu_writeq(value, pmu->reg + RISCV_IOMMU_REG_IOHPMCYCLES); + } else { + /* Event counter start from idx 1 */ + value =3D riscv_iommu_pmu_readq(pmu->reg + RISCV_IOMMU_REG_IOHPMEVT(idx = - 1)) & + ~RISCV_IOMMU_IOHPMEVT_OF; + riscv_iommu_pmu_writeq(value, pmu->reg + RISCV_IOMMU_REG_IOHPMEVT(idx - = 1)); + } +} + +static void riscv_iommu_pmu_start_all(struct riscv_iommu_pmu *pmu, u32 inh= ibit) +{ + writel(inhibit, pmu->reg + RISCV_IOMMU_REG_IOCOUNTINH); +} + +/* Returns the inhibit state prior to stopping, so callers can restore it = later */ +static u32 riscv_iommu_pmu_stop_all(struct riscv_iommu_pmu *pmu) +{ + void __iomem *addr =3D pmu->reg + RISCV_IOMMU_REG_IOCOUNTINH; + u32 inhibit =3D readl(addr); + + writel(GENMASK_U32(pmu->num_counters - 1, 0), addr); + + return inhibit; +} + +/* PMU APIs */ +static void riscv_iommu_pmu_set_period(struct perf_event *event) +{ + struct riscv_iommu_pmu *pmu =3D to_riscv_iommu_pmu(event->pmu); + struct hw_perf_event *hwc =3D &event->hw; + u64 counter_mask =3D hwc->idx ? pmu->event_cntr_mask : pmu->cycle_cntr_ma= sk; + u64 period; + + /* + * Limit the maximum period to prevent the counter value + * from overtaking the one we are about to program. + * In effect we are reducing max_period to account for + * interrupt latency (and we are being very conservative). + */ + period =3D counter_mask >> 1; + riscv_iommu_pmu_set_counter(pmu, hwc->idx, period); + local64_set(&hwc->prev_count, period); +} + +/* + * Tally @config against what the hardware implements: one cycle counter p= lus + * pmu->num_counters - 1 event counters. Returns false once the group woul= d need + * more of either than exist, so that groups which could never be schedule= d are + * rejected in ->event_init() instead of failing with -EAGAIN in ->add() f= orever. + */ +static bool riscv_iommu_pmu_claim_counter(struct riscv_iommu_pmu *pmu, u64= config, + unsigned int *nr_cycles, + unsigned int *nr_events) +{ + if (is_cycle_event(config)) + return ++(*nr_cycles) <=3D 1; + + return ++(*nr_events) <=3D pmu->num_counters - 1; +} + +static int riscv_iommu_pmu_event_init(struct perf_event *event) +{ + struct riscv_iommu_pmu *pmu =3D to_riscv_iommu_pmu(event->pmu); + struct hw_perf_event *hwc =3D &event->hw; + struct perf_event *sibling; + unsigned int nr_cycles =3D 0; + unsigned int nr_events =3D 0; + int on_cpu; + + if (event->attr.type !=3D event->pmu->type) + return -ENOENT; + + if (is_sampling_event(event)) + return -EOPNOTSUPP; + + if (event->cpu < 0) + return -EOPNOTSUPP; + + /* + * Reject event IDs this driver does not know about. Programming one + * into IOHPMEVT would be accepted by the hardware but would never + * count anything, which is indistinguishable from an idle counter. + */ + if (get_event(event) >=3D RISCV_IOMMU_HPMEVENT_MAX) + return -EINVAL; + + /* + * There is no IOHPMEVT register associated with IOHPMCYCLES, so none + * of the filtering fields can be programmed for the cycle event. + * Reject them here instead of counting unfiltered cycles behind the + * user's back. + */ + if (is_cycle_event(event->attr.config) && + (event->attr.config & ~RISCV_IOMMU_IOHPMEVT_EVENTID)) + return -EINVAL; + + /* + * All events are bound to the CPU the interrupt is affine to. That + * CPU is unset while no online CPU could be found for this PMU, and + * assigning -1 here would turn this into a task bound event, which is + * not something this PMU can serve. + */ + on_cpu =3D pmu->on_cpu; + if (on_cpu < 0) + return -ENODEV; + + event->cpu =3D on_cpu; + + hwc->idx =3D -1; + hwc->config =3D event->attr.config; + + /* + * Account for this event itself first. It has to be done before the + * check below, otherwise an event which is on its own would never be + * matched against the number of counters the hardware implements. + */ + if (!riscv_iommu_pmu_claim_counter(pmu, event->attr.config, + &nr_cycles, &nr_events)) + return -EINVAL; + + /* On its own, so there is no group to validate */ + if (event->group_leader =3D=3D event) + return 0; + + /* + * Software events never occupy a hardware counter, so they do not have + * to sit on this pmu and do not consume any of its budget. Anything + * else in the group does, starting with the leader. + */ + if (!is_software_event(event->group_leader)) { + /* A hardware leader has to share this pmu's counters */ + if (event->group_leader->pmu !=3D event->pmu) + return -EINVAL; + + if (!riscv_iommu_pmu_claim_counter(pmu, + event->group_leader->attr.config, + &nr_cycles, &nr_events)) + return -EINVAL; + } + + /* + * Then the rest of the group. This walks group_leader->sibling_list, + * which does not contain the event being initialised yet - hence + * accounting for it separately above. + */ + for_each_sibling_event(sibling, event->group_leader) { + if (is_software_event(sibling)) + continue; + + if (sibling->pmu !=3D event->pmu) + return -EINVAL; + + if (!riscv_iommu_pmu_claim_counter(pmu, sibling->attr.config, + &nr_cycles, &nr_events)) + return -EINVAL; + } + + return 0; +} + +static void riscv_iommu_pmu_update(struct perf_event *event) +{ + struct hw_perf_event *hwc =3D &event->hw; + struct riscv_iommu_pmu *pmu =3D to_riscv_iommu_pmu(event->pmu); + u64 delta, prev, now; + u32 idx =3D hwc->idx; + u64 counter_mask =3D idx ? pmu->event_cntr_mask : pmu->cycle_cntr_mask; + + do { + prev =3D local64_read(&hwc->prev_count); + now =3D riscv_iommu_pmu_get_counter(pmu, idx); + } while (local64_cmpxchg(&hwc->prev_count, prev, now) !=3D prev); + + delta =3D (now - prev) & counter_mask; + local64_add(delta, &event->count); +} + +static void riscv_iommu_pmu_start(struct perf_event *event, int flags) +{ + struct riscv_iommu_pmu *pmu =3D to_riscv_iommu_pmu(event->pmu); + struct hw_perf_event *hwc =3D &event->hw; + + if (WARN_ON_ONCE(!(event->hw.state & PERF_HES_STOPPED))) + return; + + if (flags & PERF_EF_RELOAD) + WARN_ON_ONCE(!(event->hw.state & PERF_HES_UPTODATE)); + + hwc->state =3D 0; + riscv_iommu_pmu_set_period(event); + riscv_iommu_pmu_set_event(pmu, hwc->idx, hwc->config); + riscv_iommu_pmu_enable_counter(pmu, hwc->idx); + + perf_event_update_userpage(event); +} + +static void riscv_iommu_pmu_stop(struct perf_event *event, int flags) +{ + struct riscv_iommu_pmu *pmu =3D to_riscv_iommu_pmu(event->pmu); + struct hw_perf_event *hwc =3D &event->hw; + int idx =3D hwc->idx; + + if (hwc->state & PERF_HES_STOPPED) + return; + + riscv_iommu_pmu_disable_counter(pmu, idx); + + if ((flags & PERF_EF_UPDATE) && !(hwc->state & PERF_HES_UPTODATE)) + riscv_iommu_pmu_update(event); + + hwc->state |=3D PERF_HES_STOPPED | PERF_HES_UPTODATE; +} + +static int riscv_iommu_pmu_add(struct perf_event *event, int flags) +{ + struct riscv_iommu_pmu *pmu =3D to_riscv_iommu_pmu(event->pmu); + struct hw_perf_event *hwc =3D &event->hw; + unsigned int num_counters =3D pmu->num_counters; + unsigned int idx; + + /* Reserve index zero for iohpmcycles */ + if (is_cycle_event(event->attr.config)) + idx =3D RISCV_IOMMU_HPM_CYCLE_IDX; + else + idx =3D find_next_zero_bit(pmu->used_counters, num_counters, 1); + + /* All event counters or cycle counter are in use */ + if (idx =3D=3D num_counters || pmu->events[idx]) + return -EAGAIN; + + set_bit(idx, pmu->used_counters); + + pmu->events[idx] =3D event; + hwc->idx =3D idx; + hwc->state =3D PERF_HES_STOPPED | PERF_HES_UPTODATE; + local64_set(&hwc->prev_count, 0); + + if (flags & PERF_EF_START) + riscv_iommu_pmu_start(event, flags); + + /* Propagate changes to the userspace mapping. */ + perf_event_update_userpage(event); + + return 0; +} + +static void riscv_iommu_pmu_read(struct perf_event *event) +{ + riscv_iommu_pmu_update(event); +} + +static void riscv_iommu_pmu_del(struct perf_event *event, int flags) +{ + struct riscv_iommu_pmu *pmu =3D to_riscv_iommu_pmu(event->pmu); + struct hw_perf_event *hwc =3D &event->hw; + int idx =3D hwc->idx; + + riscv_iommu_pmu_stop(event, PERF_EF_UPDATE); + pmu->events[idx] =3D NULL; + clear_bit(idx, pmu->used_counters); + + perf_event_update_userpage(event); +} + +/* + * Bind the events and the interrupt to the same CPU: the overflow handler + * relies on running on pmu->on_cpu to be serialised against the perf + * callbacks, so the two must not drift apart. Only consider CPUs which the + * irqchip actually accepts - committing on_cpu to a CPU which + * irq_set_affinity() then rejects would leave the handler permanently on = the + * wrong CPU with nothing left to correct it. + * + * cpumask_local_spread() walks the online CPUs in order of NUMA distance = from + * the iommu, so the closest usable one wins. + * + * Returns the chosen CPU, or nr_cpu_ids if none could be used. + */ +static unsigned int riscv_iommu_pmu_bind_cpu(struct riscv_iommu_pmu *pmu, + unsigned int skip_cpu) +{ + unsigned int cpu, i; + + for (i =3D 0; i < num_online_cpus(); i++) { + cpu =3D cpumask_local_spread(i, pmu->numa_node); + if (cpu =3D=3D skip_cpu) + continue; + if (!irq_set_affinity(pmu->irq, cpumask_of(cpu))) + return cpu; + } + + return nr_cpu_ids; +} + +static int riscv_iommu_pmu_online_cpu(unsigned int cpu, struct hlist_node = *node) +{ + struct riscv_iommu_pmu *iommu_pmu; + unsigned int target_cpu; + + iommu_pmu =3D hlist_entry_safe(node, struct riscv_iommu_pmu, node); + + if (READ_ONCE(iommu_pmu->on_cpu) !=3D -1) + return 0; + + target_cpu =3D riscv_iommu_pmu_bind_cpu(iommu_pmu, nr_cpu_ids); + if (target_cpu >=3D nr_cpu_ids) { + /* on_cpu stays unset, so a later callback tries again */ + pr_debug("failed to point irq %u at any online cpu\n", + iommu_pmu->irq); + return 0; + } + + WRITE_ONCE(iommu_pmu->on_cpu, target_cpu); + + return 0; +} + +static int riscv_iommu_pmu_offline_cpu(unsigned int cpu, struct hlist_node= *node) +{ + struct riscv_iommu_pmu *iommu_pmu; + unsigned int target_cpu; + + iommu_pmu =3D hlist_entry_safe(node, struct riscv_iommu_pmu, node); + + if (READ_ONCE(iommu_pmu->on_cpu) !=3D (int)cpu) + return 0; + + /* + * The last online CPU cannot be taken offline, and this callback runs + * before __cpu_disable() clears the outgoing CPU from cpu_online_mask, + * so there is always another online CPU to move to. + * + * Should that ever fail, unset on_cpu rather than leaving it pointing at + * the CPU which is going away. That keeps events from being bound to a + * dead CPU and lets riscv_iommu_pmu_online_cpu() pick again once a CPU + * comes back. + */ + target_cpu =3D riscv_iommu_pmu_bind_cpu(iommu_pmu, cpu); + if (WARN_ON_ONCE(target_cpu >=3D nr_cpu_ids)) { + WRITE_ONCE(iommu_pmu->on_cpu, -1); + } else { + WRITE_ONCE(iommu_pmu->on_cpu, target_cpu); + perf_pmu_migrate_context(&iommu_pmu->pmu, cpu, target_cpu); + } + + return 0; +} + +/* + * Must run on pmu->on_cpu with interrupts disabled. + * That is what serialises it against the perf callbacks, and is why none + * of the state below needs a lock. + */ +static void riscv_iommu_pmu_process_overflow(struct riscv_iommu_pmu *pmu) +{ + DECLARE_BITMAP(ovf_bitmap, BITS_PER_TYPE(u64)); + u32 ovf, idx, inhibit; + + inhibit =3D riscv_iommu_pmu_stop_all(pmu); + + ovf =3D readl(pmu->reg + RISCV_IOMMU_REG_IOCOUNTOVF); + if (ovf) { + bitmap_from_u64(ovf_bitmap, ovf); + for_each_set_bit(idx, ovf_bitmap, pmu->num_counters) { + struct perf_event *event =3D pmu->events[idx]; + + /* + * pmu->events[idx] only means the counter is allocated, + * not that it is counting. A counter which has not been + * started has no valid prev_count to compute a delta + * against, and one which has been stopped must not be + * reprogrammed here. The overflow bit still has to be + * cleared below in either case, including when the event + * was already removed by riscv_iommu_pmu_del(), + * otherwise the interrupt would stay pending forever. + */ + if (event && !(event->hw.state & PERF_HES_STOPPED)) { + riscv_iommu_pmu_update(event); + riscv_iommu_pmu_set_period(event); + } + + riscv_iommu_pmu_clear_ovf(pmu, idx); + } + } + + /* Clear performance monitoring interrupt pending bit */ + writel_relaxed(RISCV_IOMMU_IPSR_PMIP, pmu->reg + RISCV_IOMMU_REG_IPSR); + + riscv_iommu_pmu_start_all(pmu, inhibit); +} + +static void riscv_iommu_pmu_work(struct irq_work *work) +{ + struct riscv_iommu_pmu *pmu =3D container_of(work, struct riscv_iommu_pmu= , work); + + riscv_iommu_pmu_process_overflow(pmu); +} + +static irqreturn_t riscv_iommu_pmu_irq_handler(int irq, void *dev_id) +{ + struct riscv_iommu_pmu *pmu =3D (struct riscv_iommu_pmu *)dev_id; + int target_cpu; + + /* Check whether this interrupt is for PMU */ + if (!(readl_relaxed(pmu->reg + RISCV_IOMMU_REG_IPSR) & RISCV_IOMMU_IPSR_P= MIP)) + return IRQ_NONE; + + /* + * PCI MSI/MSI-X on IMSIC sets IRQCHIP_MOVE_DEFERRED, so + * irq_set_affinity() reports success while only recording the request + * and the move is applied later, in interrupt context. Until then the + * interrupt is still routed to the CPU the irqchip picked initially, + * which is not the CPU the events are bound to. + * + * Hand the counters over to that CPU instead of touching them here. + * Deliberately leave the hardware alone: the iommu only raises a new + * interrupt on a 0->1 transition of both the counter's overflow bit and + * ipsr.pmip, so clearing either one here would lose the overflow, and + * leaving them set means no further interrupt can arrive while the work + * is pending. + */ + target_cpu =3D READ_ONCE(pmu->on_cpu); + if (target_cpu !=3D smp_processor_id() && target_cpu >=3D 0 && + cpu_online(target_cpu)) { + irq_work_queue_on(&pmu->work, target_cpu); + return IRQ_HANDLED; + } + + /* + * Either this is the bound CPU, or there is no usable one to defer to. + * In the latter case handle it here anyway: nobody else is going to + * clear ipsr.pmip, and without that the pmu would stop reporting for + * good. + */ + riscv_iommu_pmu_process_overflow(pmu); + + return IRQ_HANDLED; +} + +static unsigned int riscv_iommu_pmu_get_irq_num(struct riscv_iommu_device = *iommu) +{ + /* Reuse ICVEC.CIV mask for all interrupt vectors mapping */ + int vec =3D (iommu->icvec >> (RISCV_IOMMU_INTR_PM * 4)) & RISCV_IOMMU_ICV= EC_CIV; + + return iommu->irqs[vec]; +} + +static int riscv_iommu_pmu_request_irq(struct auxiliary_device *auxdev, + struct riscv_iommu_device *iommu, + struct riscv_iommu_pmu *pmu) +{ + /* + * Bind the handler to the auxiliary device, which is the same devres + * scope that frees @pmu. Requesting it on the parent iommu device + * would keep the handler registered with a dangling dev_id once @pmu + * is freed, either on a later probe failure or on device removal. + * + * IRQF_SHARED is required because ICVEC maps the performance + * monitoring source onto one of the vectors the iommu driver already + * requested for its queues whenever fewer than RISCV_IOMMU_INTR_COUNT + * vectors are available. Both requesters have to agree on sharing, or + * this one is rejected with -EBUSY. IRQF_ONESHOT does not have to be + * matched by hand: devm_request_irq() adds IRQF_COND_ONESHOT, so this + * handler adopts whatever the first requester picked. + */ + return devm_request_irq(&auxdev->dev, pmu->irq, riscv_iommu_pmu_irq_handl= er, + IRQF_SHARED | IRQF_NOBALANCING, + dev_name(iommu->dev), pmu); +} + +static void riscv_iommu_pmu_flush_work(void *data) +{ + struct riscv_iommu_pmu *pmu =3D data; + + irq_work_sync(&pmu->work); +} + +static void riscv_iommu_pmu_remove_cpuhp_instance(void *data) +{ + struct riscv_iommu_pmu *pmu =3D data; + + cpuhp_state_remove_instance_nocalls(cpuhp_state, &pmu->node); +} + +static void riscv_iommu_pmu_do_unregister(void *data) +{ + struct riscv_iommu_pmu *pmu =3D data; + + perf_pmu_unregister(&pmu->pmu); +} + +static int riscv_iommu_pmu_probe(struct auxiliary_device *auxdev, + const struct auxiliary_device_id *id) +{ + struct riscv_iommu_device *iommu_dev =3D dev_get_platdata(&auxdev->dev); + struct riscv_iommu_pmu *iommu_pmu; + void __iomem *addr; + char *name; + int ret; + + iommu_pmu =3D devm_kzalloc(&auxdev->dev, sizeof(*iommu_pmu), GFP_KERNEL); + if (!iommu_pmu) + return -ENOMEM; + + iommu_pmu->reg =3D iommu_dev->reg; + + /* + * Counter number and width are hardware-implemented, detect them by + * writing 1s and reading back which bits stuck. + * + * The specification requires a minimum of one programmable event + * counter besides the cycles counter when capabilities.HPM is 1, + * which is the condition under which this device is created. So a + * compliant implementation always reports at least two counters, and + * both IOHPMCYCLES and the first IOHPMCTR are always present. A + * readback which says otherwise is non-compliant hardware and is + * rejected rather than worked around. + * + * The implemented counters are assumed to be consecutive, so that + * hweight32() of the IOCOUNTINH readback can be used as the bound on + * valid counter indices. + * + * The counter masks are also assumed to be a contiguous run of bits + * starting at bit 0, which is what riscv_iommu_pmu_update() relies on + * when it masks the difference of two samples to handle wraparound. + */ + addr =3D iommu_pmu->reg + RISCV_IOMMU_REG_IOCOUNTINH; + writel(RISCV_IOMMU_IOCOUNTINH_HPM, addr); + iommu_pmu->num_counters =3D hweight32(readl(addr)); + if (iommu_pmu->num_counters < 2) { + dev_err(&auxdev->dev, "hardware reports %u counter(s)\n", + iommu_pmu->num_counters); + return -ENODEV; + } + + /* Bit 63 of IOHPMCYCLES is the OF bit, not part of the counter */ + addr =3D iommu_pmu->reg + RISCV_IOMMU_REG_IOHPMCYCLES; + riscv_iommu_pmu_writeq(RISCV_IOMMU_IOHPMCYCLES_COUNTER, addr); + iommu_pmu->cycle_cntr_mask =3D riscv_iommu_pmu_readq(addr) & + RISCV_IOMMU_IOHPMCYCLES_COUNTER; + if (!iommu_pmu->cycle_cntr_mask) { + dev_err(&auxdev->dev, "cycles counter is not implemented\n"); + return -ENODEV; + } + + /* Assume the width of all event counters are the same */ + addr =3D iommu_pmu->reg + RISCV_IOMMU_REG_IOHPMCTR_BASE; + riscv_iommu_pmu_writeq(RISCV_IOMMU_IOHPMCTR_COUNTER, addr); + iommu_pmu->event_cntr_mask =3D riscv_iommu_pmu_readq(addr); + if (!iommu_pmu->event_cntr_mask) { + dev_err(&auxdev->dev, "event counter is not implemented\n"); + return -ENODEV; + } + + iommu_pmu->pmu =3D (struct pmu) { + .module =3D THIS_MODULE, + .parent =3D &auxdev->dev, + .task_ctx_nr =3D perf_invalid_context, + .event_init =3D riscv_iommu_pmu_event_init, + .add =3D riscv_iommu_pmu_add, + .del =3D riscv_iommu_pmu_del, + .start =3D riscv_iommu_pmu_start, + .stop =3D riscv_iommu_pmu_stop, + .read =3D riscv_iommu_pmu_read, + .attr_groups =3D riscv_iommu_pmu_attr_grps, + .capabilities =3D PERF_PMU_CAP_NO_EXCLUDE, + }; + + auxiliary_set_drvdata(auxdev, iommu_pmu); + + name =3D devm_kasprintf(&auxdev->dev, GFP_KERNEL, + "riscv_iommu_pmu_%u", auxdev->id); + if (!name) { + dev_err(&auxdev->dev, "Failed to create name riscv_iommu_pmu_%u\n", + auxdev->id); + return -ENOMEM; + } + + iommu_pmu->numa_node =3D dev_to_node(iommu_dev->dev); + iommu_pmu->irq =3D riscv_iommu_pmu_get_irq_num(iommu_dev); + + /* + * IRQ_WORK_INIT_HARD, not IRQ_WORK_INIT: without IRQ_WORK_HARD_IRQ, + * irq_work_queue_on() puts the work on a lazy list to be run by a + * thread on PREEMPT_RT, which would lose the interrupts-disabled + * context this driver relies on for serialisation. + */ + iommu_pmu->work =3D IRQ_WORK_INIT_HARD(riscv_iommu_pmu_work); + + /* + * Registered before the interrupt so that devres, which releases in + * reverse order, frees the interrupt first and only then waits for an + * in-flight work item. + */ + ret =3D devm_add_action_or_reset(&auxdev->dev, riscv_iommu_pmu_flush_work, + iommu_pmu); + if (ret) + return ret; + + ret =3D riscv_iommu_pmu_request_irq(auxdev, iommu_dev, iommu_pmu); + if (ret) { + dev_err(&auxdev->dev, "Failed to request irq %s: %d\n", name, ret); + return ret; + } + + /* + * Bind all events to the same cpu context to avoid race enabling. + * riscv_iommu_pmu_online_cpu() picks the CPU and sets the irq + * affinity for us once the instance is registered below. + */ + iommu_pmu->on_cpu =3D -1; + + ret =3D cpuhp_state_add_instance(cpuhp_state, &iommu_pmu->node); + if (ret) { + dev_err(&auxdev->dev, "Failed to register hotplug %s: %d\n", name, ret); + return ret; + } + + ret =3D devm_add_action_or_reset(&auxdev->dev, + riscv_iommu_pmu_remove_cpuhp_instance, + iommu_pmu); + if (ret) + return ret; + + ret =3D perf_pmu_register(&iommu_pmu->pmu, name, -1); + if (ret) { + dev_err(&auxdev->dev, "Failed to register %s: %d\n", name, ret); + return ret; + } + + ret =3D devm_add_action_or_reset(&auxdev->dev, + riscv_iommu_pmu_do_unregister, + iommu_pmu); + if (ret) + return ret; + + /* + * The PMU name only carries the aux dev id, not the iommu dev name, so + * find the iommu dev name here to map this PMU back to its iommu dev. + */ + dev_info(&auxdev->dev, "%s: Registered with %u counters (iommu %s)\n", + name, iommu_pmu->num_counters, dev_name(iommu_dev->dev)); + + return 0; +} + +static const struct auxiliary_device_id riscv_iommu_pmu_id_table[] =3D { + { .name =3D "riscv-iommu.pmu" }, + {} +}; +MODULE_DEVICE_TABLE(auxiliary, riscv_iommu_pmu_id_table); + +static struct auxiliary_driver iommu_pmu_driver =3D { + .driver =3D { + .suppress_bind_attrs =3D true, + }, + .probe =3D riscv_iommu_pmu_probe, + .id_table =3D riscv_iommu_pmu_id_table, +}; + +static int __init riscv_iommu_pmu_init(void) +{ + int ret; + + cpuhp_state =3D cpuhp_setup_state_multi(CPUHP_AP_ONLINE_DYN, + "perf/riscv/iommu:online", + riscv_iommu_pmu_online_cpu, + riscv_iommu_pmu_offline_cpu); + if (cpuhp_state < 0) + return cpuhp_state; + + ret =3D auxiliary_driver_register(&iommu_pmu_driver); + if (ret) + cpuhp_remove_multi_state(cpuhp_state); + + return ret; +} +module_init(riscv_iommu_pmu_init); + +MODULE_DESCRIPTION("RISC-V IOMMU PMU"); +MODULE_LICENSE("GPL"); --=20 2.43.7 From nobody Fri Sep 25 22:19:33 2026 Received: from mail-pg1-f171.google.com (mail-pg1-f171.google.com [209.85.215.171]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 45F0B33938F for ; Tue, 8 Sep 2026 02:03:54 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.171 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788833035; cv=none; b=tF4GN0vH5dUZrA3oRtzz6yJYkd6Ikx4qno1VgPnXjbWtof4FUWzQECV/hrk+6YW+smVhRdp2/HqzgAL2+6wB+T2YT6u6jvp3tkev4nN+ialPTfrq8kjTS/WD7JWW2DKpsNyAN3u9cXTkboTkBxAc5DZGOo9tLJsKCyHgRBXauM8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788833035; c=relaxed/simple; bh=1N0u4QGDF0E8zTrVUubLP7yatSB6NAmP0HgLOmukhmU=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=oQKk4RPFxeLptoNPb6K7I9bl/hKOS+aSyu8siO/Qp08GsDro6Y0AuB6NHmtHf6//el8U1AEJB4D5N6zA+mXQIHK0eKy9B1yZGJfOR9d1aYUHlUMxYt5U9P8gNJm2XHNzBIva01kONNfiVhG5RS00O2s+zA1X1blMjejIfbVnN/U= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=sifive.com; spf=pass smtp.mailfrom=sifive.com; dkim=pass (2048-bit key) header.d=sifive.com header.i=@sifive.com header.b=XUBC2h8q; arc=none smtp.client-ip=209.85.215.171 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=sifive.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=sifive.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=sifive.com header.i=@sifive.com header.b="XUBC2h8q" Received: by mail-pg1-f171.google.com with SMTP id 41be03b00d2f7-cc1a4c62804so3024319a12.3 for ; Mon, 07 Sep 2026 19:03:54 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=sifive.com; s=google; t=1788833033; x=1789437833; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=Db3wlFQ7tWb65H2cUA6T2ug5PcDOcaldDbKEP2DhiPU=; b=XUBC2h8qaUcTRJIHRC57Y/p1PkHDBZG9Fdaf6o9xyyxA/pNcxXHs36rnCsLd+3QnkL 4gx7D5+okQeuoCn7hB9bMYqrZdd+jB0vBGd3Gv9ynuGxAZCa3gIxQxO/mdmfV2F488Qk UEyTifsT6y7+IH0P1OHLEsDOBOmMKy0oGarzGyo8KBJ2/0/Y3oLOBqGDTnbvs+9IptB1 FdNYVvjUqAJmwDyhph5mM1nIcb5fYNxe+ipf0sE/vzC2ahhcaLBv4e2D3E62ks58YfiA Vu0ZvO0qsdYsLOtuTYEPil35KYBuGIrZ4EDIwV4nPg6HKJaFdroBmcWMeLzwlRYvLtlo Nqgg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788833033; x=1789437833; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=Db3wlFQ7tWb65H2cUA6T2ug5PcDOcaldDbKEP2DhiPU=; b=ji8a0qx5bFFKJxUkc4fZRhyOOHXPppXPKt9oJV1wWb/66i+TKycnXxAnl4TUn1zKeN J5G3skfPuZyyYhOWMQaFulDBY2fMp69P7UuJbIAAr8JjXhp1XJW4l44/NUj2j2q0sV9o 0CR0Zbwx/V7bLbA1tjptngMyhY5SY3vShzxLqqTcCqzbAskEMUn4kvYIxvAH+hvC/GS7 sBy6N401rzsX5e1nARhQ77Gaj/1cKbxdnYsVJNliZxsDX+PQ0kd+AWPyVUNrfbtuvEKM PkHvNOzqLDRmB5zXZ6FJKnePvGy8QV0/WMhAJ4/L+D7jT9qNsqcohbcVJ3hJfYvya0Ru W/wQ== X-Forwarded-Encrypted: i=1; AKwUvBx/ZMITaPQe8/IIJt87jJXYpyZ8BMoyPdBqzMMz3zPRiT7k7yTFAUbPZW3JgCedV6Ll7i9F+R1sCOzlVk4=@vger.kernel.org X-Gm-Message-State: AFuF++kYhkcivVobehPxJqQed82jFQSm5YzTqFMULUKyp6hHWA+b3vw2 JAZnnWWXvwb0NTiqPtPMshfiy1sPs8PERMx2z1jKNes+Kz6aROgxkn2Hwb2Wld3X/Mc= X-Gm-Gg: AYBFou1v5QG8SSHbnaMaMZlTQf0w04c2ZZrN19qt8FhwgLPTo/+Vy4SDinnYbu/7ULY akvIUD0A+cnKXtQt/jGC9NGiz9HGM/kqUXs1ipF0/LyYNqZjCeeZGsTX+/BZt6p68ORnzsWAULt cD0T9oG1AXSRyK9a+J5hx3bL4a1eQp+as0wrCGdE8WPlbv5oCMgDJT9l/KsumMkuPcwv3hy8nA7 pdlGD9u3Z0mtOjS4dQFcHWkgYuw+7MgSE8tPMsTVYaJQs99/+G6V7NdoRysaL8q8RW3xnGJisjo n/aY4BjWcpSxH/Fsf/Kw9h0+0niBsYCb+1HWHuLlmn65HrdLESztsGQuTolCP51pwIwMHMXSjEJ OR2zc+lSme0ZXk8VvdKDqVIW+WLwpYlhGO85nHnw0V2dwB2s9npnNZuL60DeIT3vUjsJ7yNj+TE Qt/b0n1921n81dMrPlTp4jS4yBPRU0+ot/Rpg7cJasvpolsBAPwLrps6ez+apjSbLTD/eX5iPy X-Received: by 2002:a17:90b:3d89:b0:396:5fce:8e24 with SMTP id 98e67ed59e1d1-39b260d2ea8mr34307454a91.4.1788833033541; Mon, 07 Sep 2026 19:03:53 -0700 (PDT) Received: from sw04.internal.sifive.com ([4.53.31.132]) by smtp.gmail.com with ESMTPSA id a92af1059eb24-1434c745a09sm893671c88.7.2026.09.07.19.03.52 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 07 Sep 2026 19:03:53 -0700 (PDT) From: Zong Li To: tomasz.jeznach@linux.dev, joro@8bytes.org, will@kernel.org, robin.murphy@arm.com, pjw@kernel.org, palmer@dabbelt.com, aou@eecs.berkeley.edu, alex@ghiti.fr, mark.rutland@arm.com, andrew.jones@oss.qualcomm.com, guoren@kernel.org, david.laight.linux@gmail.com, zhangzhanpeng.jasper@bytedance.com, yang.yicong@picoheart.com, nutty.liu@hotmail.com, iommu@lists.linux.dev, linux-riscv@lists.infradead.org, linux-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org Cc: Zong Li , Chen Pei , Fangyu Yu , Samuel Holland Subject: [PATCH v9 2/2] iommu/riscv: create a auxiliary device for HPM Date: Mon, 7 Sep 2026 19:03:45 -0700 Message-ID: <20260908020347.1836653-3-zong.li@sifive.com> X-Mailer: git-send-email @GIT_VERSION@ In-Reply-To: <20260908020347.1836653-1-zong.li@sifive.com> References: <20260908020347.1836653-1-zong.li@sifive.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Create an auxiliary device for HPM when the IOMMU supports a hardware performance monitor. Tested-by: Chen Pei Tested-by: Fangyu Yu Reviewed-by: Nutty Liu Reviewed-by: Guo Ren Reviewed-by: Yicong Yang Suggested-by: Samuel Holland Signed-off-by: Zong Li --- drivers/iommu/riscv/Kconfig | 1 + drivers/iommu/riscv/iommu.c | 37 +++++++++++++++++++++++++++++++++++++ 2 files changed, 38 insertions(+) diff --git a/drivers/iommu/riscv/Kconfig b/drivers/iommu/riscv/Kconfig index b86e5ab94183..8025bf0fb67f 100644 --- a/drivers/iommu/riscv/Kconfig +++ b/drivers/iommu/riscv/Kconfig @@ -10,6 +10,7 @@ config RISCV_IOMMU select GENERIC_PT select IOMMU_PT select IOMMU_PT_RISCV64 + select AUXILIARY_BUS help Support for implementations of the RISC-V IOMMU architecture that complements the RISC-V MMU capabilities, providing similar address diff --git a/drivers/iommu/riscv/iommu.c b/drivers/iommu/riscv/iommu.c index cec3ddd7ab10..7f619971bb70 100644 --- a/drivers/iommu/riscv/iommu.c +++ b/drivers/iommu/riscv/iommu.c @@ -14,6 +14,7 @@ =20 #include #include +#include #include #include #include @@ -48,6 +49,9 @@ static DEFINE_IDA(riscv_iommu_pscids); #define RISCV_IOMMU_MAX_PSCID (BIT(20) - 1) =20 +/* IOMMU PMU auxiliary device id allocation namespace. */ +static DEFINE_IDA(riscv_iommu_pmu_ida); + /* Device resource-managed allocations */ struct riscv_iommu_devres { void *addr; @@ -565,6 +569,36 @@ static irqreturn_t riscv_iommu_fltq_process(int irq, v= oid *data) return IRQ_HANDLED; } =20 +/* + * IOMMU Hardware performance monitor + */ +static void riscv_iommu_pmu_id_free(void *data) +{ + ida_free(&riscv_iommu_pmu_ida, (unsigned long)data); +} + +static int riscv_iommu_hpm_enable(struct riscv_iommu_device *iommu) +{ + struct auxiliary_device *auxdev; + int id, ret; + + id =3D ida_alloc(&riscv_iommu_pmu_ida, GFP_KERNEL); + if (id < 0) + return id; + + ret =3D devm_add_action_or_reset(iommu->dev, riscv_iommu_pmu_id_free, + (void *)(unsigned long)id); + if (ret) + return ret; + + auxdev =3D __devm_auxiliary_device_create(iommu->dev, "riscv-iommu", + "pmu", iommu, id); + if (!auxdev) + return -ENODEV; + + return 0; +} + /* Lookup and initialize device context info structure. */ static struct riscv_iommu_dc *riscv_iommu_get_dc(struct riscv_iommu_device= *iommu, unsigned int devid) @@ -1613,6 +1647,9 @@ int riscv_iommu_init(struct riscv_iommu_device *iommu) goto err_remove_sysfs; } =20 + if (iommu->caps & RISCV_IOMMU_CAPABILITIES_HPM) + riscv_iommu_hpm_enable(iommu); + return 0; =20 err_remove_sysfs: --=20 2.43.7