From nobody Mon Feb 9 10:27:30 2026 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by smtp.lore.kernel.org (Postfix) with ESMTP id C41CAC77B73 for ; Wed, 31 May 2023 04:07:07 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S234096AbjEaEHG (ORCPT ); Wed, 31 May 2023 00:07:06 -0400 Received: from lindbergh.monkeyblade.net ([23.128.96.19]:55648 "EHLO lindbergh.monkeyblade.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S234228AbjEaEGa (ORCPT ); Wed, 31 May 2023 00:06:30 -0400 Received: from foss.arm.com (foss.arm.com [217.140.110.172]) by lindbergh.monkeyblade.net (Postfix) with ESMTP id 2D7D01BE; Tue, 30 May 2023 21:05:56 -0700 (PDT) Received: from usa-sjc-imap-foss1.foss.arm.com (unknown [10.121.207.14]) by usa-sjc-mx-foss1.foss.arm.com (Postfix) with ESMTP id 30F4B15BF; Tue, 30 May 2023 21:06:41 -0700 (PDT) Received: from a077893.arm.com (unknown [10.163.73.163]) by usa-sjc-imap-foss1.foss.arm.com (Postfix) with ESMTPA id B980F3F6C4; Tue, 30 May 2023 21:05:50 -0700 (PDT) From: Anshuman Khandual To: linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, will@kernel.org, catalin.marinas@arm.com, mark.rutland@arm.com Cc: Anshuman Khandual , Mark Brown , James Clark , Rob Herring , Marc Zyngier , Suzuki Poulose , Peter Zijlstra , Ingo Molnar , Arnaldo Carvalho de Melo , linux-perf-users@vger.kernel.org Subject: [PATCH V11 10/10] arm64/perf: Implement branch records save on PMU IRQ Date: Wed, 31 May 2023 09:34:28 +0530 Message-Id: <20230531040428.501523-11-anshuman.khandual@arm.com> X-Mailer: git-send-email 2.25.1 In-Reply-To: <20230531040428.501523-1-anshuman.khandual@arm.com> References: <20230531040428.501523-1-anshuman.khandual@arm.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Precedence: bulk List-ID: X-Mailing-List: linux-kernel@vger.kernel.org Content-Type: text/plain; charset="utf-8" This modifies armv8pmu_branch_read() to concatenate live entries along with task context stored entries and then process the resultant buffer to create perf branch entry array for perf_sample_data. It follows the same principle like task sched out. Cc: Catalin Marinas Cc: Will Deacon Cc: Mark Rutland Cc: linux-arm-kernel@lists.infradead.org Cc: linux-kernel@vger.kernel.org Tested-by: James Clark Signed-off-by: Anshuman Khandual --- drivers/perf/arm_brbe.c | 75 +++++++++++++++++------------------------ 1 file changed, 30 insertions(+), 45 deletions(-) diff --git a/drivers/perf/arm_brbe.c b/drivers/perf/arm_brbe.c index 0678ebf0a896..e3efc1563111 100644 --- a/drivers/perf/arm_brbe.c +++ b/drivers/perf/arm_brbe.c @@ -693,41 +693,45 @@ void armv8pmu_branch_reset(void) isb(); } =20 -static bool capture_branch_entry(struct pmu_hw_events *cpuc, - struct perf_event *event, int idx) +static void brbe_regset_branch_entries(struct pmu_hw_events *cpuc, struct = perf_event *event, + struct brbe_regset *regset, int idx) { struct perf_branch_entry *entry =3D &cpuc->branches->branch_entries[idx]; - u64 brbinf =3D get_brbinf_reg(idx); - - /* - * There are no valid entries anymore on the buffer. - * Abort the branch record processing to save some - * cycles and also reduce the capture/process load - * for the user space as well. - */ - if (brbe_invalid(brbinf)) - return false; + u64 brbinf =3D regset[idx].brbinf; =20 perf_clear_branch_entry_bitfields(entry); if (brbe_record_is_complete(brbinf)) { - entry->from =3D get_brbsrc_reg(idx); - entry->to =3D get_brbtgt_reg(idx); + entry->from =3D regset[idx].brbsrc; + entry->to =3D regset[idx].brbtgt; } else if (brbe_record_is_source_only(brbinf)) { - entry->from =3D get_brbsrc_reg(idx); + entry->from =3D regset[idx].brbsrc; entry->to =3D 0; } else if (brbe_record_is_target_only(brbinf)) { entry->from =3D 0; - entry->to =3D get_brbtgt_reg(idx); + entry->to =3D regset[idx].brbtgt; } capture_brbe_flags(entry, event, brbinf); - return true; +} + +static void process_branch_entries(struct pmu_hw_events *cpuc, struct perf= _event *event, + struct brbe_regset *regset, int nr_regset) +{ + int idx; + + for (idx =3D 0; idx < nr_regset; idx++) + brbe_regset_branch_entries(cpuc, event, regset, idx); + + cpuc->branches->branch_stack.nr =3D nr_regset; + cpuc->branches->branch_stack.hw_idx =3D -1ULL; } =20 void armv8pmu_branch_read(struct pmu_hw_events *cpuc, struct perf_event *e= vent) { struct brbe_hw_attr *brbe_attr =3D (struct brbe_hw_attr *)cpuc->percpu_pm= u->private; + struct arm64_perf_task_context *task_ctx =3D event->pmu_ctx->task_ctx_dat= a; + struct brbe_regset live[BRBE_MAX_ENTRIES]; + int nr_live, nr_store; u64 brbfcr, brbcr; - int idx, loop1_idx1, loop1_idx2, loop2_idx1, loop2_idx2, count; =20 brbcr =3D read_sysreg_s(SYS_BRBCR_EL1); brbfcr =3D read_sysreg_s(SYS_BRBFCR_EL1); @@ -739,35 +743,16 @@ void armv8pmu_branch_read(struct pmu_hw_events *cpuc,= struct perf_event *event) write_sysreg_s(brbfcr | BRBFCR_EL1_PAUSED, SYS_BRBFCR_EL1); isb(); =20 - /* Determine the indices for each loop */ - loop1_idx1 =3D BRBE_BANK0_IDX_MIN; - if (brbe_attr->brbe_nr <=3D BRBE_BANK_MAX_ENTRIES) { - loop1_idx2 =3D brbe_attr->brbe_nr - 1; - loop2_idx1 =3D BRBE_BANK1_IDX_MIN; - loop2_idx2 =3D BRBE_BANK0_IDX_MAX; + nr_live =3D capture_brbe_regset(brbe_attr, live); + if (event->ctx->task) { + nr_store =3D task_ctx->nr_brbe_records; + nr_store =3D stitch_stored_live_entries(task_ctx->store, live, nr_store, + nr_live, brbe_attr->brbe_nr); + process_branch_entries(cpuc, event, task_ctx->store, nr_store); + task_ctx->nr_brbe_records =3D 0; } else { - loop1_idx2 =3D BRBE_BANK0_IDX_MAX; - loop2_idx1 =3D BRBE_BANK1_IDX_MIN; - loop2_idx2 =3D brbe_attr->brbe_nr - 1; - } - - /* Loop through bank 0 */ - select_brbe_bank(BRBE_BANK_IDX_0); - for (idx =3D 0, count =3D loop1_idx1; count <=3D loop1_idx2; idx++, count= ++) { - if (!capture_branch_entry(cpuc, event, idx)) - goto skip_bank_1; - } - - /* Loop through bank 1 */ - select_brbe_bank(BRBE_BANK_IDX_1); - for (count =3D loop2_idx1; count <=3D loop2_idx2; idx++, count++) { - if (!capture_branch_entry(cpuc, event, idx)) - break; + process_branch_entries(cpuc, event, live, nr_live); } - -skip_bank_1: - cpuc->branches->branch_stack.nr =3D idx; - cpuc->branches->branch_stack.hw_idx =3D -1ULL; process_branch_aborts(cpuc); =20 /* Unpause the buffer */ --=20 2.25.1