From nobody Mon Aug 24 04:16:54 2026 Received: from mail-pl1-f178.google.com (mail-pl1-f178.google.com [209.85.214.178]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 743B3215F42 for ; Sun, 16 Aug 2026 14:22:47 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.178 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786890168; cv=none; b=h18XhBtXzPKDlUm4b82zLjQbJS1G+I/2M9x25L0UqCPSG8Y5sPQtyL2pnPnR57mCrw1mDzUtHeVEDjMy98cTGEypuXm1eGovdlD1TR+CiI8l2KYxavBoGRl1ZyB/LzcjFV9886OVp2i0fD/qmTQx1TKTGnC38jXwavaWjSw3odE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786890168; c=relaxed/simple; bh=zEbOdsYLmiABDyfXCev+B7XhB/QfOxcoi0kPMyjH2LQ=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=s/Fi9YaUvfVP4Tl/HvFkitnNAbNUoUFeEABn6yYMv4qhSR7zH4EWD8z7bv8HX07WNE0nLpuWInTA71MZSJ82O0z+yk/UFcCBjS9axZSbAGV3Mo4naYU59mbCo24Gn46tj8pCp8GLBIcK5bSzZsvLfumTfQ8sj6dtxVd8moeHf38= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=cpyJ930B; arc=none smtp.client-ip=209.85.214.178 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="cpyJ930B" Received: by mail-pl1-f178.google.com with SMTP id d9443c01a7336-2cf50c6f235so31010195ad.0 for ; Sun, 16 Aug 2026 07:22:47 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786890166; x=1787494966; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=m31MQUkSQS2jVChso5vZVbxBylBeL+g3s9FxURuBxhg=; b=cpyJ930BSxhcCWkTSZ69W2ejP6gLdp9VSu7Qrh5xl0rLdfHH5doU1VjjSs5u60WiVv CGdMJkmd67X/dw1RmTGmoJZh1D73pFYtNz7PH7wYgppNjlIHG8WVtgZO1Pd6pJz/fkCP 2UJbirjrtzGTzecUxTCld856mX5JX1IjOnaYboWwM94THXWNH5AZgD33SpQHLsDV1SLa t4CxhOLJrZR+rpCv7PQsBua4v3TGlZVkrBmcO3RpaNhaQZhKw5JzENdKdsCgZ4p8WJcx Sbws5b1WitLnPDycNW/J0cFPgyhStH+qLR1Oboa2NBuh4lNoeKoWzoG1AIBylGEqzZQK lMIQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786890166; x=1787494966; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=m31MQUkSQS2jVChso5vZVbxBylBeL+g3s9FxURuBxhg=; b=Ia0lm29CCJxbOXpPI21MZO6Ol8ZX22JAlyREiwFP4lbe3+GGYqdqimNhsDX+4ijczg ngbjPT0oGbDvVJNnFljx376vL03Vn4bxeKudlC8CUziT/N8w7DACTA5VHFGRTd9+WXLm 2NUoMKSuh6L4j02KP2YJ5k/StNZGxsKgvBm6ivh35kdsI0nkQ7SnRojCxs/MGJNr+3eR IMi2gh0SaugMABw2KfWldSd4Taux/sW3s7EvMRFIAPFpMu+HOenitOc1kpopFcUKzH5I oH1sGxsjc3/IZ+BFKH+vAO4kZMcTrUG517ya+tWuls+EwM0xckeMda/0YPfBKlwuSAGA 8T+w== X-Gm-Message-State: AOJu0YwfnRMltFlOty98xBa16tuV+1dj5TrbnHrUgchXcbsfSDKfVnI5 m7rasQBoUyr1HRjU7QfjYSq5wLVVfaYgcTVcSU7/FiCp6iMy1Gicim4b X-Gm-Gg: AR+sD139Ixx4uj4PXB+dJqUTH0RPX6kOBP/zfJz0rnNj9XkJ9KGbe8mZavbjXSa4xdv 0z3MRDye/pmQui2kchSlmsxZ38h7AtTvc/nzS+5B88rPbaMGFu6bbFXA8uL8CPhfCpE2vmm/sjb TM3x6gaa1dufXlPXFV1YCF48C/jmcBdCOLir65nbzrrfUo09nxIe6hUQLNQKlRen1UIpdsbm1Pz qL5xue6mLyknrp7dXL2w5J9cY9+thJjmMq50da+w+2NYaQPJRKvwbDhNN8MREQsBZ6D21DVdq74 tVjW9mz40TK01ZPbBkXPp+CY6Wn7p/QwVwxD518asaXD/4bSYbEctuuiLq0D0m0AuscPLnrAGdu ll1IMyIv8l8Ydu8iMtfy+f7uUb8XpARvEsLBPPrJtvFHENK1RLKvLFRrNVpZ4OceZ4dJfnISVny mWgXOCl9KwzwoK0PuTzecsVORVuhdsYNCdFsYIso01F9UUXh7tGoZlrtnr6AgQ1X3q4NTG+0KC1 UX9r6o= X-Received: by 2002:a17:902:e80b:b0:2c9:fb11:1bf4 with SMTP id d9443c01a7336-2d3b08825e1mr196775075ad.7.1786890166204; Sun, 16 Aug 2026 07:22:46 -0700 (PDT) Received: from kernel.tail6741c6.ts.net ([116.128.244.169]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2d3aebc8199sm26102365ad.74.2026.08.16.07.22.38 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sun, 16 Aug 2026 07:22:44 -0700 (PDT) From: Kunwu Chan X-Google-Original-From: Kunwu Chan To: will@kernel.org, mark.rutland@arm.com, sj@kernel.org, akpm@linux-foundation.org, shuah@kernel.org, kunwu.chan@linux.dev Cc: linux-kernel@vger.kernel.org, linux-arm-kernel@lists.infradead.org, linux-perf-users@vger.kernel.org, damon@lists.linux.dev, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, Kunwu Chan , Lian Wang Subject: [RFC PATCH 1/4] mm/damon/perf: introduce AUX backend interface and Kconfig Date: Sun, 16 Aug 2026 22:22:18 +0800 Message-ID: <20260816142222.689624-2-kunwu.chan@linux.dev> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260816142222.689624-1-kunwu.chan@linux.dev> References: <20260816142222.689624-1-kunwu.chan@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Kunwu Chan Add the backend operations used by AUX trace-buffer PMUs, backend state to damon_perf_event, and stubs for configurations without AUX support. Add CONFIG_DAMON_PERF_AUX for the ARM SPE transport and a separate CONFIG_DAMON_PERF_SPE_KUNIT_TEST option. The current transport requires ARM SPE to be built into the kernel; it remains independent of the optional observability counters and tracepoints. Co-developed-by: Lian Wang (ProcessMission) Signed-off-by: Lian Wang (ProcessMission) Signed-off-by: Kunwu Chan --- include/linux/damon.h | 5 +++ mm/damon/Kconfig | 28 +++++++++++++++ mm/damon/perf/aux_backend.h | 72 +++++++++++++++++++++++++++++++++++++ 3 files changed, 105 insertions(+) create mode 100644 mm/damon/perf/aux_backend.h diff --git a/include/linux/damon.h b/include/linux/damon.h index c191c065b0e4..a27c5f6c459b 100644 --- a/include/linux/damon.h +++ b/include/linux/damon.h @@ -8,6 +8,7 @@ #ifndef _DAMON_H_ #define _DAMON_H_ =20 +#include #include #include #include @@ -127,6 +128,7 @@ struct damon_target { enum damon_report_source { DAMON_REPORT_SRC_PERF_OVERFLOW =3D 0, /* overflow_handler (IBS, PEBS) */ DAMON_REPORT_SRC_PAGE_FAULT, /* damon_report_page_fault() */ + DAMON_REPORT_SRC_PERF_AUX, /* AUX trace-buffer backend */ }; =20 /** @@ -538,6 +540,7 @@ struct damos_filter { struct damon_ctx; struct damon_target_lookup; struct damos; +struct damon_perf_backend_ops; =20 /** * struct damos_walk_control - Control damos_walk(). @@ -1056,6 +1059,8 @@ struct damon_perf_event_attr { struct damon_perf_event { struct damon_perf_event_attr attr; void *priv; + const struct damon_perf_backend_ops *ops; + cpumask_t aux_cpumask; struct list_head list; struct hlist_node hlist_node; bool init_complete; diff --git a/mm/damon/Kconfig b/mm/damon/Kconfig index 9fac286df124..e9fb62ec186f 100644 --- a/mm/damon/Kconfig +++ b/mm/damon/Kconfig @@ -149,4 +149,32 @@ config DAMON_PERF_OBSERVE =20 If unsure, say N. =20 + +config DAMON_PERF_AUX + bool "DAMON AUX trace-buffer backend support" + depends on DAMON + depends on PERF_EVENTS + depends on ARM_SPE_PMU=3Dy + default n + help + Enable the functional AUX trace-buffer transport for DAMON perf + events. This provides backend selection and drain scheduling for + ARM SPE, which does not deliver samples through an overflow + callback. ARM SPE must be built into the kernel. + + This transport is independent of CONFIG_DAMON_PERF_OBSERVE. + + If unsure, say N. + +config DAMON_PERF_SPE_KUNIT_TEST + bool "Test the DAMON perf SPE parser" if !KUNIT_ALL_TESTS + depends on DAMON_PERF_AUX && KUNIT=3Dy + default KUNIT_ALL_TESTS + help + Test the ARM SPE parser with byte-exact packet streams. The cases + cover valid records, truncation, alignment, error resynchronization + and consumed-length accounting. Results can be exposed through + debugfs when CONFIG_KUNIT_DEBUGFS is enabled. + + If unsure, say N. endmenu diff --git a/mm/damon/perf/aux_backend.h b/mm/damon/perf/aux_backend.h new file mode 100644 index 000000000000..70e6ac6771bc --- /dev/null +++ b/mm/damon/perf/aux_backend.h @@ -0,0 +1,72 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +#ifndef _DAMON_PERF_AUX_BACKEND_H +#define _DAMON_PERF_AUX_BACKEND_H + +#include +#include +#include +#include + +struct damon_ctx; +struct damon_perf_event; + +/** + * struct damon_perf_backend_ops - PMU-specific AUX backend operations + * @name: Human-readable backend name. + * @flags: Bitmask of DAMON_PERF_BACKEND_* flags. + * @match_pmu: Return true when this backend claims @perf_event. + * @init: Allocate per-event, per-CPU resources. + * @cleanup: Release resources initialized for one CPU. + * @arm: Prepare the AUX producer before perf_event_enable(). + * @disarm: Quiesce backend state after perf_event_disable(). + * @drain: Parse pending AUX data into DAMON access reports. + * + * All callbacks run in process context. The caller serializes resource + * lifetime against CPU hotplug and invokes @drain only for CPUs recorded + * in damon_perf_event::aux_cpumask. + */ +struct damon_perf_backend_ops { + const char *name; + u32 flags; + bool (*match_pmu)(struct perf_event *perf_event); + int (*init)(struct damon_perf_event *event, int cpu, + struct perf_event *perf_event); + void (*cleanup)(struct damon_perf_event *event, int cpu); + int (*arm)(struct damon_perf_event *event, int cpu); + void (*disarm)(struct damon_perf_event *event, int cpu); + unsigned int (*drain)(struct damon_perf_event *event, int cpu); +}; + +#define DAMON_PERF_BACKEND_AUX BIT(0) + +#ifdef CONFIG_DAMON_PERF_AUX +void damon_perf_aux_drain(struct damon_ctx *ctx); +int damon_perf_aux_register_backend(const struct damon_perf_backend_ops *o= ps); +const struct damon_perf_backend_ops * +damon_perf_aux_find_backend(struct perf_event *perf_event); +void damon_perf_aux_select(struct damon_perf_event *event, + struct perf_event *perf_event); +#else +static inline void damon_perf_aux_drain(struct damon_ctx *ctx) +{ +} + +static inline int +damon_perf_aux_register_backend(const struct damon_perf_backend_ops *ops) +{ + return -EOPNOTSUPP; +} + +static inline const struct damon_perf_backend_ops * +damon_perf_aux_find_backend(struct perf_event *perf_event) +{ + return NULL; +} + +static inline void damon_perf_aux_select(struct damon_perf_event *event, + struct perf_event *perf_event) +{ +} +#endif + +#endif /* _DAMON_PERF_AUX_BACKEND_H */ --=20 2.43.0 From nobody Mon Aug 24 04:16:54 2026 Received: from mail-pl1-f179.google.com (mail-pl1-f179.google.com [209.85.214.179]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BA7241C84D0 for ; Sun, 16 Aug 2026 14:22:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.179 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786890180; cv=none; b=ngkw8a0HWCwynK+kBvKenaczYQ2axqQi49WQ7+d7QYw1fRp5oy7TCTJxtbiO245PescT+cz9HIjvs84RrIpK7IChXm6yHgqcg40WaWHQXMduCxDIEaMr7l6W3CSGpH3EjG3E4hlI+qxf+NWk3TclcKebBaiTI+5ZnwgiCV6aWag= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786890180; c=relaxed/simple; bh=R3eB8Hh3ZhzNxvlGuoiM50K0wQne80n9XHAAXq4LuU8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=X8Oq6OWOlIWHZeUhw2F5va3XXpzWab4JWRYgN1U6gI65459dk9z+3OGFsFbL+8qTyXTLagFsLq+/OFs7GU9kAnpccH6TeqLMuqM7P8ttzCSnfhChM+pjmyVddGvFqUR/BbUPzHra3OGByn7TWEY1Xw+Ij52eCmVDfjINVKOVEwY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=mcL0jAGW; arc=none smtp.client-ip=209.85.214.179 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="mcL0jAGW" Received: by mail-pl1-f179.google.com with SMTP id d9443c01a7336-2cee9b74ee1so18742915ad.3 for ; Sun, 16 Aug 2026 07:22:57 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786890177; x=1787494977; darn=vger.kernel.org; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:from:to:cc:subject :date:message-id:reply-to:content-type; bh=XFvz6C60IHzA1UYwdcXz9Guq3Gbi/MTHeZWN6fp2F3A=; b=mcL0jAGW6NFvAYGCMuM83nh0vISi3y2HG+d+bA0GYJj4BZSXxWAX8O0Vs5Usw9ze3Y ohyytSMKCduI0z1zN1eDhBBkMmq0TU/aXWtJczxaCCxhYB4QpdXHiRqgIJW0Ix61MnkC /QnT2RHvQw3jGWo4sqvwLjm/1FaUrBFAHSMGUqaNARgi0naMjQ5uRgVNbwQUnl0Xx4B4 iBQrkcZeVRTPMkAgc3Z2RWWu7hrO909a/1HuR4WdQZJJQm2VnLvgwQLTuo4LW/ywnfE0 UgBzvTyai5a/37yDVojPva9acyowgIahqcnRkERtWcVaXIB1TZeWaygqYEfqGKQC/b4N Mlgg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786890177; x=1787494977; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=XFvz6C60IHzA1UYwdcXz9Guq3Gbi/MTHeZWN6fp2F3A=; b=WdvlfNrMrbhi4jGiTc5NQfirt6tO3O/WgTuKqNA68RB21xEW9WWsofbIvqyeqAF167 Hr1Tr2qS+ZrEeFtUO1/DUPAL1qDF+I2VgHg5RvyI9Y0bja63naMbmAVA/MGE/kMDVYJR zu2hVvGKXJ95RPdR7ljFX9DNtrDRKA0pHnpjOt9kVdgIKOQZlHgKK2GbAIuH8F7wQ5OS bpnfPKXYOc1c5VOR0L0hKg0SH4C7/KaI8359S7NV1uvqqjlW0XkHobQkPSgUkrFZVnQa yn67Kg5YTDQXtE1DqBZbvfF0hKOg6wHBL0b3Icowct+AU/WmHCF9nwRB+RekhGpygN8N YCVw== X-Gm-Message-State: AOJu0YwBfe8NB6frRRcn8opoGSfHlFrTE0hr4anT8BcRMRxSLllY8Cg6 DMMTqd1da0FuDEzDaNg+FgVaW82VkcgEIk93T12ETsX6/xaSRh+GaEBk X-Gm-Gg: AR+sD12itDGHd5ZY7sGTN6ife+8MjLe3dBLFUp2fF5/5PQkV9YOm4GullsnbJmsjY6v WnlseX5maRLRN80iX0qxx+19hhtGw11pR+HSOczzHR+lZ8THZSlu6gT+l/D+3jwUgjitcbf30ee 6bOWog09tq1NURB13E+RCs5GTcmS1b9T05bC0ObU+X1tvPqxsSKV7XDSQie6R1OyFihPhzDJzv/ aUij+iZjT9thYULFEgiNqZV1JmRrVyQhhirTx5CPl0dnxiUmwoZ153kbDVqTzvhxR899ivON5G/ XqKvCpqUyE4NjoDKuk25uj35RditehALpslOWNulUwgHR8qKrghBGFvN3RWEiL0dsIHvCT3agA9 YOpMEY8EkY77dwZHHmS/RczaYrH1LEf6jFEHjmvmRu6loKOk0U+edbtYe3kPizdgKZ1rRCOkRvW wGN7iI87e8cWin3BWZUQw/4kRFFcJvERFuEZmue3A1IPDAYCY9GyxKvUpWThg5bB78QNovwpvlj TgL1Vo= X-Received: by 2002:a17:902:fd90:b0:2c8:248a:5dbb with SMTP id d9443c01a7336-2d3b0d16339mr208743955ad.7.1786890176678; Sun, 16 Aug 2026 07:22:56 -0700 (PDT) Received: from kernel.tail6741c6.ts.net ([116.128.244.169]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2d3aebc8199sm26102365ad.74.2026.08.16.07.22.46 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sun, 16 Aug 2026 07:22:55 -0700 (PDT) From: Kunwu Chan X-Google-Original-From: Kunwu Chan To: will@kernel.org, mark.rutland@arm.com, sj@kernel.org, akpm@linux-foundation.org, shuah@kernel.org, kunwu.chan@linux.dev Cc: linux-kernel@vger.kernel.org, linux-arm-kernel@lists.infradead.org, linux-perf-users@vger.kernel.org, damon@lists.linux.dev, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, Kunwu Chan , Lian Wang Subject: [RFC PATCH 2/4] mm/damon/perf: add AUX trace-buffer PMU backend for ARM SPE Date: Sun, 16 Aug 2026 22:22:19 +0800 Message-ID: <20260816142222.689624-3-kunwu.chan@linux.dev> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260816142222.689624-1-kunwu.chan@linux.dev> References: <20260816142222.689624-1-kunwu.chan@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable From: Kunwu Chan Add an ARM SPE backend that owns a per-CPU AUX buffer, drains its packet stream from kdamond process context, and publishes synthesized access reports through DAMON's existing SPSC report ring. Drain AUX before consuming the SPSC ring on each monitoring interval. After disabling the events, run the same AUX-then-ring sequence once more so records finalized by perf_event_disable() are not lost. Decode CONTEXTIDR_EL1 as a sampled tid and resolve its tgid under RCU. Drop records whose context is missing or stale instead of assigning them to an arbitrary DAMON target. Retain incomplete trailing records across drains, guarantee progress for aligned ALIGNMENT packets, and reject unsupported extended packet classes. Match ARM SPE events by the PMU event_init callback, avoiding a fixed-size PMU registry and its device-lifetime problems. Roll back partially armed CPU events on failure and release per-CPU AUX state on all error paths. This depends on the perf AUX kernel-consumer API series. Co-developed-by: Lian Wang (ProcessMission) Signed-off-by: Lian Wang (ProcessMission) Signed-off-by: Kunwu Chan --- drivers/perf/arm_spe_pmu.c | 8 + include/linux/perf/arm_spe_pmu.h | 11 + mm/damon/core.c | 22 +- mm/damon/ops-common.h | 6 +- mm/damon/perf/Makefile | 5 + mm/damon/perf/aux_backend.c | 121 ++++++++ mm/damon/perf/spe_backend.c | 479 +++++++++++++++++++++++++++++++ mm/damon/perf/spe_parser.h | 109 +++++++ mm/damon/vaddr.c | 68 ++++- 9 files changed, 819 insertions(+), 10 deletions(-) create mode 100644 include/linux/perf/arm_spe_pmu.h create mode 100644 mm/damon/perf/aux_backend.c create mode 100644 mm/damon/perf/spe_backend.c create mode 100644 mm/damon/perf/spe_parser.h diff --git a/drivers/perf/arm_spe_pmu.c b/drivers/perf/arm_spe_pmu.c index dbd0da111639..50430342475c 100644 --- a/drivers/perf/arm_spe_pmu.c +++ b/drivers/perf/arm_spe_pmu.c @@ -27,6 +27,8 @@ #include #include #include +#include + #include #include #include @@ -1097,6 +1099,12 @@ static int arm_spe_pmu_perf_init(struct arm_spe_pmu = *spe_pmu) return perf_pmu_register(&spe_pmu->pmu, name, -1); } =20 +bool arm_spe_pmu_match(struct perf_event *perf_event) +{ + return perf_event->pmu->event_init =3D=3D arm_spe_pmu_event_init; +} +EXPORT_SYMBOL_GPL(arm_spe_pmu_match); + static void arm_spe_pmu_perf_destroy(struct arm_spe_pmu *spe_pmu) { perf_pmu_unregister(&spe_pmu->pmu); diff --git a/include/linux/perf/arm_spe_pmu.h b/include/linux/perf/arm_spe_= pmu.h new file mode 100644 index 000000000000..da5fb3d7b080 --- /dev/null +++ b/include/linux/perf/arm_spe_pmu.h @@ -0,0 +1,11 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +#ifndef _LINUX_PERF_ARM_SPE_PMU_H +#define _LINUX_PERF_ARM_SPE_PMU_H + +#include + +struct perf_event; + +bool arm_spe_pmu_match(struct perf_event *perf_event); + +#endif /* _LINUX_PERF_ARM_SPE_PMU_H */ diff --git a/mm/damon/core.c b/mm/damon/core.c index 92b21f9484c9..aded25624e2d 100644 --- a/mm/damon/core.c +++ b/mm/damon/core.c @@ -22,6 +22,7 @@ /* for damon_get_folio() used by node eligible memory metrics */ #include "ops-common.h" #include "perf/perf.h" +#include "perf/aux_backend.h" =20 #define CREATE_TRACE_POINTS #include @@ -1763,8 +1764,14 @@ static int damon_commit_perf_events(struct damon_ctx= *dst, * the kdamond runs. Arm now if we are committing into a * running ctx whose substrate is already armed. */ - if (dst->perf_events_active) - damon_perf_event_arm(new_event); + if (dst->perf_events_active) { + err =3D damon_perf_event_arm(new_event); + if (err) { + damon_perf_cleanup(dst, new_event); + kfree(new_event); + goto out; + } + } } list_add_tail(&new_event->list, &dst->perf_events); } @@ -4077,6 +4084,9 @@ static unsigned int kdamond_check_reported_accesses(s= truct damon_ctx *ctx) unsigned int i; unsigned int total_reports =3D 0, matched_reports =3D 0; =20 + /* AUX backends publish into the same ring consumed below. */ + damon_perf_aux_drain(ctx); + tbl =3D damon_build_target_lookup(ctx, &nr_targets); if (!tbl) { pr_warn_ratelimited( @@ -4194,8 +4204,10 @@ static int kdamond_fn(void *data) struct damon_perf_event *event; =20 WRITE_ONCE(ctx->perf_events_active, true); - list_for_each_entry(event, &ctx->perf_events, list) - damon_perf_event_arm(event); + list_for_each_entry(event, &ctx->perf_events, list) { + if (damon_perf_event_arm(event)) + goto done; + } } =20 if (ctx->ops.init) @@ -4323,7 +4335,7 @@ static int kdamond_fn(void *data) WRITE_ONCE(ctx->perf_events_active, false); list_for_each_entry(event, &ctx->perf_events, list) damon_perf_event_disarm(event); - /* Drain any in-flight reports queued before disarm took effect. */ + /* Final AUX drain and ring drain after perf_event_disable(). */ kdamond_check_reported_accesses(ctx); } damon_destroy_targets(ctx); diff --git a/mm/damon/ops-common.h b/mm/damon/ops-common.h index 35da400a67ec..7142ffa8878a 100644 --- a/mm/damon/ops-common.h +++ b/mm/damon/ops-common.h @@ -33,11 +33,12 @@ bool damos_ops_has_filter(struct damos *s); */ struct damon_perf { struct perf_event * __percpu *event; + void *aux_priv; }; =20 int damon_perf_init(struct damon_ctx *ctx, struct damon_perf_event *event); void damon_perf_cleanup(struct damon_ctx *ctx, struct damon_perf_event *ev= ent); -void damon_perf_event_arm(struct damon_perf_event *event); +int damon_perf_event_arm(struct damon_perf_event *event); void damon_perf_event_disarm(struct damon_perf_event *event); =20 #else /* !CONFIG_PERF_EVENTS */ @@ -53,8 +54,9 @@ static inline void damon_perf_cleanup(struct damon_ctx *c= tx, { } =20 -static inline void damon_perf_event_arm(struct damon_perf_event *event) +static inline int damon_perf_event_arm(struct damon_perf_event *event) { + return 0; } =20 static inline void damon_perf_event_disarm(struct damon_perf_event *event) diff --git a/mm/damon/perf/Makefile b/mm/damon/perf/Makefile index 150cbaa875fa..76857e443234 100644 --- a/mm/damon/perf/Makefile +++ b/mm/damon/perf/Makefile @@ -3,3 +3,8 @@ # Observability: per-CPU counters, tracepoints, debugfs perf_stats obj-$(CONFIG_DAMON_PERF_OBSERVE) +=3D damon-perf.o damon-perf-objs :=3D stats.o debugfs.o + +obj-$(CONFIG_DAMON_PERF_AUX) +=3D damon-perf-aux.o +damon-perf-aux-objs :=3D aux_backend.o spe_backend.o +damon-perf-aux-$(CONFIG_DAMON_PERF_SPE_KUNIT_TEST) +=3D \ + spe_parser_test.o diff --git a/mm/damon/perf/aux_backend.c b/mm/damon/perf/aux_backend.c new file mode 100644 index 000000000000..c6c55c7845a5 --- /dev/null +++ b/mm/damon/perf/aux_backend.c @@ -0,0 +1,121 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * DAMON Perf AUX Backend =E2=80=94 Generic Drain Scheduler + * + * Provides damon_perf_aux_drain() which is called by kdamond_fn() each + * tick before the SPSC ring drain. It iterates all perf events on the + * context and for any event whose backend ops carry DAMON_PERF_BACKEND_AU= X, + * invokes the per-CPU drain callback to parse PMU records and feed them + * into damon_report_access(). + * + * Also provides the auto-selection helper that assigns a backend to an + * event based on the PMU object of the created perf_event. + */ + +#include +#include +#include +#include + +#include "aux_backend.h" + +/* Registered backends (populated at initcall time). */ +#define AUX_BACKEND_MAX 4 + +static const struct damon_perf_backend_ops *aux_backends[AUX_BACKEND_MAX]; +static int nr_aux_backends; + +/** + * damon_perf_aux_register_backend - Register an AUX backend ops table + * @ops: Backend callbacks to register. + * + * Called at initcall time by each backend. Returns 0 on success, + * -ENOSPC if the static table is full. + */ +int damon_perf_aux_register_backend(const struct damon_perf_backend_ops *o= ps) +{ + if (nr_aux_backends >=3D AUX_BACKEND_MAX) + return -ENOSPC; + + aux_backends[nr_aux_backends++] =3D ops; + return 0; +} + +/** + * damon_perf_aux_find_backend - Find the backend that claims the PMU + * @perf_event: Created perf event whose PMU should be matched. + * + * Returns the ops table whose match_pmu() claims @perf_event->pmu + * (exact PMU object comparison), or NULL if no backend matches. + */ +const struct damon_perf_backend_ops * +damon_perf_aux_find_backend(struct perf_event *perf_event) +{ + int i; + + for (i =3D 0; i < nr_aux_backends; i++) { + if (aux_backends[i]->match_pmu && + aux_backends[i]->match_pmu(perf_event)) + return aux_backends[i]; + } + return NULL; +} + +/** + * damon_perf_aux_select - Auto-select a backend for @event + * @event: DAMON perf event that will own the selected backend. + * @perf_event: Created perf event used for capability and PMU matching. + * + * Called from damon_perf_cpu_online() after the first perf_event is + * created. Checks the PMU capabilities of the created event; if it has + * PERF_PMU_CAP_ITRACE, looks up a registered AUX backend that claims + * the PMU object of the created event. + * Overflow-handler PMUs (IBS, PEBS, generic counters) keep ops =3D=3D NUL= L. + */ +void damon_perf_aux_select(struct damon_perf_event *event, + struct perf_event *perf_event) +{ + if (event->ops) + return; /* already assigned */ + + if (!(perf_event->pmu->capabilities & PERF_PMU_CAP_ITRACE)) + return; /* not an ITRACE / AUX PMU */ + + event->ops =3D damon_perf_aux_find_backend(perf_event); +} + +/** + * damon_perf_aux_drain - Drain all AUX backends into the SPSC ring + * @ctx: DAMON context whose AUX events should be drained. + * + * Must be called BEFORE kdamond_check_reported_accesses() each tick + * so that freshly-parsed records are available for the ring drain. + * Also called at kdamond stop for a final flush. + */ +void damon_perf_aux_drain(struct damon_ctx *ctx) +{ + struct damon_perf_event *event; + int cpu; + + /* + * Hold the CPU hotplug read lock so that a concurrent CPU offline + * callback cannot free the per-CPU backend resources (st->win, + * AUX buffer) while drain is accessing them. The offline path + * runs under the write-side hotplug lock and clears aux_cpumask + * before freeing, so once cpus_read_lock() is held any CPU still + * in the mask has live resources. + */ + cpus_read_lock(); + list_for_each_entry(event, &ctx->perf_events, list) { + if (!event->ops || + !(event->ops->flags & DAMON_PERF_BACKEND_AUX)) + continue; + if (!event->ops->drain) + continue; + + /* Only CPUs with initialized AUX resources. */ + for_each_cpu(cpu, &event->aux_cpumask) + event->ops->drain(event, cpu); + } + cpus_read_unlock(); +} diff --git a/mm/damon/perf/spe_backend.c b/mm/damon/perf/spe_backend.c new file mode 100644 index 000000000000..42dd62cd0fcf --- /dev/null +++ b/mm/damon/perf/spe_backend.c @@ -0,0 +1,479 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * DAMON Perf - ARM SPE Backend + * + * Implements the damon_perf_backend_ops for ARM SPE (Statistical + * Profiling Extension). SPE delivers samples into an AUX trace buffer + * owned by the perf core: perf_event_setup_aux() allocates it at init, + * the arm_spe_pmu driver writes it through perf_aux_output_begin()/ + * perf_aux_output_end() (updating rb->aux_head), and each kdamond tick + * this backend copies the newly written window [aux_tail, aux_head) + * into a linear scratch buffer, parses the SPE record stream (same + * packet encoding as tools/perf/util/arm-spe-decoder), and feeds + * synthesized damon_access_report entries into the per-CPU SPSC ring + * via damon_report_access(). + * + * The AUX ring runs in non-overwrite streaming mode: when the buffer + * is full the PMU pauses itself, and this backend's drain both frees + * the space (advancing aux_tail) and re-enables the paused event. + * + * Parser and record layout live in spe_parser.h; the parsing function + * itself is kept in this module so the KUnit test (spe_parser_test.c) + * can exercise it without a perf PMU. + */ + +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "../ops-common.h" +#include "aux_backend.h" +#include "perf.h" +#include "spe_parser.h" + +/* + * Parse one SPE record from the linear drain window. + * + * st->win holds a window of st->win_size bytes starting at the stream + * position st->aux_tail. On success the cursor is advanced past the + * record. An incomplete trailing record is retained in the AUX ring so + * the producer can finish it after the consumer releases any complete + * prefix records. + * SPE records are padded with PAD packets by the driver + * (arm_spe_pmu_pad_buf), so a complete record followed by padding + * reaches the window end only after the last record was committed. + */ +int spe_parse_one_record(struct spe_parser_state *st, struct spe_record *r= ec) +{ + unsigned int pos =3D 0; + unsigned int record_start =3D 0; + u32 local_tid =3D 0; + bool have_tid =3D false; + bool record_started =3D false; + + memset(rec, 0, sizeof(*rec)); + + while (pos < st->win_size) { + u8 hdr =3D st->win[pos]; + u8 hdr1; + u64 payload; + int width, index; + + /* PAD - 1 byte, skip. */ + if (hdr =3D=3D SPE_HDR_PAD) { + pos++; + if (!record_started) + record_start =3D pos; + continue; + } + /* END - record terminator, no payload. */ + if (hdr =3D=3D SPE_HDR_END) { + record_started =3D true; + pos++; + goto record_done; + } + /* TIMESTAMP - record terminator, always 8-byte payload. */ + if (hdr =3D=3D SPE_HDR_TIMESTAMP) { + record_started =3D true; + if (pos + 1 + 8 > st->win_size) + goto truncated; + pos +=3D 1 + 8; + goto record_done; + } + /* EVENTS / DATA-SOURCE - payload is not needed by DAMON. */ + if ((hdr & SPE_HDR_MASK1) =3D=3D SPE_HDR_EVENTS || + (hdr & SPE_HDR_MASK1) =3D=3D SPE_HDR_SOURCE) { + record_started =3D true; + width =3D 1 << ((hdr >> 4) & 0x3); + if (pos + 1 + width > st->win_size) + goto truncated; + pos +=3D 1 + width; + continue; + } + /* CONTEXT / OP-TYPE / EXTENDED share MASK2. */ + if ((hdr & SPE_HDR_MASK2) =3D=3D SPE_HDR_CONTEXT || + (hdr & SPE_HDR_MASK2) =3D=3D SPE_HDR_OP_TYPE) { + record_started =3D true; + index =3D hdr & 0x3; + width =3D 1 << ((hdr >> 4) & 0x3); + pos +=3D 1; + } else if ((hdr & SPE_HDR_MASK2) =3D=3D SPE_HDR_EXTENDED) { + /* + * Extended header: a second byte carries the + * width and (for data packets) the upper index + * bits. hdr1 =3D=3D 0 is the ALIGNMENT pseudo-packet: + * consume bytes up to the next 2^(hdr[3:0]+1) + * aligned position, exactly like + * arm_spe_get_alignment(). hdr is the extended + * marker (0x20), so this alignment is 2 bytes. + */ + if (pos + 2 > st->win_size) + goto truncated; + hdr1 =3D st->win[pos + 1]; + if (hdr1 =3D=3D SPE_HDR1_ALIGNMENT) { + /* + * Alignment is relative to the original + * AUX stream position, not the scratch + * window. st->aux_tail is the absolute + * offset of the window start. + */ + unsigned long stream =3D st->aux_tail + pos; + unsigned int align =3D 1U << ((hdr & 0xf) + 1); + unsigned int skip =3D align - + (stream & (align - 1)); + + if (pos + skip > st->win_size) + goto truncated; + pos +=3D skip; + if (!record_started) + record_start =3D pos; + continue; + } + if ((hdr1 & SPE_HDR_MASK3) !=3D SPE_HDR_ADDRESS && + (hdr1 & SPE_HDR_MASK3) !=3D SPE_HDR_COUNTER) + goto bad_packet; + record_started =3D true; + index =3D ((hdr & 0x3) << 3) | (hdr1 & 0x7); + width =3D 1 << ((hdr1 >> 4) & 0x3); + hdr =3D hdr1; + pos +=3D 2; + } else if ((hdr & SPE_HDR_MASK3) =3D=3D SPE_HDR_ADDRESS || + (hdr & SPE_HDR_MASK3) =3D=3D SPE_HDR_COUNTER) { + record_started =3D true; + index =3D hdr & 0x7; + width =3D 1 << ((hdr >> 4) & 0x3); + pos +=3D 1; + } else { + goto bad_packet; + } + + if (pos + width > st->win_size) + goto truncated; + + /* Little-endian payload. */ + { + const u8 *src =3D st->win + pos; + int j; + + payload =3D 0; + for (j =3D width - 1; j >=3D 0; j--) + payload =3D (payload << 8) | src[j]; + } + pos +=3D width; + + if ((hdr & SPE_HDR_MASK2) =3D=3D SPE_HDR_CONTEXT) { + /* + * Bits 3-2 encode the context format: + * 0 =3D 32-bit CONTEXTIDR_EL1, + * 1 =3D 64-bit CONTEXTIDR_EL1 (FEAT_CONTEXTIDR_EL1_64). + * Both carry the task pid in the lower 32 bits. + */ + if (((hdr >> 2) & 0x3) <=3D 1) { + local_tid =3D (u32)payload; + have_tid =3D true; + } + } else if ((hdr & SPE_HDR_MASK2) =3D=3D SPE_HDR_OP_TYPE) { + /* LD/ST/ATOMIC class: payload bit 0 =3D store. */ + if ((index & SPE_OP_CLASS_MASK) =3D=3D SPE_OP_CLASS_LDST) + rec->is_write =3D !!(payload & SPE_OP_PKT_ST); + } else if ((hdr & SPE_HDR_MASK3) =3D=3D SPE_HDR_ADDRESS) { + if (index =3D=3D SPE_ADDR_DATA_VIRT) { + rec->va =3D payload & GENMASK_ULL(55, 0); + rec->have_addr =3D true; + } + } + /* EVENTS/COUNTER/TIMESTAMP payloads are ignored. */ + } + +truncated: + /* + * PAD and ALIGNMENT packets before the record are independently + * consumable. Keep the record itself in the AUX ring; the next + * drain copies it again together with newly produced bytes. + */ + st->aux_tail +=3D record_start; + st->bytes +=3D record_start; + return SPE_PARSE_NEED_MORE; + +record_done: + st->aux_tail +=3D pos; + st->bytes +=3D pos; + if (!rec->have_addr) + return SPE_PARSE_SKIP; + rec->tid =3D have_tid ? local_tid : 0; + st->records++; + return SPE_PARSE_REPORT; + +bad_packet: + /* Resync one byte forward: drop the offending byte. */ + pos++; + st->aux_tail +=3D pos; + st->bytes +=3D pos; + return SPE_PARSE_ERROR; +} + +/* ---- Report synthesis ---------------------------------------------- */ + +static void spe_submit(struct spe_record *rec, int cpu, + struct perf_event *perf_event) +{ + struct damon_access_report report =3D { + .vaddr =3D rec->va & PAGE_MASK, + .size =3D PAGE_SIZE, + .cpu =3D cpu, + .is_write =3D rec->is_write, +#ifdef CONFIG_DAMON_PERF_OBSERVE + .source =3D DAMON_REPORT_SRC_PERF_AUX, +#endif + }; + + /* + * CONTEXTIDR_EL1 carries the sampled task's pid, not its tgid. + * Resolve pid -> tgid here under RCU (no lifetime pin) so + * kdamond_check_reported_accesses() can match the report against + * DAMON's pid targets. A task that exits between sampling and + * this lookup yields no match (tgid =3D=3D 0), which mirrors perf's + * own CONTEXTIDR semantics; pid-reuse can misattribute a sample + * the same way it would in any hardware-context-based tool. + */ + if (rec->tid) { + struct task_struct *task; + + rcu_read_lock(); + task =3D find_task_by_vpid(rec->tid); + if (task) { + report.tid =3D rec->tid; + report.tgid =3D task_tgid_nr(task); + } + rcu_read_unlock(); + } + + /* + * A missing or stale CONTEXTID cannot be attributed safely. Do not + * turn it into an access by assigning an arbitrary DAMON target. + */ + if (!report.tgid) { + damon_perf_observe_miss(rec->va, cpu, + DAMON_REPORT_MISS_TGID); + return; + } + + /* reason 0 denotes a valid sample that is queued to the ring. */ + damon_perf_observe_sample(rec->va, 0, + 0, cpu, 0, 0, perf_event->attr.sample_type); + damon_report_access(&report); +} + +/* ---- Backend ops ---------------------------------------------------- */ + +static bool spe_match_pmu(struct perf_event *perf_event) +{ + return arm_spe_pmu_match(perf_event); +} + +static struct spe_parser_state *spe_state(struct damon_perf_event *event, + int cpu) +{ + struct damon_perf *perf =3D event->priv; + unsigned long addr =3D (unsigned long)perf->aux_priv + + per_cpu_offset(cpu); + + return (struct spe_parser_state *)addr; +} + +static int spe_backend_init(struct damon_perf_event *event, int cpu, + struct perf_event *perf_event) +{ + struct damon_perf *perf =3D event->priv; + struct spe_parser_state *st; + + int ret; + + /* + * Allocate the AUX buffer the PMU writes into. Must happen + * before the event is enabled (arm() ordering, see + * damon_perf_event_arm()). The buffer is non-overwrite + * streaming: the PMU pauses itself when it fills up and this + * backend re-enables it once the drain has freed space. + */ + ret =3D perf_event_setup_aux(perf_event, SPE_BUFFER_PAGES, 0); + if (ret) { + pr_warn_ratelimited("damon-perf: cpu %u aux setup failed: %d\n", + cpu, ret); + return ret; + } + + if (!perf->aux_priv) { + perf->aux_priv =3D alloc_percpu(struct spe_parser_state); + if (!perf->aux_priv) + return -ENOMEM; + } + + st =3D spe_state(event, cpu); + memset(st, 0, sizeof(*st)); + st->win =3D kzalloc(SPE_BUFFER_PAGES * PAGE_SIZE, GFP_KERNEL); + if (!st->win) + return -ENOMEM; + + return 0; +} + +static void spe_backend_cleanup(struct damon_perf_event *event, int cpu) +{ + struct spe_parser_state *st =3D spe_state(event, cpu); + + kfree(st->win); + st->win =3D NULL; +} + +static int spe_backend_arm(struct damon_perf_event *event, int cpu) +{ + struct damon_perf *perf =3D event->priv; + struct spe_parser_state *st =3D spe_state(event, cpu); + struct perf_event *perf_event; + unsigned long head; + + perf_event =3D *per_cpu_ptr(perf->event, cpu); + if (!perf_event) + return -ENODEV; + + /* + * Start consuming from the current head so data written before + * this session (e.g. a previous arm/disarm cycle) is discarded. + * The consumer cursor must be moved too: perf_aux_output_begin() + * computes free space from user_page->aux_tail. + */ + head =3D perf_event_aux_head(perf_event); + st->aux_tail =3D head; + if (perf_event_aux_tail_set(perf_event, head) < 0) + return -EIO; + st->records =3D 0; + st->bytes =3D 0; + st->armed =3D true; + return 0; +} + +static void spe_backend_disarm(struct damon_perf_event *event, int cpu) +{ + struct spe_parser_state *st =3D spe_state(event, cpu); + + /* + * Prevent spe_backend_drain() from calling perf_event_enable() + * during the final drain that follows perf_event_disable(). + * The resume-after-full logic is correct only while armed. + */ + st->armed =3D false; +} + +static unsigned int spe_backend_drain(struct damon_perf_event *event, int = cpu) +{ + struct damon_perf *perf =3D event->priv; + struct spe_parser_state *st =3D spe_state(event, cpu); + struct perf_event *perf_event; + unsigned long head, size, consumed; + long copied; + unsigned int drained =3D 0; + + perf_event =3D *per_cpu_ptr(perf->event, cpu); + if (!perf_event) + return 0; + + head =3D perf_event_aux_head(perf_event); + size =3D head - st->aux_tail; + /* + * The AUX head/tail are absolute cursors that can span many + * buffer sizes while the hardware runs. If the window exceeds + * the buffer size, skip the overwritten prefix (discarded data) + * and clamp to one buffer worth of data. + */ + if (size > SPE_BUFFER_PAGES * PAGE_SIZE) { + st->aux_tail =3D head - SPE_BUFFER_PAGES * PAGE_SIZE; + size =3D SPE_BUFFER_PAGES * PAGE_SIZE; + } + if (!size) + return 0; + + /* + * Linearize the (possibly wrapped) window into the scratch buffer. + * Do not parse unless the accessor copied the complete snapshot. + */ + copied =3D perf_event_aux_copy(perf_event, st->aux_tail, + st->aux_tail + size, st->win); + if (copied < 0 || (unsigned long)copied !=3D size) + return 0; + st->win_size =3D size; + + while (drained < SPE_BUFFER_MAX_RECORDS) { + struct spe_record rec; + unsigned long tail0 =3D st->aux_tail; + int ret; + + ret =3D spe_parse_one_record(st, &rec); + if (ret =3D=3D SPE_PARSE_NEED_MORE) + break; + + /* Drop the consumed prefix from the window. */ + consumed =3D st->aux_tail - tail0; + st->win_size -=3D consumed; + memmove(st->win, st->win + consumed, st->win_size); + + drained++; + if (ret =3D=3D SPE_PARSE_REPORT) + spe_submit(&rec, cpu, perf_event); + } + + /* + * Release the consumed space back to the AUX ring. If the tail + * advance fails, do not resume a paused producer. + */ + if (perf_event_aux_tail_set(perf_event, st->aux_tail) < 0) + return drained; + + /* + * ARM SPE stops its hardware when a non-overwrite ring is full, + * but leaves event->state ACTIVE. perf_event_enable() therefore + * cannot restart it by itself. Move the event through OFF after + * releasing space, then enable it again. Do this only for a ring + * snapshot that was exactly full and while the backend is armed; + * the final drain after disarm must not restart the producer. + */ + if (st->armed && size =3D=3D SPE_BUFFER_PAGES * PAGE_SIZE) { + perf_event_pause(perf_event, false); + perf_event_enable(perf_event); + } + + return drained; +} + +/* ---- Backend registration ------------------------------------------ */ + +static const struct damon_perf_backend_ops spe_backend_ops =3D { + .name =3D "arm_spe", + .flags =3D DAMON_PERF_BACKEND_AUX, + .match_pmu =3D spe_match_pmu, + .init =3D spe_backend_init, + .cleanup =3D spe_backend_cleanup, + .arm =3D spe_backend_arm, + .disarm =3D spe_backend_disarm, + .drain =3D spe_backend_drain, +}; + +static int __init spe_backend_initcall(void) +{ + int ret; + + ret =3D damon_perf_aux_register_backend(&spe_backend_ops); + if (ret) + pr_warn("damon-perf: SPE backend registration failed: %d\n", + ret); + return ret; +} +late_initcall(spe_backend_initcall); diff --git a/mm/damon/perf/spe_parser.h b/mm/damon/perf/spe_parser.h new file mode 100644 index 000000000000..ef9035ee7784 --- /dev/null +++ b/mm/damon/perf/spe_parser.h @@ -0,0 +1,109 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * ARM SPE Record Format Definitions + * + * Constants and parser interface for the byte stream produced by the + * Statistical Profiling Extension (ARM ARM DDI 0487, Chapter D8). + * Packet encodings and decode order match + * tools/perf/util/arm-spe-decoder/arm-spe-pkt-decoder.c. + */ + +#ifndef _DAMON_PERF_SPE_PARSER_H +#define _DAMON_PERF_SPE_PARSER_H + +#include + +/* AUX buffer geometry. */ +#define SPE_BUFFER_MIN_PAGES 2 /* arm_spe_pmu_setup_aux() minimum */ +#define SPE_BUFFER_PAGES 16 /* 64 KiB per-CPU trace buffer */ +#define SPE_BUFFER_MAX_RECORDS 8192 /* per-drain record budget */ + +/* Packet header masks/values (arm-spe-pkt-decoder.h). */ +#define SPE_HDR_MASK1 0xcf +#define SPE_HDR_MASK2 0xfc +#define SPE_HDR_MASK3 0xf8 +#define SPE_HDR_PAD 0x00 +#define SPE_HDR_END 0x01 +#define SPE_HDR_TIMESTAMP 0x71 +#define SPE_HDR_EVENTS 0x42 +#define SPE_HDR_SOURCE 0x43 +#define SPE_HDR_CONTEXT 0x64 +#define SPE_HDR_OP_TYPE 0x48 +#define SPE_HDR_EXTENDED 0x20 +#define SPE_HDR_ADDRESS 0xb0 +#define SPE_HDR_COUNTER 0x98 +#define SPE_HDR1_ALIGNMENT 0x00 + +/* Address packet index for the virtual data address. */ +#define SPE_ADDR_DATA_VIRT 2 + +/* OP-TYPE index bits[1:0] =3D=3D 1 selects the LD/ST/ATOMIC class. */ +#define SPE_OP_CLASS_MASK 0x3 +#define SPE_OP_CLASS_LDST 0x1 +#define SPE_OP_PKT_ST 0x1 + +/* + * Parser return values. The distinction between NEED_MORE and SKIP is + * essential: NEED_MORE means the window ran out mid-record and the record + * remains unconsumed, while SKIP means a complete record was consumed that + * has no virtual address. + */ +enum spe_parse_ret { + SPE_PARSE_NEED_MORE =3D 0, /* incomplete record retained */ + SPE_PARSE_REPORT =3D 1, /* complete record with a VA */ + SPE_PARSE_SKIP =3D 2, /* complete record without a VA */ + SPE_PARSE_ERROR =3D -EINVAL, /* bad packet; parser resynced */ +}; + +/** + * struct spe_record - One parsed SPE record. + * @va: DATA_VIRT address (bits[55:0]). + * @is_write: Store/load class from OP-TYPE payload bit 0. + * @tid: CONTEXTIDR_EL1 payload (the sampled task's pid). + * @have_addr: An ADDRESS/DATA_VIRT packet was seen. + */ +struct spe_record { + unsigned long va; + bool is_write; + u32 tid; + bool have_addr; +}; + +/** + * struct spe_parser_state - Per-(event,cpu) parser state. + * @win: Linear copy of the drained AUX window (read-only for + * the parser; the caller owns the buffer). + * @win_size: Current window length (bytes). + * @aux_tail: Absolute consumption cursor (rb->aux_head domain). + * @records: Records parsed with a VA (session total). + * @bytes: Bytes consumed (session total). + * + * spe_parse_one_record() never modifies @win. It advances @aux_tail + * and @bytes by the bytes it consumes; the caller is responsible for + * dropping the consumed prefix from @win (memmove + win_size shrink) + * before calling again. + */ +struct spe_parser_state { + u8 *win; + unsigned int win_size; + unsigned long aux_tail; + unsigned int records; + unsigned long bytes; + bool armed; +}; + +/** + * spe_parse_one_record() - Parse one record from the linear window. + * + * Reads the next record from @st->win (a window of @st->win_size bytes + * starting at the stream position @st->aux_tail) and advances + * @st->aux_tail past it. Returns one of enum spe_parse_ret. + * + * An incomplete trailing record is retained. The next drain reparses it + * from its first packet after the producer appends more bytes. Leading P= AD + * and ALIGNMENT packets can be consumed independently. + */ +int spe_parse_one_record(struct spe_parser_state *st, + struct spe_record *rec); + +#endif /* _DAMON_PERF_SPE_PARSER_H */ diff --git a/mm/damon/vaddr.c b/mm/damon/vaddr.c index a68c7262d533..814571c4db78 100644 --- a/mm/damon/vaddr.c +++ b/mm/damon/vaddr.c @@ -18,6 +18,7 @@ #include =20 #include "perf/perf.h" +#include "perf/aux_backend.h" =20 #include "../internal.h" #include "ops-common.h" @@ -1159,6 +1160,21 @@ static int damon_perf_cpu_online(unsigned int cpu, s= truct hlist_node *node) *per_cpu_ptr(perf->event, cpu) =3D perf_event; =20 damon_perf_observe_event_bound(event, cpu, perf_event); + if (!event->ops) + cpumask_clear(&event->aux_cpumask); + damon_perf_aux_select(event, perf_event); + if (event->ops && event->ops->init) { + int ret =3D event->ops->init(event, cpu, perf_event); + + if (ret) { + pr_warn_ratelimited("damon-perf: cpu %u AUX init failed: %d\n", cpu, re= t); + perf_event_release_kernel(perf_event); + *per_cpu_ptr(perf->event, cpu) =3D NULL; + event->any_cpu_failed =3D true; + return 0; + } + cpumask_set_cpu(cpu, &event->aux_cpumask); + } =20 /* * Late-online CPU after the substrate is armed: events are created @@ -1167,6 +1183,16 @@ static int damon_perf_cpu_online(unsigned int cpu, s= truct hlist_node *node) * already-online CPUs. */ if (event->ctx && READ_ONCE(event->ctx->perf_events_active)) { + if (event->ops && event->ops->arm && + event->ops->arm(event, cpu)) { + if (cpumask_test_and_clear_cpu(cpu, &event->aux_cpumask) && + event->ops->cleanup) + event->ops->cleanup(event, cpu); + perf_event_release_kernel(perf_event); + *per_cpu_ptr(perf->event, cpu) =3D NULL; + event->any_cpu_failed =3D true; + return 0; + } perf_event_enable(perf_event); damon_perf_observe_event_enabled(event, cpu, perf_event->state, perf_event->oncpu); @@ -1188,30 +1214,58 @@ static int damon_perf_cpu_offline(unsigned int cpu,= struct hlist_node *node) if (perf_event) { damon_perf_observe_event_destroyed(event, cpu); perf_event_disable(perf_event); + if (event->ops && event->ops->disarm) + event->ops->disarm(event, cpu); + if (cpumask_test_and_clear_cpu(cpu, &event->aux_cpumask) && + event->ops && event->ops->cleanup) + event->ops->cleanup(event, cpu); perf_event_release_kernel(perf_event); *per_cpu_ptr(perf->event, cpu) =3D NULL; } return 0; } =20 -void damon_perf_event_arm(struct damon_perf_event *event) +int damon_perf_event_arm(struct damon_perf_event *event) { struct damon_perf *perf =3D event->priv; struct perf_event *perf_event; - int cpu; + int cpu, failed_cpu =3D nr_cpu_ids; =20 if (!perf) - return; + return -EINVAL; =20 for_each_online_cpu(cpu) { perf_event =3D *per_cpu_ptr(perf->event, cpu); if (perf_event) { + if (event->ops && event->ops->arm && + event->ops->arm(event, cpu)) { + event->any_cpu_failed =3D true; + failed_cpu =3D cpu; + break; + } perf_event_enable(perf_event); damon_perf_observe_event_enabled(event, cpu, perf_event->state, perf_event->oncpu); } } + if (failed_cpu =3D=3D nr_cpu_ids) + return 0; + + /* Roll back CPUs enabled by this arm attempt. */ + for_each_online_cpu(cpu) { + if (cpu >=3D failed_cpu) + break; + perf_event =3D *per_cpu_ptr(perf->event, cpu); + if (!perf_event) + continue; + perf_event_disable(perf_event); + if (event->ops && event->ops->disarm) + event->ops->disarm(event, cpu); + damon_perf_observe_event_disabled(event, cpu, + perf_event->state); + } + return -EIO; } =20 void damon_perf_event_disarm(struct damon_perf_event *event) @@ -1227,6 +1281,8 @@ void damon_perf_event_disarm(struct damon_perf_event = *event) perf_event =3D *per_cpu_ptr(perf->event, cpu); if (perf_event) { perf_event_disable(perf_event); + if (event->ops && event->ops->disarm) + event->ops->disarm(event, cpu); damon_perf_observe_event_disabled(event, cpu, perf_event->state); } @@ -1272,6 +1328,8 @@ int damon_perf_init(struct damon_ctx *ctx, struct dam= on_perf_event *event) =20 free_event: damon_perf_observe_event_free(event); + if (perf->aux_priv) + free_percpu((void __percpu *)perf->aux_priv); free_percpu(perf->event); free_perf: kfree(perf); @@ -1291,6 +1349,10 @@ void damon_perf_cleanup(struct damon_ctx *ctx, struc= t damon_perf_event *event) cpuhp_state_remove_instance(damon_perf_cpuhp_state, &event->hlist_node); =20 + if (perf->aux_priv) { + free_percpu((void __percpu *)perf->aux_priv); + perf->aux_priv =3D NULL; + } free_percpu(perf->event); kfree(perf); event->priv =3D NULL; --=20 2.43.0 From nobody Mon Aug 24 04:16:54 2026 Received: from mail-pl1-f169.google.com (mail-pl1-f169.google.com [209.85.214.169]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 261103B19B7 for ; Sun, 16 Aug 2026 14:23:06 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.169 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786890188; cv=none; b=dq8oWTc6zaI0CIAh/WDW8KIAvNGzm4klnOcIhsVnISDq60IUgge+dnbwMtkMbhJSYqsS2FNJ2EJE6vMg9jhJmdd8WpetF83xfoBOdYF9rOH/lZYfEPstjHWKlxm6fj/qm+fRNI27N3vtrdI3Sw8+2e3pkswnDbo0/JjwnkoEeXE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786890188; c=relaxed/simple; bh=aJCxxTWPtwoBZ6fKwboPkrJD4uPFgZMJY9FISCbeF2E=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=NqD+QS13IgAQLAuC4654gzObPJrjQ1eBRi/bI1jLgGM9TfyMmVi8FsyAe4X+7i0x4fAmFvSsKOBT1DhA154psOBqvWuZ3DxVJ1kNJ4WB3i13bjKlSgUWyZugkAUJKrAW9uNH2jjCoZlF5IcM09d+y0DVKolch2bmzbLACWojMW8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=UJg+ZAF3; arc=none smtp.client-ip=209.85.214.169 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="UJg+ZAF3" Received: by mail-pl1-f169.google.com with SMTP id d9443c01a7336-2cacb8416a1so24353045ad.1 for ; Sun, 16 Aug 2026 07:23:06 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786890185; x=1787494985; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=V80z3viJi/1rXBOV+LmIMGbRsmAIYZLeea8uFd5X4Y0=; b=UJg+ZAF3BDujDCDoTykBAQz6DyhPA9t9N/AtoJF7azN3/OFtHApOpEfG91UammfiuM 58k9lqRCDPLogUTGFqFmezui/N7b4pi3TIWaNxBmsKDcb9Dtz25x9hMWnxRPra4UCqC4 pGyl3HmhvY3BKxrIBtRaaaUj3quOWmgMo+Tj9bGwCCrcmY9z9wDM5LuPUVUGi86rrOWK o8nemciv2A776TeSZ+Jx3MbqVrRqLKIu5XW+i39OfThK9F7G2iLxKOoJZlQa6bpi8kbn z803oZF6M2+hyY4AdqG8CtUj8VO2r4fenEP8jTus+1HIyRD5W2wl3CmcVZjUBpSKiE3Y Dxhg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786890185; x=1787494985; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=V80z3viJi/1rXBOV+LmIMGbRsmAIYZLeea8uFd5X4Y0=; b=WnQwtN9Ug5hQKcRMKyblpdlIobJjx5VeBnaApQaOr1YGDpm0uE4myciGHCICnSl9Vm 4KiPGY3aT28K8x48iuJgdvLZbAuOvVNCg8cuRDXyB0bE+ymvF6PVceBPc65eBj6MHMxK go8L+tt1+WjgekXemzTCny31v9s0cZKRSxmDZMYFfTvQ6cG4E4hnQF9S5UVXdiyDebSD g/UkL9oCmGxaXAgrVoSxHOjTlusZdjq8ECJNCCnZWpoUjsQMa6qUXWK6QpOUtccfXBlW pdwb0f4mbxsUumZ/i3FEbU8/DI40aKMcSqK1AeJ1OcoX0I6lbJuJrtMjqw8y8DU0J5Z2 Cguw== X-Gm-Message-State: AOJu0YzC2OpV7zlZCklIDcudsiYzuBVgnI2S9ViDaJWTD3qZYBFf9gqZ 260QKw2IOeN+FY59nuPQqW5b3RxqTypAsRE04z+jqtO3ow0ZQoJZKk2+ X-Gm-Gg: AR+sD11EROQ+rPYKYVajLV/57y2ymEGlLyJCbNx5eIaJ2NCMXxOe/rQcAWr2JbrrAQu KgEsxirX1nTDCuRrpG/ciC/JBR9WHmTN6VmxYmP/7Qdf2ApuRGJFhVCmNJ94L/8tD3fcVKoxsWC Luur4v9gbOynDo9qmh8TnGSISHuSCujCapgVm9iqTOxKS5jNZzCW5YNLrkVcETLhWEP8ZJBkBTW BCmuKrONktmKEY3NWH6whlwt5K39tjl6YhbiX+tHCymCpk158VFdRTxXplH0DQ3FqhDMWDvxlao pwFJAXFSHGRJxZHYzkZt+3lwQ/Ro9XraI0ZPKlZU+WlzKI2cObPu9MeVaFttzi4e1ja/x6nrmxk SGGR8oHd+mHbh/7TxCjiI2aqExzkAt4ohklTXqjErUTUNln7n7DCp/wUQ5uzx68IFQtDE4psY/Q Pq+ec01MunzO99YHCVPjXWFUMDJfrkb2NOAnF33lUmeGPtKxKKuxIFe8tOjk4hT30gSX27UCM4j biedTQHmXQaGv8a4Q== X-Received: by 2002:a17:902:ce8c:b0:2cf:afa5:b19a with SMTP id d9443c01a7336-2d3b0d13bb8mr212666695ad.11.1786890185194; Sun, 16 Aug 2026 07:23:05 -0700 (PDT) Received: from kernel.tail6741c6.ts.net ([116.128.244.169]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2d3aebc8199sm26102365ad.74.2026.08.16.07.22.57 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sun, 16 Aug 2026 07:23:04 -0700 (PDT) From: Kunwu Chan X-Google-Original-From: Kunwu Chan To: will@kernel.org, mark.rutland@arm.com, sj@kernel.org, akpm@linux-foundation.org, shuah@kernel.org, kunwu.chan@linux.dev Cc: linux-kernel@vger.kernel.org, linux-arm-kernel@lists.infradead.org, linux-perf-users@vger.kernel.org, damon@lists.linux.dev, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, "Lian Wang (ProcessMission)" , Kunwu Chan Subject: [RFC PATCH 3/4] mm/damon/perf: add KUnit tests for the SPE record parser Date: Sun, 16 Aug 2026 22:22:20 +0800 Message-ID: <20260816142222.689624-4-kunwu.chan@linux.dev> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260816142222.689624-1-kunwu.chan@linux.dev> References: <20260816142222.689624-1-kunwu.chan@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Lian Wang (ProcessMission)" Add byte-exact tests for load and store records, timestamp terminators, multiple records, PAD and ALIGNMENT packets, extended addresses, invalid extended headers, error resynchronization, and records without a virtual address. Cover ALIGNMENT packets at both odd and already aligned stream offsets. Also verify that a record split across two AUX snapshots leaves the tail unchanged until the terminating packet becomes available. Co-developed-by: Kunwu Chan Signed-off-by: Kunwu Chan Signed-off-by: Lian Wang (ProcessMission) --- mm/damon/perf/spe_parser_test.c | 365 ++++++++++++++++++++++++++++++++ 1 file changed, 365 insertions(+) create mode 100644 mm/damon/perf/spe_parser_test.c diff --git a/mm/damon/perf/spe_parser_test.c b/mm/damon/perf/spe_parser_tes= t.c new file mode 100644 index 000000000000..598d9fd7cdc7 --- /dev/null +++ b/mm/damon/perf/spe_parser_test.c @@ -0,0 +1,365 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * KUnit tests for the DAMON perf ARM SPE record parser. + * + * Each parameterized case feeds a byte-exact SPE packet stream (same + * encodings and decode order as tools/perf/util/arm-spe-decoder) into + * spe_parse_one_record() and verifies the synthesized records, the + * return values, and the aux_tail accounting. The loop mirrors + * spe_backend_drain() including the caller-side window shrink. + */ + +#include +#include + +#include "spe_parser.h" + +/** + * struct spe_parse_case - One parser test case. + * @name: Parameter description (shown on failure). + * @stream: Byte-exact SPE packet stream. + * @len: @stream length. + * @exp_reports: Expected SPE_PARSE_REPORT count. + * @exp_skips: Expected SPE_PARSE_SKIP count. + * @exp_errors: Expected SPE_PARSE_ERROR count. + * @exp_tail: Expected st->aux_tail after the stream. + * @exp_va: Expected virtual address of the first REPORT record. + * @exp_tid: Expected tid of the first REPORT record. + * @exp_is_write: Expected access type of the first REPORT record. + */ +struct spe_parse_case { + const char *name; + const u8 *stream; + size_t len; + unsigned int exp_reports; + unsigned int exp_skips; + unsigned int exp_errors; + unsigned long exp_tail; + unsigned long exp_va; + u32 exp_tid; + bool exp_is_write; +}; + +static const u8 stream_store[] =3D { + 0x66, 0x2a, 0x00, 0x00, 0x00, /* CONTEXT: 64-bit EL1 tid=3D42 */ + 0xb2, 0x00, 0x10, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* VA 0x1000 */ + 0x49, 0x01, /* OP-TYPE: ST */ + 0x01, /* END */ +}; + +static const u8 stream_ts_load[] =3D { + 0xb2, 0x00, 0x20, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* VA 0x2000 */ + 0x49, 0x00, /* OP-TYPE: load */ + 0x71, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* TIMESTAMP end */ +}; + +static const u8 stream_two[] =3D { + 0x66, 0x2a, 0x00, 0x00, 0x00, /* CONTEXT: 64-bit EL1 tid=3D42 */ + 0xb2, 0x00, 0x10, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* VA 0x1000 */ + 0x49, 0x01, /* OP-TYPE: ST */ + 0x01, /* END */ + 0xb2, 0x00, 0x20, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* VA 0x2000 */ + 0x49, 0x00, /* OP-TYPE: load */ + 0x01, /* END */ +}; + +static const u8 stream_pad[] =3D { + 0x00, 0x00, /* PAD prefix */ + 0x66, 0x2a, 0x00, 0x00, 0x00, /* CONTEXT: 64-bit EL1 tid=3D42 */ + 0xb2, 0x00, 0x10, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* VA 0x1000 */ + 0x49, 0x01, /* OP-TYPE: ST */ + 0x01, /* END */ + 0x00, 0x00, 0x00, /* PAD padding */ +}; + +static const u8 stream_alignment[] =3D { + 0x66, 0x2a, 0x00, 0x00, 0x00, /* CONTEXT: tid 42, pos 0-4 */ + 0x20, 0x00, /* ALIGNMENT at odd pos 5 */ + 0xb2, 0x00, 0x10, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* VA 0x1000 */ + 0x49, 0x00, /* OP-TYPE: load */ + 0x01, /* END */ +}; + +static const u8 stream_alignment_aligned[] =3D { + 0x20, 0x00, /* ALIGNMENT at even pos 0 */ + 0x66, 0x2a, 0x00, 0x00, 0x00, /* CONTEXT: 64-bit EL1 tid=3D42 */ + 0xb2, 0x00, 0x10, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* VA 0x1000 */ + 0x49, 0x00, /* OP-TYPE: load */ + 0x01, /* END */ +}; + +static const u8 stream_bad[] =3D { + 0xff, /* unknown header */ + 0x66, 0x2a, 0x00, 0x00, 0x00, /* CONTEXT: 64-bit EL1 tid=3D42 */ + 0xb2, 0x00, 0x10, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* VA 0x1000 */ + 0x49, 0x01, /* OP-TYPE: ST */ + 0x01, /* END */ +}; + +static const u8 stream_skip[] =3D { + 0x49, 0x00, /* OP-TYPE: load */ + 0x01, /* END, no address */ +}; + +static const u8 stream_other_pkts[] =3D { + 0x42, 0x05, /* EVENTS (width 1) */ + 0x43, 0x06, /* DATA-SOURCE (width 1) */ + 0x98, 0x00, 0x00, /* COUNTER (width 2) */ + 0xb2, 0x00, 0x30, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* VA 0x3000 */ + 0x49, 0x00, /* OP-TYPE: load */ + 0x01, /* END */ +}; + +static const u8 stream_truncated[] =3D { + 0x66, 0x2a, 0x00, 0x00, 0x00, /* CONTEXT: 64-bit EL1 tid=3D42 */ + 0xb2, 0x00, 0x10, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* VA 0x1000 */ + 0x49, 0x01, /* OP-TYPE: ST, no END */ +}; + +static const u8 stream_truncated_packet[] =3D { + 0xb2, 0x00, 0x10, /* short 8-byte address */ +}; + +static const u8 stream_ext_addr[] =3D { + 0x20, 0xb2, /* EXTENDED ADDRESS, DATA_VIRT */ + 0x00, 0x40, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* VA 0x4000 */ + 0x49, 0x00, /* OP-TYPE: load */ + 0x01, /* END */ +}; + +static const u8 stream_invalid_extended[] =3D { + 0x20, 0x42, 0x00, /* invalid extended EVENTS */ + 0x01, /* END after resync */ +}; + +static const u8 stream_pad_only[] =3D { + 0x00, 0x00, 0x00, +}; + +static const u8 stream_empty[] =3D { 0x00 }; + +static const struct spe_parse_case spe_parse_cases[] =3D { + { + .name =3D "store record", + .stream =3D stream_store, + .len =3D sizeof(stream_store), + .exp_reports =3D 1, + .exp_tail =3D 17, + .exp_va =3D 0x1000, + .exp_tid =3D 42, + .exp_is_write =3D true, + }, + { + .name =3D "load record with timestamp terminator", + .stream =3D stream_ts_load, + .len =3D sizeof(stream_ts_load), + .exp_reports =3D 1, + .exp_tail =3D 20, + .exp_va =3D 0x2000, + .exp_tid =3D 0, + .exp_is_write =3D false, + }, + { + .name =3D "two records in one window", + .stream =3D stream_two, + .len =3D sizeof(stream_two), + .exp_reports =3D 2, + .exp_tail =3D 29, + .exp_va =3D 0x1000, + .exp_tid =3D 42, + .exp_is_write =3D true, + }, + { + .name =3D "pad-wrapped record", + .stream =3D stream_pad, + .len =3D sizeof(stream_pad), + .exp_reports =3D 1, + .exp_tail =3D 22, + .exp_va =3D 0x1000, + .exp_tid =3D 42, + .exp_is_write =3D true, + }, + { + .name =3D "alignment packet at odd position", + .stream =3D stream_alignment, + .len =3D sizeof(stream_alignment), + .exp_reports =3D 1, + .exp_tail =3D 19, + .exp_va =3D 0x1000, + .exp_tid =3D 42, + .exp_is_write =3D false, + }, + { + .name =3D "alignment packet at aligned position", + .stream =3D stream_alignment_aligned, + .len =3D sizeof(stream_alignment_aligned), + .exp_reports =3D 1, + .exp_tail =3D sizeof(stream_alignment_aligned), + .exp_va =3D 0x1000, + .exp_tid =3D 42, + .exp_is_write =3D false, + }, + { + .name =3D "bad packet resync", + .stream =3D stream_bad, + .len =3D sizeof(stream_bad), + .exp_reports =3D 1, + .exp_errors =3D 1, + .exp_tail =3D 18, + .exp_va =3D 0x1000, + .exp_tid =3D 42, + .exp_is_write =3D true, + }, + { + .name =3D "record without address", + .stream =3D stream_skip, + .len =3D sizeof(stream_skip), + .exp_skips =3D 1, + .exp_tail =3D 3, + }, + { + .name =3D "events/source/counter packets ignored", + .stream =3D stream_other_pkts, + .len =3D sizeof(stream_other_pkts), + .exp_reports =3D 1, + .exp_tail =3D 19, + .exp_va =3D 0x3000, + .exp_is_write =3D false, + }, + { + .name =3D "truncated trailing record retained", + .stream =3D stream_truncated, + .len =3D sizeof(stream_truncated), + .exp_tail =3D 0, + }, + { + .name =3D "truncated packet retained", + .stream =3D stream_truncated_packet, + .len =3D sizeof(stream_truncated_packet), + .exp_tail =3D 0, + }, + { + .name =3D "extended address packet", + .stream =3D stream_ext_addr, + .len =3D sizeof(stream_ext_addr), + .exp_reports =3D 1, + .exp_tail =3D 13, + .exp_va =3D 0x4000, + .exp_is_write =3D false, + }, + { + .name =3D "invalid extended header resync", + .stream =3D stream_invalid_extended, + .len =3D sizeof(stream_invalid_extended), + .exp_skips =3D 1, + .exp_errors =3D 1, + .exp_tail =3D sizeof(stream_invalid_extended), + }, + { + .name =3D "pad-only window", + .stream =3D stream_pad_only, + .len =3D sizeof(stream_pad_only), + .exp_tail =3D 3, + }, + { + .name =3D "empty window", + .stream =3D stream_empty, + .len =3D 0, + .exp_tail =3D 0, + }, +}; + +KUNIT_ARRAY_PARAM_DESC(spe_parse, spe_parse_cases, name); + +static void spe_parse_case_test(struct kunit *test) +{ + const struct spe_parse_case *tc =3D test->param_value; + struct spe_parser_state st =3D { 0 }; + struct spe_record rec; + u8 *buf; + unsigned int reports =3D 0, skips =3D 0, errors =3D 0, guard =3D 0; + bool first_checked =3D false; + + buf =3D kunit_kmalloc(test, tc->len ?: 1, GFP_KERNEL); + KUNIT_ASSERT_NOT_ERR_OR_NULL(test, buf); + memcpy(buf, tc->stream, tc->len); + st.win =3D buf; + st.win_size =3D tc->len; + + while (guard++ < SPE_BUFFER_MAX_RECORDS) { + unsigned long tail0 =3D st.aux_tail; + unsigned long consumed; + int ret =3D spe_parse_one_record(&st, &rec); + + if (ret =3D=3D SPE_PARSE_NEED_MORE) + break; + + /* caller-side window shrink, mirrors spe_backend_drain() */ + consumed =3D st.aux_tail - tail0; + st.win_size -=3D consumed; + memmove(st.win, st.win + consumed, st.win_size); + + switch (ret) { + case SPE_PARSE_REPORT: + reports++; + if (!first_checked) { + KUNIT_EXPECT_EQ(test, tc->exp_va, rec.va); + KUNIT_EXPECT_EQ(test, tc->exp_tid, rec.tid); + KUNIT_EXPECT_EQ(test, tc->exp_is_write, + rec.is_write); + first_checked =3D true; + } + break; + case SPE_PARSE_SKIP: + skips++; + break; + case SPE_PARSE_ERROR: + errors++; + break; + } + } + + KUNIT_EXPECT_EQ(test, tc->exp_reports, reports); + KUNIT_EXPECT_EQ(test, tc->exp_skips, skips); + KUNIT_EXPECT_EQ(test, tc->exp_errors, errors); + KUNIT_EXPECT_EQ(test, tc->exp_tail, st.aux_tail); + KUNIT_EXPECT_EQ(test, tc->exp_tail, st.bytes); + KUNIT_EXPECT_EQ(test, tc->exp_reports, st.records); +} + +static void spe_split_record_test(struct kunit *test) +{ + struct spe_parser_state st =3D { + .win =3D (u8 *)stream_store, + .win_size =3D sizeof(stream_store) - 1, + }; + struct spe_record rec; + int ret; + + ret =3D spe_parse_one_record(&st, &rec); + KUNIT_ASSERT_EQ(test, SPE_PARSE_NEED_MORE, ret); + KUNIT_EXPECT_EQ(test, 0UL, st.aux_tail); + KUNIT_EXPECT_EQ(test, 0UL, st.bytes); + + /* The next AUX copy starts at the unchanged tail and includes END. */ + st.win_size =3D sizeof(stream_store); + ret =3D spe_parse_one_record(&st, &rec); + KUNIT_ASSERT_EQ(test, SPE_PARSE_REPORT, ret); + KUNIT_EXPECT_EQ(test, (unsigned long)sizeof(stream_store), + st.aux_tail); + KUNIT_EXPECT_EQ(test, 0x1000UL, rec.va); + KUNIT_EXPECT_EQ(test, 42U, rec.tid); +} + +static struct kunit_case spe_parser_test_cases[] =3D { + KUNIT_CASE_PARAM(spe_parse_case_test, spe_parse_gen_params), + KUNIT_CASE(spe_split_record_test), + {}, +}; + +static struct kunit_suite spe_parser_test_suite =3D { + .name =3D "damon_perf_spe_parser", + .test_cases =3D spe_parser_test_cases, +}; + +kunit_test_suite(spe_parser_test_suite); --=20 2.43.0 From nobody Mon Aug 24 04:16:54 2026 Received: from mail-pl1-f176.google.com (mail-pl1-f176.google.com [209.85.214.176]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CD2A1329E5A for ; Sun, 16 Aug 2026 14:23:13 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.176 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786890195; cv=none; b=R02wD81OiFdQ8l3PSEcetFEPkkWWNLdLckIL1Dmc/xVBKl+UEhu+f1EWz6O5Msu3ElY6uJ8vII1+wjV33TSF+1l4MUCB7/n0bJOhCrJqVh8cbSql3Q6ZwEXPyZTFLpMl8zSXbB75Vz4c1LymbUQKb8gVSqxhnn/9oM9h03EJsZM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786890195; c=relaxed/simple; bh=1SMO3jUvVdF+0ZsnAaZeuLRBqr2AGnWazoTU4AFxdcs=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=D79aAsP6MosDCH6uqyPUfxo+xnEWuCh/sAVq4jjQEXeuzWi7TFnokkWzyuMZrbMQD70E+qEU7+X9IuD6D0ZqRjGtDvs8hwIlPIZs1Pq1kBKRjd2XqiRSfe3TiMrPwZUKRD9H8tgqJhGXRJXz16vN3B4iYpenhrrba3K3/BLgpgQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=H6ndfzXk; arc=none smtp.client-ip=209.85.214.176 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="H6ndfzXk" Received: by mail-pl1-f176.google.com with SMTP id d9443c01a7336-2cf52d15d88so24626235ad.2 for ; Sun, 16 Aug 2026 07:23:13 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786890193; x=1787494993; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=VZh5CgmasB3FNwZhE5a0RmyqcaBUm5A+ddeEPpNzaA4=; b=H6ndfzXkM25PxJf+61CsdEftdWn1QolgXXX//LABwIFN/00He1bcWNIIWNaMA+CIMF laXlfJ/kPs4F4fK+1jA2PcQ+imwM6CL+bmG0gOF6ceoGrX7/r8KXcB60+j+1UAIQ8o2l /wlEdPKEDMNiui/tDUhi+Fy1murrhbg+So4gnLxuaczVYZkIcN1vIj1z4LmxuVIU0WrE Un4gYg5L9zZXUKdC1B8qBC3JPJ1w+yYE8uqThF1yZ1ssxImIdffPxSWApQW7nKOqXhh4 3a6+ApGnp18UQF7pfwFvpxACGg2NQDda1u+viwMKI8YwhKzs6BpHoZaofafJTpwZMi8q E91A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786890193; x=1787494993; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=VZh5CgmasB3FNwZhE5a0RmyqcaBUm5A+ddeEPpNzaA4=; b=bGnfUWbQWgqxd1Rw4piL+jETYTytLOGjx0Fx94+wK49caul6rTfAVsaQc4IqjEg9F6 HH51Z66gxegPRrVAgCtBaH0+ynBegq6u9cYTj4bCt9Mv2UnjMqsdFESIg4oCXi0I1jZl tg7MHKkXpF8YytzRKAiBTEfdqg6Q0q2W5spAtRdETP+SxnWNdFDEyOCVE5RMuHAf6WiO nZWcnES8ryEf6gcNWnQlXx+nWtttwGSHXuh9LCdq2r03At/8PH4iDLyK7FDoEMRoNW7C Vk+6P4NttnR4UC4pSmffRDjAfNKRCof53J4iSvpk8nnI5XBwAhJSyqgd6sjy3EVHS1BJ VSAw== X-Gm-Message-State: AOJu0YzhkzwJLSoskT33F7o3lwntIbZABD1zstPEuZ6egv1xei70rcSX eKTvYEwTsCcH7RccJdyQW5bxQeHhYOfu5Jt1svZ/0/dZ07IXq+QjHifD X-Gm-Gg: AR+sD1043DPxkxLxe9wuG8Zq/BaVjUdrcFRTeJi16QtzCpeWFF9eeInBYECDSxzyI+i H3pLPm339WXVW9wLTxVW3rAxZStmxYbM/pAH3uryRjTfTLhzCA47PbeBwVT9JD7K3ZoJ0MEs/1q iguO5KqtF1Vj4VxbemQhqSCpbV/z65FyoaiyveAdj7jaEPuelbpDqDwUJzCILCVkdmARack3FMP MR5UpnUXqEnCWTOWi/GxTZf7TqxpLgQhA0682rtWTS9pRbg6fU8hQOQztlzIKK//De36p20ccd4 SuQn2JYQi5RtV11KtPSi+jJKo6c7JaVtRlXHCQTXkW6oAypRljYtqfMQWLNVTLnfqI5QXm6HcQw yE8CGLh7o/vsh9GeSUTB2PuLel7HXYILnBQ5eEGDtK8NBSQzIvlYt1TTorlLrs/jP1PR2mmw1Ah 33TJBHnMczAFLp4JhPXd7djvtqgZ5WscwlSbMMor6BsTTVJXrfBDP1mF/QuoVyyVtIy+hhSqKZC OLDjOY= X-Received: by 2002:a17:903:8d0:b0:2cc:aa36:c04c with SMTP id d9443c01a7336-2d3b0ad0fa1mr214403475ad.1.1786890192602; Sun, 16 Aug 2026 07:23:12 -0700 (PDT) Received: from kernel.tail6741c6.ts.net ([116.128.244.169]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2d3aebc8199sm26102365ad.74.2026.08.16.07.23.05 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sun, 16 Aug 2026 07:23:12 -0700 (PDT) From: Kunwu Chan X-Google-Original-From: Kunwu Chan To: will@kernel.org, mark.rutland@arm.com, sj@kernel.org, akpm@linux-foundation.org, shuah@kernel.org, kunwu.chan@linux.dev Cc: linux-kernel@vger.kernel.org, linux-arm-kernel@lists.infradead.org, linux-perf-users@vger.kernel.org, damon@lists.linux.dev, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, "Lian Wang (ProcessMission)" , Kunwu Chan Subject: [RFC PATCH 4/4] selftests/damon: add DAMON perf AUX backend test Date: Sun, 16 Aug 2026 22:22:21 +0800 Message-ID: <20260816142222.689624-5-kunwu.chan@linux.dev> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260816142222.689624-1-kunwu.chan@linux.dev> References: <20260816142222.689624-1-kunwu.chan@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Lian Wang (ProcessMission)" Register a DAMON kselftest for the ARM SPE AUX backend. When kernel sources are available, check the backend integration and the required AUX-before-ring lifecycle ordering. Independently run the SPE parser KUnit suite through debugfs when it is built. On an ARM SPE system with no pre-existing kdamond, create one controlled userspace target and collect counter deltas from a five-second session. Stop the session explicitly, then verify positive end-to-end pipeline counters and final enqueue/dequeue closure. Use a private mktemp result directory and restore only the DAMON state created by this test. This is a smoke and lifecycle test; full-ring pause/resume validation remains a separate hardware stress test. Co-developed-by: Kunwu Chan Signed-off-by: Kunwu Chan Signed-off-by: Lian Wang (ProcessMission) --- tools/testing/selftests/damon/Makefile | 1 + .../selftests/damon/damon_perf_aux_test.sh | 414 ++++++++++++++++++ 2 files changed, 415 insertions(+) create mode 100755 tools/testing/selftests/damon/damon_perf_aux_test.sh diff --git a/tools/testing/selftests/damon/Makefile b/tools/testing/selftes= ts/damon/Makefile index 1db8fa95ba2d..204a0b60d84a 100644 --- a/tools/testing/selftests/damon/Makefile +++ b/tools/testing/selftests/damon/Makefile @@ -24,4 +24,5 @@ TEST_PROGS +=3D sysfs_no_op_commit_break.py EXTRA_CLEAN =3D __pycache__ =20 TEST_PROGS +=3D damon_perf_obs_test.sh +TEST_PROGS +=3D damon_perf_aux_test.sh include ../lib.mk diff --git a/tools/testing/selftests/damon/damon_perf_aux_test.sh b/tools/t= esting/selftests/damon/damon_perf_aux_test.sh new file mode 100755 index 000000000000..86ccaca6310c --- /dev/null +++ b/tools/testing/selftests/damon/damon_perf_aux_test.sh @@ -0,0 +1,414 @@ +#!/bin/bash +# SPDX-License-Identifier: GPL-2.0 +# +# DAMON Perf AUX Backend - Automated Test +# +# Validates the AUX trace-buffer backend framework and the ARM SPE +# backend: +# 1. Source-level structure checks (files, Makefile, ops interface) +# 2. SPE parser: runs the KUnit suite against byte-exact binary +# fixtures (mm/damon/perf/spe_parser_test.c) through debugfs +# 3. Architecture checks (exact PMU match, SPSC ring contract, +# lifecycle ordering) +# 4. Runtime smoke test: live DAMON session configured with an ARM +# SPE event. Only runs when no kdamonds exist; the pre-test +# sysfs state is restored on exit. +# +# Usage: sudo ./damon_perf_aux_test.sh [pmu-name] +# +# Requirements: +# - CONFIG_DAMON_PERF_OBSERVE=3Dy +# - CONFIG_DAMON_PERF_SPE_KUNIT_TEST=3Dy and CONFIG_KUNIT_DEBUGFS=3Dy +# for section 2 (skipped otherwise) +# - CONFIG_ARM_SPE_PMU=3Dy for section 4 (skipped otherwise) +# - Root privileges + +set -e +PASSED=3D0; FAILED=3D0; SKIPPED=3D0 +pass() { echo " [PASS] $1"; PASSED=3D$((PASSED + 1)); } +fail() { echo " [FAIL] $1"; FAILED=3D$((FAILED + 1)); } +skip() { echo " [SKIP] $*"; SKIPPED=3D$((SKIPPED + 1)); } + +PMU_NAME=3D"${1:-arm_spe_0}" +RESULTS_DIR=3D$(mktemp -d "${TMPDIR:-/tmp}/damon_aux_test.XXXXXX") || exit= 1 +exec > >(tee "$RESULTS_DIR/output.log") 2>&1 + +ROOT=3D"$(git rev-parse --show-toplevel 2>/dev/null || true)" +if [[ -z "$ROOT" ]]; then + ROOT=3D"/usr/src/linux" +fi +SRC_DIR=3D"$ROOT/mm/damon/perf" + +ADMIN=3D"/sys/kernel/mm/damon/admin" + +echo "=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D" +echo " DAMON Perf AUX Backend Test" +echo "=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D" + +# ---- Section 1: Source structure ---- +if [[ -d "$SRC_DIR" && -f "$ROOT/mm/damon/core.c" ]]; then +echo "" +echo "--- 1. Backend Source Files ---" +for f in aux_backend.h aux_backend.c spe_parser.h spe_backend.c spe_parser= _test.c; do + if [[ -f "$SRC_DIR/$f" ]]; then + pass "Source file: $f" + else + fail "Source file: $f" + fi +done + +echo "" +echo "--- 2. Makefile Registration ---" +for obj in aux_backend.o spe_backend.o spe_parser_test.o; do + if grep -q "$obj" "$SRC_DIR/Makefile" 2>/dev/null; then + pass "Makefile registers: $obj" + else + fail "Makefile registers: $obj" + fi +done + +echo "" +echo "--- 3. Backend Ops Interface ---" +OPS_COUNT=3D$(grep -c '(\*match_pmu)\|(\*init)\|(\*cleanup)\|(\*arm)\|(\*d= isarm)\|(\*drain)' \ + "$SRC_DIR/aux_backend.h" 2>/dev/null || echo 0) +if [[ "$OPS_COUNT" -ge 6 ]]; then + pass "Backend ops interface: 6 ops defined ($OPS_COUNT found)" +else + fail "Backend ops interface: expected 6 ops, found $OPS_COUNT" +fi +if grep -q 'DAMON_PERF_BACKEND_AUX' "$SRC_DIR/aux_backend.h" 2>/dev/null; = then + pass "DAMON_PERF_BACKEND_AUX flag defined" +else + fail "DAMON_PERF_BACKEND_AUX flag" +fi + +echo "" +echo "--- 4. SPE Parser Constants ---" +for c in SPE_BUFFER_MIN_PAGES SPE_BUFFER_MAX_RECORDS SPE_HDR_MASK1 SPE_HDR= _PAD \ + SPE_HDR_END SPE_HDR_TIMESTAMP SPE_HDR_EXTENDED SPE_ADDR_DATA_VIRT \ + SPE_OP_CLASS_LDST; do + if grep -q "$c" "$SRC_DIR/spe_parser.h" 2>/dev/null; then + pass "Parser constant: $c" + else + fail "Parser constant: $c" + fi +done +else + skip "source structure" "kernel source tree not available" +fi + +# ---- Section 2: Parser correctness (KUnit, binary fixtures) ---- +echo "" +echo "--- 5. SPE Parser: KUnit suite (binary fixtures) ---" +KUNIT_DIR=3D"/sys/kernel/debug/kunit/damon_perf_spe_parser" +if [[ -d "$KUNIT_DIR" ]]; then + # Writing to "run" triggers a fresh run; "results" then holds the + # TAP output of the byte-exact fixture cases. + if [[ -f "$KUNIT_DIR/run" ]]; then + echo run > "$KUNIT_DIR/run" 2>/dev/null || true + fi + RES=3D"$KUNIT_DIR/results" + if [[ -f "$RES" ]]; then + if grep -Eq "^[[:space:]]*not ok" "$RES"; then + fail "KUnit parser suite has failing cases" + grep -E "^[[:space:]]*not ok" "$RES" | sed 's/^/ /' + elif grep -Eq "^[[:space:]]*ok " "$RES"; then + pass "KUnit parser suite: all cases pass" + else + fail "KUnit parser suite produced no completed cases" + fi + echo " $(grep '# Totals:' "$RES" | tail -1)" + # Spot-check the fixtures that cover the review-critical paths: + # record content, consumed-length accounting, truncation, + # alignment, extended packets, and error resync. + for c in "store record" "bad packet resync" \ + "alignment packet at odd position" \ + "alignment packet at aligned position" \ + "extended address packet" \ + "invalid extended header resync" \ + "truncated trailing record retained" \ + "truncated packet retained" \ + "pad-only window"; do + if grep -Eq "^[[:space:]]*ok .*$c" "$RES"; then + pass "fixture: $c" + else + fail "fixture: $c" + fi + done + else + skip "KUnit results" "no results file (suite did not run)" + fi +else + skip "KUnit parser suite" "CONFIG_KUNIT_DEBUGFS missing or suite not b= uilt" +fi + +# ---- Section 3: Architecture checks ---- +echo "" +echo "--- 6. Architecture Correctness ---" + +if [[ ! -d "$SRC_DIR" || ! -f "$ROOT/mm/damon/core.c" ]]; then + skip "architecture source checks" "kernel source tree not available" +else + +# Exact PMU match by pmu object, not by name-prefix matching. +if grep -q 'arm_spe_pmu_match' "$SRC_DIR/spe_backend.c" 2>/dev/null; then + pass "PMU match: exact match via arm_spe_pmu_match()" +else + fail "PMU match: arm_spe_pmu_match() missing" +fi +if grep -q 'strncmp.*arm_spe' "$SRC_DIR/spe_backend.c" 2>/dev/null; then + fail "PMU match: strncmp name-prefix matching must not be used" +else + pass "PMU match: no strncmp name-prefix matching" +fi + +# Backend selection: gated on the ITRACE capability, then matched by PMU. +if grep -q 'PERF_PMU_CAP_ITRACE' "$SRC_DIR/aux_backend.c" 2>/dev/null; then + pass "PMU selection: gated by PERF_PMU_CAP_ITRACE" +else + fail "PMU selection: ITRACE gate missing" +fi + +# Lifecycle ordering: arm() before perf_event_enable(), disarm() after +# perf_event_disable(). +ARM_LINE=3D$(awk '/^int damon_perf_event_arm\(/ { in_fn=3D1 } \ + in_fn && /ops->arm/ { print NR; exit }' "$ROOT/mm/damon/vaddr.c") +ENABLE_LINE=3D$(awk '/^int damon_perf_event_arm\(/ { in_fn=3D1 } \ + in_fn && /perf_event_enable/ { print NR; exit }' "$ROOT/mm/damon/vaddr= .c") +if [[ -n "$ARM_LINE" && -n "$ENABLE_LINE" && "$ARM_LINE" -lt "$ENABLE_LINE= " ]]; then + pass "Lifecycle: ops->arm() before perf_event_enable()" +else + fail "Lifecycle: arm() must precede perf_event_enable()" +fi + +DISABLE_LINE=3D$(awk '/^void damon_perf_event_disarm\(/ { in_fn=3D1 } \ + in_fn && /perf_event_disable/ { print NR; exit }' "$ROOT/mm/damon/vadd= r.c") +DISARM_LINE=3D$(awk '/^void damon_perf_event_disarm\(/ { in_fn=3D1 } \ + in_fn && /ops->disarm/ { print NR; exit }' "$ROOT/mm/damon/vaddr.c") +if [[ -n "$DISABLE_LINE" && -n "$DISARM_LINE" && "$DISABLE_LINE" -lt "$DIS= ARM_LINE" ]]; then + pass "Lifecycle: ops->disarm() after perf_event_disable()" +else + fail "Lifecycle: disarm() must follow perf_event_disable()" +fi + +# The stop path calls the common check after disarming. That common check +# must drain AUX before it starts consuming the SPSC rings. +DISARM_LOOP=3D$(grep -n 'damon_perf_event_disarm' "$ROOT/mm/damon/core.c" = 2>/dev/null | \ + tail -1 | cut -d: -f1) +FINAL_CHECK=3D$(awk -v start=3D"$DISARM_LOOP" 'NR > start && \ + /kdamond_check_reported_accesses\(ctx\)/ { print NR; exit }' \ + "$ROOT/mm/damon/core.c") +if [[ -n "$DISARM_LOOP" && -n "$FINAL_CHECK" && \ + "$DISARM_LOOP" -lt "$FINAL_CHECK" ]]; then + pass "Lifecycle: final common drain after disarm" +else + fail "Lifecycle: final common drain must follow disarm" +fi + +# Per-tick and final: AUX must publish before the SPSC loop reads rings. +CHECK_FN=3D$(grep -n '^static unsigned int kdamond_check_reported_accesses= ' \ + "$ROOT/mm/damon/core.c" | cut -d: -f1) +DRAIN_TICK=3D$(awk -v start=3D"$CHECK_FN" 'NR > start && \ + /damon_perf_aux_drain\(ctx\)/ { print NR; exit }' "$ROOT/mm/damon/core= .c") +RING_DRAIN=3D$(awk -v start=3D"$CHECK_FN" 'NR > start && \ + /for_each_online_cpu\(cpu\)/ { print NR; exit }' "$ROOT/mm/damon/core.= c") +if [[ -n "$DRAIN_TICK" && -n "$RING_DRAIN" && \ + "$DRAIN_TICK" -lt "$RING_DRAIN" ]]; then + pass "Tick: AUX drain before SPSC ring drain" +else + fail "Tick: AUX drain must precede the ring drain" +fi + +# Per-(event,cpu) parser state scoped via aux_priv. +if grep -q 'aux_priv' "$ROOT/mm/damon/ops-common.h" 2>/dev/null; then + pass "State scoping: per-(event,cpu) via aux_priv" +else + fail "State scoping: aux_priv in ops-common.h" +fi + +# SPSC ring contract: kdamond is the only process-context producer and +# writes only to its own CPU's ring. A remote-write helper must not +# exist. +if grep -q 'damon_report_access_on_cpu' "$ROOT/mm/damon/core.c" 2>/dev/nul= l; then + fail "SPSC: damon_report_access_on_cpu() must not exist" +else + pass "SPSC: no remote-writer helper (current-CPU NMI-safe enqueue only= )" +fi +if grep -q '^void damon_report_access' "$ROOT/mm/damon/core.c" 2>/dev/null= ; then + pass "SPSC: damon_report_access() present" +else + fail "SPSC: damon_report_access() missing" +fi +fi + +# ---- Section 4: Runtime smoke test ---- +echo "" +echo "--- 7. Runtime Smoke Test (live DAMON + SPE) ---" + +CREATED=3D0 +WORK_PID=3D"" +NR_SAVED=3D"" +if [[ -f "$ADMIN/kdamonds/nr_kdamonds" ]]; then + NR_SAVED=3D$(cat "$ADMIN/kdamonds/nr_kdamonds") +fi + +restore_damon_state() { + # Tear down only what this test created, then restore the saved + # number of kdamonds. A kdamond that was running before the test + # is never touched (section 7 skips in that case). + if [[ "$CREATED" =3D=3D "1" ]]; then + echo off > "$ADMIN/kdamonds/0/state" 2>/dev/null || true + echo 0 > "$ADMIN/kdamonds/nr_kdamonds" 2>/dev/null || true + if [[ -n "$NR_SAVED" && "$NR_SAVED" -gt 0 ]]; then + echo "$NR_SAVED" > "$ADMIN/kdamonds/nr_kdamonds" 2>/dev/null |= | true + fi + fi + if [[ -n "$WORK_PID" ]]; then + kill "$WORK_PID" 2>/dev/null || true + wait "$WORK_PID" 2>/dev/null || true + WORK_PID=3D"" + fi +} +trap restore_damon_state EXIT + +if [[ -z "$NR_SAVED" ]]; then + skip "runtime smoke" "DAMON admin interface not available" +elif [[ "$NR_SAVED" -gt 0 ]]; then + skip "runtime smoke" "existing kdamonds present (nr_kdamonds=3D$NR_SAV= ED);" \ + "refusing to disturb them" +elif [[ ! -d "/sys/bus/event_source/devices/$PMU_NAME" ]]; then + skip "runtime smoke" "no $PMU_NAME PMU (kernel without ARM SPE or not = booted with it)" +else + SPE_TYPE=3D$(cat "/sys/bus/event_source/devices/$PMU_NAME/type") + pass "SPE PMU $PMU_NAME present (type=3D$SPE_TYPE)" + DMESG_LINES_BEFORE=3D$(dmesg 2>/dev/null | wc -l) + + # Keep a userspace memory workload alive as both the DAMON target and + # an SPE data source. The EXIT trap owns and terminates only this PID. + dd if=3D/dev/zero of=3D/dev/null bs=3D1M 2>/dev/null & + WORK_PID=3D$! + + # arm_spe_pmu_event_init() rejects freq mode, so the event must be + # configured in period mode. + echo 1 > "$ADMIN/kdamonds/nr_kdamonds" + CREATED=3D1 + echo 1 > "$ADMIN/kdamonds/0/contexts/nr_contexts" + echo 1 > "$ADMIN/kdamonds/0/contexts/0/targets/nr_targets" + echo "$WORK_PID" > "$ADMIN/kdamonds/0/contexts/0/targets/0/pid_target" + echo 1 > "$ADMIN/kdamonds/0/contexts/0/monitoring_attrs/sample/perf_ev= ents/nr_perf_events" + echo "$SPE_TYPE" > "$ADMIN/kdamonds/0/contexts/0/monitoring_attrs/samp= le/perf_events/0/type" + echo 0 > "$ADMIN/kdamonds/0/contexts/0/monitoring_attrs/sample/perf_ev= ents/0/freq" + echo 256 > "$ADMIN/kdamonds/0/contexts/0/monitoring_attrs/sample/perf_= events/0/sample_period" + echo 3 > "$ADMIN/kdamonds/0/contexts/0/monitoring_attrs/sample/perf_ev= ents/0/config" + PE=3D"$ADMIN/kdamonds/0/contexts/0/monitoring_attrs/sample/perf_events" + FREQ_VAL=3D$(cat "$PE/0/freq" 2>/dev/null || true) + if [[ "$FREQ_VAL" =3D=3D "0" ]]; then + pass "SPE event in period mode (freq=3D0, period=3D256)" + else + fail "SPE event must use period mode (freq=3D0), read back $FREQ_V= AL" + fi + + STATS=3D"/sys/kernel/debug/damon/perf_stats" + if [[ -r "$STATS" ]]; then + cp "$STATS" "$RESULTS_DIR/perf_stats-before.log" + fi + + echo on > "$ADMIN/kdamonds/0/state" + sleep 5 + + STATE_NOW=3D$(cat "$ADMIN/kdamonds/0/state" 2>/dev/null || true) + if [[ "$STATE_NOW" =3D=3D "on" ]]; then + pass "kdamond with SPE event is running" + else + fail "kdamond with SPE event failed to start (state=3D$STATE_NOW)" + fi + + if [[ -r "$STATS" && -f "$RESULTS_DIR/perf_stats-before.log" ]]; then + cp "$STATS" "$RESULTS_DIR/perf_stats-running.log" + fi + + # Stop explicitly so the final snapshot covers disable -> AUX drain -> + # SPSC ring drain, rather than leaving that path to the EXIT trap. + echo off > "$ADMIN/kdamonds/0/state" + for _ in $(seq 1 50); do + [[ "$(cat "$ADMIN/kdamonds/0/state" 2>/dev/null || true)" =3D=3D "= off" ]] && break + sleep 0.1 + done + STATE_NOW=3D$(cat "$ADMIN/kdamonds/0/state" 2>/dev/null || true) + if [[ "$STATE_NOW" =3D=3D "off" ]]; then + pass "kdamond stopped after final drain" + else + fail "kdamond did not stop (state=3D$STATE_NOW)" + fi + + # Check only messages added during this run, not stale boot history. + DMESG_AFTER=3D"$RESULTS_DIR/dmesg-after.log" + dmesg 2>/dev/null > "$DMESG_AFTER" || true + DMESG_LINES_AFTER=3D$(wc -l < "$DMESG_AFTER") + if [[ "$DMESG_LINES_AFTER" -ge "$DMESG_LINES_BEFORE" ]]; then + NEW_DMESG=3D$(tail -n "+$((DMESG_LINES_BEFORE + 1))" "$DMESG_AFTER= ") + else + NEW_DMESG=3D$(cat "$DMESG_AFTER") + fi + FAIL_LINES=3D$(printf '%s\n' "$NEW_DMESG" | \ + grep -Ei "damon-perf.*(fail|warn|error)|WARNING:|BUG:|Oops:|lockde= p" || true) + if [[ -n "$FAIL_LINES" ]]; then + fail "no damon-perf failures in dmesg" + echo "$FAIL_LINES" | sed 's/^/ /' + else + pass "no damon-perf failures in dmesg" + fi + + if [[ -r "$STATS" && -f "$RESULTS_DIR/perf_stats-before.log" ]]; then + cp "$STATS" "$RESULTS_DIR/perf_stats-final.log" + pass "debugfs perf_stats readable" + + stat_value() { + awk -v name=3D"$2" '$1 =3D=3D name { print $2; exit }' "$1" + } + for counter in callback valid enqueue dequeue match update; do + before=3D$(stat_value "$RESULTS_DIR/perf_stats-before.log" "$c= ounter") + after=3D$(stat_value "$RESULTS_DIR/perf_stats-final.log" "$cou= nter") + if [[ "$before" =3D~ ^[0-9]+$ && "$after" =3D~ ^[0-9]+$ ]]; th= en + delta=3D$((after - before)) + else + delta=3D"" + fi + if [[ "$delta" =3D~ ^[0-9]+$ && "$delta" -gt 0 ]]; then + pass "AUX pipeline: $counter delta=3D$delta" + else + fail "AUX pipeline: expected positive $counter delta, got = ${delta:-missing}" + fi + done + + enqueue_before=3D$(stat_value "$RESULTS_DIR/perf_stats-before.log"= enqueue) + enqueue_final=3D$(stat_value "$RESULTS_DIR/perf_stats-final.log" e= nqueue) + dequeue_before=3D$(stat_value "$RESULTS_DIR/perf_stats-before.log"= dequeue) + dequeue_final=3D$(stat_value "$RESULTS_DIR/perf_stats-final.log" d= equeue) + enqueue_delta=3D$((enqueue_final - enqueue_before)) + dequeue_delta=3D$((dequeue_final - dequeue_before)) + if [[ "$enqueue_delta" -eq "$dequeue_delta" ]]; then + pass "final ring closure: enqueue=3D$enqueue_delta dequeue=3D$= dequeue_delta" + else + fail "final ring closure: enqueue=3D$enqueue_delta dequeue=3D$= dequeue_delta" + fi + else + fail "debugfs perf_stats unavailable; cannot validate AUX data pat= h" + fi + # Cleanup happens in the EXIT trap (restore_damon_state). +fi + +# ---- Summary ---- +echo "" +echo "=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D" +echo " SUMMARY: $PASSED passed, $FAILED failed, $SKIPPED skipped" +echo "=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D" +echo "Results saved to: $RESULTS_DIR" + +if [[ "$FAILED" -gt 0 ]]; then + echo "Overall: FAIL" + exit 1 +else + echo "Overall: PASS" + exit 0 +fi --=20 2.43.0