From nobody Mon Sep 28 20:12:34 2026 Received: from mail-pj1-f50.google.com (mail-pj1-f50.google.com [209.85.216.50]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E42F93E4510 for ; Tue, 18 Aug 2026 06:10:52 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.50 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787033454; cv=none; b=K7AKsHpjgPb4AlQNRSknAzlTzw4pXWC0pAJGsMd8gTgA7+5dJb/uN0NJcTEsTItUnssfu/uQ8h93HpXHXy9bUMzAxmxtm22aUKVVuRos3vDKqko1V0Lpu16jMyggAifRrqUyN7gDux0HaND8gIniqdKstOPVBBf90gDRHfjUViA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787033454; c=relaxed/simple; bh=J+lH5aIxoKZrtLS3G91f/QlNeqCm03+jL/fZVbeUkqA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=J1wHnRSe1JXnSjgWr2vt19Sr/Sn8s62+L/KXvAPDHO4FnU1dOJ+yHiLxYGGL4+sOdz+huoIvALYcZPdcKjKeMgYnuqGdFG1cc911ZBIOamzC/A65smf9JXhisGsdkl1eDWXaBenG/uKr5mbZFYemOrDg5nc3TuLxzfq02sh/hFY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=eatYc1Vw; arc=none smtp.client-ip=209.85.216.50 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="eatYc1Vw" Received: by mail-pj1-f50.google.com with SMTP id 98e67ed59e1d1-38e08baf860so4251829a91.2 for ; Mon, 17 Aug 2026 23:10:52 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787033452; x=1787638252; darn=vger.kernel.org; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:from:to:cc:subject :date:message-id:reply-to:content-type; bh=A673oz0E/o2MwRaOE5XvkcWpC/QdFWh299uzrbmuC70=; b=eatYc1Vw7M0ffV6mpH3kKhaZwUsBqmo9pKhWsoX0EpwQnq+Z4FpRF8rP7AxmzYuAcu GP2civu9peg7SIcaBywnQv6w5PqqZUIBf/MmbYjSPjN6W/AhJ71UCBllDzYILTVASi9H wjMq1nwMT9jqcQuzt7ch291bl5zOvcPL8EZGXD6CEYdpHPVMWnYvFtf+IXiElvnS/5Z/ hznZjQmMqXPP/1Ctk+oP1qbfFG0VQssKRaSqzlTpCCGAmGFxF0oUWyXGJVYW76jd1ivh 2Jsw83kEkcsiPSMs3LnOH4SfN9STZZmAKRNu/3AJ7OLOcsPW2ZZN0bGuB3JzqpXwO8pU BW2A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787033452; x=1787638252; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=A673oz0E/o2MwRaOE5XvkcWpC/QdFWh299uzrbmuC70=; b=kQ8Ib8RNmuHXo0tRZf+Pylq208dxIZRR7lBulaQahCwLOmBvycrOyo0vkAYNyQVhVs 47enJsDragoXyZrVRAbrvxAkyjIjHhRZdciIHnvomsm/0przcAHfMieJKo/8G4rLZKo+ pgcSZjqL6OzAZq20KBNu/nkqzEP4mIgI/L8zrPSnFlRZhtLNI9TcfpaP3uyAwryo1Iup PMBWV9Ji4JL4d3mxCt3qCh6sqZA5TdwDkMAvRA6rg8rbF0oeyOwwIMn4bcPDJc6XtnSf rRJpLwUPpX2qOhJ3i13+/OQeiU5p34C+lKwcCkGCgCa4MPt18OlZoohCnGxy252djqal Y8YQ== X-Forwarded-Encrypted: i=1; AHgh+RrVvQcy0TEII5vt3kTcH0jVkSr3K+X4xAgRHyKn4BeslhbdIUEQTRVmHTgQZqO5lUqv2kRL8ZXFHs/y/hs=@vger.kernel.org X-Gm-Message-State: AOJu0Yxe/lUbqzYLpdE18tt85dxFEDtBBuTOgm8BQ+AN1wveXi783+gR 0Q1OAGTju7L7t8OlnG/ly/IebqKNRU6/bjheaARWMAIucSOHPVdBvD5g X-Gm-Gg: AR+sD137edIGaRupO79+cIjQ2u3igVPPx3H1jUHlseBfDa37yUtUz7ILAIXXHZaC/Ay /oiHuLvKpWw8PNHKw+LmZmcZAef50a04Uo4zHJwZMmR+viNCirDDwrr6g4iWNuJTeT+a9rwNFg1 8DW139W2Lz7wyVQfmcfmY8OPSVIDOgr1vrf46ovNGd1T/WuAli8++eoDxpgOn9DAkMQkEBExFwE Uk9QUh8BCvVgYt+S90WxQtoCzmhrKwUgx+Ifnet+ksH2AbA7YFImwLD047oYdntIo0TVVSZo2cN jb5nKzko4utJ9osH7joV0VAdOUAHc0iDOsSJT/rtD7IIrwNvpZQSVU10QH+reV2lYzmQMyOECo3 /GMSkoMloKXuGrbZ2DDKkjGa9TXcmVP0kcpu5dAEvBcM8VCgmJAzS1Ws4z1SGbCtJhUQ4xE4/sZ OLoqo9E/yQ6B5sIoVhdoWPZn0BpDU2avCPXEXMsYYKSy+17GrjIcLjnHPsSPBkVtBbtUiKejV1M 1clVe/g+EKns2LCEQ== X-Received: by 2002:a17:90b:4f84:b0:38e:fea2:df53 with SMTP id 98e67ed59e1d1-3933bd186f4mr32421619a91.4.1787033451538; Mon, 17 Aug 2026 23:10:51 -0700 (PDT) Received: from kernel.tail6741c6.ts.net ([116.128.244.169]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-3954d38bf94sm4633942a91.9.2026.08.17.23.10.45 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 17 Aug 2026 23:10:51 -0700 (PDT) From: Kunwu Chan X-Google-Original-From: Kunwu Chan To: sj@kernel.org, akpm@linux-foundation.org Cc: damon@lists.linux.dev, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, rostedt@goodmis.org, mhiramat@kernel.org, mathieu.desnoyers@efficios.com, shuah@kernel.org, lianux.mm@gmail.com, Kunwu Chan Subject: [RFC PATCH 1/7] mm/damon/perf: add observability framework with tracepoints and CONFIG switch Date: Tue, 18 Aug 2026 14:10:25 +0800 Message-ID: <20260818061031.827057-2-kunwu.chan@linux.dev> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260818061031.827057-1-kunwu.chan@linux.dev> References: <20260818061031.827057-1-kunwu.chan@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable From: Kunwu Chan Add the observe framework headers together with the compile-time switch that gates the whole feature: the three DAMON perf tracepoints this series adds (guarded with CONFIG_DAMON_PERF_OBSERVE so they are not registered when the switch is off; damon_perf_ring_overflow is provided by the base series), the observe API declarations with static-inline no-ops for disabled builds, the access-report contract (miss-reason enum, report-source enum, report source field, per-event cpu_state member), and the Kconfig and Makefile wiring. The sample tracepoint carries what the PMU actually populated (data->sample_flags), what was requested (perf_event->attr. sample_type), and the execution context (process/softirq/hardirq/NMI) in a single line, so PMU support gaps and context expectations (e.g. IBS overflow in NMI, SPE AUX drain in process context) are verifiable at a glance. Every following commit in this series builds with CONFIG_DAMON_PERF_OBSERVE both enabled and disabled. Co-developed-by: Lian Wang Signed-off-by: Lian Wang Signed-off-by: Kunwu Chan --- include/linux/damon.h | 44 +++++++ include/trace/events/damon.h | 120 ++++++++++++++++++ mm/damon/Kconfig | 17 +++ mm/damon/Makefile | 1 + mm/damon/perf/Makefile | 3 + mm/damon/perf/perf.h | 228 +++++++++++++++++++++++++++++++++++ 6 files changed, 413 insertions(+) create mode 100644 mm/damon/perf/Makefile create mode 100644 mm/damon/perf/perf.h diff --git a/include/linux/damon.h b/include/linux/damon.h index 11f1c1071b9b..c191c065b0e4 100644 --- a/include/linux/damon.h +++ b/include/linux/damon.h @@ -116,6 +116,19 @@ struct damon_target { bool obsolete; }; =20 +/** + * enum damon_report_source - Tells which subsystem produced an access rep= ort. + * + * Ring and matching counters aggregate all sources; this enum lets callers + * tag reports so that tracepoints and future per-source breakdowns can + * distinguish NMI overflow-handler samples from + * page-fault hints. + */ +enum damon_report_source { + DAMON_REPORT_SRC_PERF_OVERFLOW =3D 0, /* overflow_handler (IBS, PEBS) */ + DAMON_REPORT_SRC_PAGE_FAULT, /* damon_report_page_fault() */ +}; + /** * struct damon_access_report - Represent single access report information. * @paddr: Start physical address of the accessed address range. @@ -125,6 +138,8 @@ struct damon_target { * @tid: The task id of the task that made the access. * @tgid: Thread group id of the task that made the access. * @is_write: Whether the access is write. + * @source: Which subsystem produced this report + * (enum damon_report_source). * * Any DAMON API callers that notified access events can report the inform= ation * to DAMON using damon_report_access(). This struct contains the reporti= ng @@ -138,10 +153,28 @@ struct damon_access_report { pid_t tid; pid_t tgid; bool is_write; +#ifdef CONFIG_DAMON_PERF_OBSERVE + int source; +#endif /* CONFIG_DAMON_PERF_OBSERVE */ /* private: */ unsigned long report_jiffies; /* when this report is made */ }; =20 +/* + * Reason codes for trace_damon_perf_report_missed. + * + * DAMON_REPORT_MISS_TGID: tgid mismatch (pid-based monitoring, + * missed at drain-loop level before the + * per-target iteration). + * DAMON_REPORT_MISS_NOREGION: binary search found no containing region. + * DAMON_REPORT_MISS_BOUNDARY: address + size straddles region boundary. + */ +enum damon_report_miss_reason { + DAMON_REPORT_MISS_TGID =3D 1, + DAMON_REPORT_MISS_NOREGION =3D 2, + DAMON_REPORT_MISS_BOUNDARY =3D 3, +}; + /** * enum damos_action - Represents an action of a Data Access Monitoring-ba= sed * Operation Scheme. @@ -1027,6 +1060,17 @@ struct damon_perf_event { struct hlist_node hlist_node; bool init_complete; bool any_cpu_failed; +#ifdef CONFIG_DAMON_PERF_OBSERVE + /* + * Per-CPU lifecycle state (enum damon_perf_event_state). + * Allocated lazily on the first observe_event_created(), + * freed on observe_event_destroyed(). Each event tracks + * its own progression through CREATED->BOUND->ENABLED, + * so destroying one event does not overwrite another"s + * state on the same CPU. + */ + int __percpu *cpu_state; +#endif /* CONFIG_DAMON_PERF_OBSERVE */ struct damon_ctx *ctx; }; =20 diff --git a/include/trace/events/damon.h b/include/trace/events/damon.h index 877627c9a1a1..c87fbeefb85a 100644 --- a/include/trace/events/damon.h +++ b/include/trace/events/damon.h @@ -91,6 +91,126 @@ TRACE_EVENT(damon_perf_ring_overflow, TP_printk("cpu=3D%d", __entry->cpu) ); =20 +#ifdef CONFIG_DAMON_PERF_OBSERVE +/* + * Fires from NMI overflow handlers on every hardware sample received, + * before any DAMON-side filtering. Records the raw address, full + * data_src (mem_op, mem_lvl, mem_snoop, mem_remote), period, and a + * reason code so userspace can distinguish: + * + * 0 =3D valid sample, queued to per-CPU ring + * 1 =3D data =3D=3D NULL + * 2 =3D addr =3D=3D 0 (PMU did not populate data->addr) + * 3 =3D kernel address (vaddr handler: addr >=3D TASK_SIZE) + * 4 =3D phys_addr not valid (paddr handler: !PERF_SAMPLE_PHYS_ADDR) + * + * data_src carries the raw union perf_mem_data_src value; use + * perf_mem__xxx macros to decode. + * + * sample_flags is what the PMU *actually* populated (from + * data->sample_flags); sample_type is what was *requested* (from + * perf_event->attr.sample_type). Comparing them immediately + * reveals whether the PMU is providing the fields DAMON asked for + * =E2=80=94 e.g. sample_type has PERF_SAMPLE_PHYS_ADDR but sample_flags + * does not =E2=86=92 the PMU does not support physical-address sampling. + */ +TRACE_EVENT(damon_perf_sample, + + TP_PROTO(unsigned long addr, u64 data_src, u64 period, int cpu, + u8 reason, u64 sample_flags, u64 sample_type, + u8 context), + + TP_ARGS(addr, data_src, period, cpu, reason, sample_flags, + sample_type, context), + + TP_STRUCT__entry( + __field(unsigned long, addr) + __field(u64, data_src) + __field(u64, period) + __field(int, cpu) + __field(u8, reason) + __field(u64, sample_flags) + __field(u64, sample_type) + __field(u8, context) + ), + + TP_fast_assign( + __entry->addr =3D addr; + __entry->data_src =3D data_src; + __entry->period =3D period; + __entry->cpu =3D cpu; + __entry->reason =3D reason; + __entry->sample_flags =3D sample_flags; + __entry->sample_type =3D sample_type; + __entry->context =3D context; + ), + + TP_printk("addr=3D0x%lx data_src=3D0x%llx period=3D%llu cpu=3D%d reason= =3D%u context=3D%u sample_flags=3D0x%llx sample_type=3D0x%llx", + __entry->addr, __entry->data_src, __entry->period, + __entry->cpu, __entry->reason, __entry->context, + __entry->sample_flags, __entry->sample_type) +); + +/* + * Fires when a report survived all ring/drain checks but could not be + * applied to any DAMON region. Reasons correspond to + * enum damon_report_miss_reason: + * + * DAMON_REPORT_MISS_TGID (1): no target matched the report's tgid + * DAMON_REPORT_MISS_NOREGION (2): binary search found no containing reg= ion + * DAMON_REPORT_MISS_BOUNDARY (3): address + size straddles region bound= ary + * + * Note: tgid mismatches are now resolved in the drain loop *before* + * iterating targets, so there is at most one trace hit per report + * (rather than one per non-matching target as in earlier revisions). + */ +TRACE_EVENT(damon_perf_report_missed, + + TP_PROTO(unsigned long addr, int cpu, int reason), + + TP_ARGS(addr, cpu, reason), + + TP_STRUCT__entry( + __field(unsigned long, addr) + __field(int, cpu) + __field(int, reason) + ), + + TP_fast_assign( + __entry->addr =3D addr; + __entry->cpu =3D cpu; + __entry->reason =3D reason; + ), + + TP_printk("addr=3D0x%lx cpu=3D%d reason=3D%d", __entry->addr, + __entry->cpu, __entry->reason) +); + +/* + * Per-tick drain summary. Fires from kdamond after draining the per-CPU + * SPSC ring, so users can observe total vs matched without polling dmesg + * or correlating individual miss tracepoints. + */ +TRACE_EVENT(damon_perf_drain, + + TP_PROTO(unsigned int total, unsigned int matched), + + TP_ARGS(total, matched), + + TP_STRUCT__entry( + __field(unsigned int, total) + __field(unsigned int, matched) + ), + + TP_fast_assign( + __entry->total =3D total; + __entry->matched =3D matched; + ), + + TP_printk("total=3D%u matched=3D%u", __entry->total, __entry->matched) +); +#endif /* CONFIG_DAMON_PERF_OBSERVE */ + /* Per-tick DAMOS_QUOTA_NODE_ELIGIBLE_MEM_BP goal evaluation. */ TRACE_EVENT(damos_node_eligible_mem_bp, =20 diff --git a/mm/damon/Kconfig b/mm/damon/Kconfig index ad629f0f31d8..9f811510760f 100644 --- a/mm/damon/Kconfig +++ b/mm/damon/Kconfig @@ -131,4 +131,21 @@ config DAMON_ACMA min/max memory for the system and maximum memory pressure stall time ratio. =20 +config DAMON_PERF_OBSERVE + bool "DAMON perf event observability framework" + depends on DAMON + depends on PERF_EVENTS + depends on DEBUG_FS + default n + help + Enable per-CPU pipeline counters, tracepoints, and a + debug-only debugfs perf_stats file for DAMON + hardware-sampled access reports. The debugfs format is + unstable and must not be used by scripts; counters and + tracepoints are the diagnostic interface. + + When disabled, all observe functions are compiled to + static-inline no-ops with zero runtime overhead. + + If unsure, say N. endmenu diff --git a/mm/damon/Makefile b/mm/damon/Makefile index 22494754f41e..04da39a9f56c 100644 --- a/mm/damon/Makefile +++ b/mm/damon/Makefile @@ -9,3 +9,4 @@ obj-$(CONFIG_DAMON_RECLAIM) +=3D modules-common.o reclaim.o obj-$(CONFIG_DAMON_LRU_SORT) +=3D modules-common.o lru_sort.o obj-$(CONFIG_DAMON_STAT) +=3D modules-common.o stat.o obj-$(CONFIG_DAMON_ACMA) +=3D modules-common.o acma.o +obj-$(CONFIG_DAMON) +=3D perf/ diff --git a/mm/damon/perf/Makefile b/mm/damon/perf/Makefile new file mode 100644 index 000000000000..cc0d4f1d1d28 --- /dev/null +++ b/mm/damon/perf/Makefile @@ -0,0 +1,3 @@ +# SPDX-License-Identifier: GPL-2.0 + +# Observability: per-CPU counters, tracepoints, debugfs perf_stats diff --git a/mm/damon/perf/perf.h b/mm/damon/perf/perf.h new file mode 100644 index 000000000000..78e23d436336 --- /dev/null +++ b/mm/damon/perf/perf.h @@ -0,0 +1,228 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * DAMON Hardware-sampled Access Report Observability Framework + * + * Single entry-point for all hardware sampling backends (ARM SPE, + * AMD IBS, Intel PEBS, =E2=80=A6). Every event, sample, ring operation, + * match decision, and region update flows through the + * damon_perf_observe_*() API, which fans out to per-CPU counters + * and tracepoints. When CONFIG_DAMON_PERF_OBSERVE=3Dn, everything + * compiles to static-inline no-ops. + * + * Author: Kunwu Chan + */ + +#ifndef _DAMON_PERF_H +#define _DAMON_PERF_H + +struct perf_event; +#include + +struct damon_perf_event; + +/* + * Per-event state machine + * + * Each per-CPU damon_perf_event transitions through these states. + * State is tracked in a per-CPU integer (damon_perf_cpu_state). + */ +enum damon_perf_event_state { + DAMON_PERF_STATE_UNINIT =3D 0, + DAMON_PERF_STATE_CREATED, /* struct allocated, cpuhp registered */ + DAMON_PERF_STATE_BOUND, /* perf_event_create_kernel_counter() ok */ + DAMON_PERF_STATE_ENABLED, /* perf_event_enable() called */ + DAMON_PERF_STATE_RUNNING, /* first overflow callback received */ + DAMON_PERF_STATE_ERROR, /* unrecoverable failure */ +}; + +/* + * Per-CPU statistics + * + * All counters are monotonic, best-effort reads. Userspace computes + * deltas between snapshots. Stored per-CPU so the NMI fast path uses + * this_cpu_inc() with no locking. Counter values are raw facts: + * interpretation (thresholds, verdicts) belongs in userspace. + * + * The kernel provides tracepoints under events/damon/ for structured, + * stable diagnostics. The debugfs perf_stats file is DEBUG ONLY and + * its format may change without notice. + */ +struct damon_perf_stats { + /* Per-CPU event state (enum damon_perf_event_state) */ + int cpu_state; + + /* Sampling pipeline */ + u64 callback; + u64 sample_valid; + u64 sample_null; + u64 sample_addr_zero; + u64 sample_kernel; + u64 sample_invalid_phys; + + /* Ring */ + u64 enqueue; + u64 dequeue; + u64 overflow; + u64 ring_peak; + + /* Matching */ + u64 match; + u64 miss_tgid; + u64 miss_region; + u64 miss_boundary; + u64 update; +}; + +#ifdef CONFIG_DAMON_PERF_OBSERVE + +/* + * damon_perf_observe_*() =E2=80=94 Unified Observability API + * + * These are the ONLY hooks that hardware-sampling backends should + * call. They are split into NMI-safe (sampling, ring-enqueue) and + * process-context (event lifecycle, drain, matching, update) groups. + * + * Counters always increment when CONFIG_DAMON_PERF_OBSERVE=3Dy. + * Tracepoints are guarded by trace_*_enabled() and incur zero + * overhead when ftrace is not attached. + */ + +/* Event lifecycle =E2=80=94 process context (kdamond / cpuhp callbacks) */ +void damon_perf_observe_event_created(struct damon_perf_event *event, + int cpu); +void damon_perf_observe_event_bound(struct damon_perf_event *event, + int cpu, struct perf_event *perf_event); +void damon_perf_observe_event_enabled(struct damon_perf_event *event, + int cpu, int state, int oncpu); +void damon_perf_observe_event_disabled(struct damon_perf_event *event, + int cpu, int state); +void damon_perf_observe_event_destroyed(struct damon_perf_event *event, + int cpu); +void damon_perf_observe_event_free(struct damon_perf_event *event); + +/* + * Sample observed =E2=80=94 NMI-safe. + * + * @reason: 0 =3D valid (queued to ring) + * 1 =3D data NULL + * 2 =3D addr =3D=3D 0 (PMU did not populate) + * 3 =3D kernel address (vaddr only) + * 4 =3D phys_addr not valid (paddr only) + */ +void damon_perf_observe_sample(unsigned long addr, u64 data_src, + u64 period, int cpu, u8 reason, + u64 sample_flags, u64 sample_type); + +/* Ring operations =E2=80=94 enqueue/overflow are NMI-safe */ +void damon_perf_observe_ring_enqueue(void); +void damon_perf_observe_ring_overflow(int cpu); +void damon_perf_observe_ring_dequeue(int cpu); +void damon_perf_observe_ring_peak(unsigned int occupancy); + +/* Matching =E2=80=94 process context (kdamond drain loop) */ +void damon_perf_observe_match(unsigned long addr, int cpu); +void damon_perf_observe_miss(unsigned long addr, int cpu, int reason); +void damon_perf_observe_update(int cpu); +void damon_perf_observe_drain(unsigned int total, unsigned int matched); + +/* Debugfs (debug-only, format unstable) */ +int damon_perf_debugfs_init(void); + +/* Per-CPU stats accessors (for debugfs) */ +void damon_perf_stats_snapshot(int cpu, struct damon_perf_stats *dst); +void damon_perf_stats_aggregate(struct damon_perf_stats *dst); + +/* Subsystem init */ +int damon_perf_framework_init(void); + +#else /* !CONFIG_DAMON_PERF_OBSERVE */ + +static inline void damon_perf_observe_event_created(struct damon_perf_even= t *e, + int c) +{ +} + +static inline void damon_perf_observe_event_bound(struct damon_perf_event = *e, + int c, struct perf_event *p) +{ +} + +static inline void damon_perf_observe_event_enabled(struct damon_perf_even= t *e, + int c, int s, int o) +{ +} + +static inline void damon_perf_observe_event_disabled(struct damon_perf_eve= nt *e, + int c, int s) +{ +} + +static inline void damon_perf_observe_event_destroyed(struct damon_perf_ev= ent *e, + int c) +{ +} + +static inline void damon_perf_observe_event_free(struct damon_perf_event *= e) +{ +} + +static inline void damon_perf_observe_sample(unsigned long a, u64 d, u64 p, + int c, u8 r, u64 f, u64 t) +{ +} + +static inline void damon_perf_observe_ring_enqueue(void) +{ +} + +static inline void damon_perf_observe_ring_overflow(int c) +{ +} + +static inline void damon_perf_observe_ring_dequeue(int c) +{ +} + +static inline void damon_perf_observe_ring_peak(unsigned int o) +{ +} + +static inline void damon_perf_observe_match(unsigned long a, int c) +{ +} + +static inline void damon_perf_observe_miss(unsigned long a, int c, int r) +{ +} + +static inline void damon_perf_observe_update(int c) +{ +} + +static inline void damon_perf_observe_drain(unsigned int t, unsigned int m) +{ +} + +static inline int damon_perf_debugfs_init(void) +{ + return 0; +} + +static inline void damon_perf_stats_snapshot(int c, struct damon_perf_stat= s *d) +{ + memset(d, 0, sizeof(*d)); +} + +static inline void damon_perf_stats_aggregate(struct damon_perf_stats *d) +{ + memset(d, 0, sizeof(*d)); +} + +static inline int damon_perf_framework_init(void) +{ + return 0; +} + +#endif /* CONFIG_DAMON_PERF_OBSERVE */ + +#endif /* _DAMON_PERF_H */ --=20 2.43.0 From nobody Mon Sep 28 20:12:34 2026 Received: from mail-pl1-f174.google.com (mail-pl1-f174.google.com [209.85.214.174]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1132F3E6392 for ; Tue, 18 Aug 2026 06:10:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.174 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787033460; cv=none; b=N702DvX5gn6Y6CHCYSmDQEOMrrjfjkg7Mbzk5ios/onY6d4xUfCcq9X9tPKb/QmECo9qZQrKGFQmQbvPuBi8bmSQdDuANm2KHi1gkmvueHUrtRArltLtdShQIsFwsbTjF9qnPqW58jzjSqweaGw8VEjUIvQhvlVxBo6gHpsoAfw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787033460; c=relaxed/simple; bh=6pTbLyujulLWUYG01436SvB+x/PQR2vhTC8I2dX7k/g=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=XwwvwEQUkhW/i5S0vlIggV6mPmJv2hi43jfs5561CFuEHb9eWRMyB+xtaLREtDL/2ESU7XHKRa9Wb/uhGwWk1MaY91ugZK0X9Eux3jCi7DhlT3tDD9BJ42Xwc2uIFrWQ1RDCA7PfjWqdTNzG5kLbUHn3WA5m5bkG+sxdsBSlYMo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=dFn+6mkV; arc=none smtp.client-ip=209.85.214.174 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="dFn+6mkV" Received: by mail-pl1-f174.google.com with SMTP id d9443c01a7336-2d5655cc850so28553295ad.3 for ; Mon, 17 Aug 2026 23:10:58 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787033458; x=1787638258; darn=vger.kernel.org; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:from:to:cc:subject :date:message-id:reply-to:content-type; bh=DOBdbjG7wg0kpAdHn038dSyaMR4LfRYkZkjGzLFsJjk=; b=dFn+6mkV+9gAQCEW/+EE8jzGTekb5Gwgb80KhwdeqjuH6IKalJqExIP2KmcjfDvPbD dKbNPXAZXHXk7S2VeOEFPyDyfas3HA8te80QA97UIsQSzA2fh/5/TTD3YzsbkNoREzPL //zVKvA1HjVePNpOOBVn2u+K3fygHnLMevIyr9BJ0KkA9iMxVsVp/8vKhCWAKv7ojCMT fVU4REJwCbvtq3HSrj8FcEPm55H1HkG6MPdLjWLzoAqvhtruZDtvESGU/LS3lbPutnGs sLrWX8k84lb41XV1H0xvVlnWKcAhAxLkT/83OYqp/DSSJtYK2sEB689VJGS53G7iYWXR fXxA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787033458; x=1787638258; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=DOBdbjG7wg0kpAdHn038dSyaMR4LfRYkZkjGzLFsJjk=; b=ZYHVb2YpzcVxAimdDJfld0JbDsXcHebipdWVfd8CHhrFgkEZGtRx2tY/J4wMhXrUBB MC6/BLksLSGOqnhQKs1uAfMR74VS5R9R7iR2zVi3Ytdlc4m5fmZ/nPosaOvQjpEAxkLq fkz63+yQtKx7iPOMVN1XdjI5h0eaPdrbYsztyaWS4DX/J82hnITigU1BJgynbMRqADWN u/y9u1udl0HLRxLwa8YTYUNhGMrhQW6OmZPTa2xuek2Bfhtqzej2CdHUqFi03WY5n1l5 cfDhX7WzR8+GfggUrSkyXZ+Kssr5m5cLVi8mbzTobY7gHombdNQh7jxm4X5tWQz7uWfp XnBQ== X-Forwarded-Encrypted: i=1; AHgh+Rqmu13l7o75B9b3xk1XTErU/KVlCpybMaugYE/z3igps7cCG4SyqrVHVdHAB+tYaH8iopqoeOClaesKhsM=@vger.kernel.org X-Gm-Message-State: AOJu0Yy+IqkCxJRukJ2NpOES5tYac94PjXZhCEBhXYlReG25RYmo3RBg 78Velxbta4vLt8BSllzUzcANRfABc+Nap/Q6+lYvuz/zdlTViXI6CAiZ X-Gm-Gg: AR+sD12RcOOaYdUZpFSUjE7OOwrhitUtsu4zIf/1g3dCI/a3gd12n5ZoEqdrXPowcFs np5Rg5wgxJwaQDOYMls+7GrIuHwHu0czTB+yyZ57TMjrrTFOLAp1gzydjasssRHpqIU2eV9soRD KD9sQkmA/BYU5FMQuRmC6quIGnAbmjUtT/ht/XSXvNb5R/KVO7d8kLKWNL2zQ9Rkd4pNXNt552s mrFgmxoP+6ZR3IjyQaGD64Ij4U67L6XGUZPV06yIZh2I+baniNJNnK3zubs7NltfU3QmPn+7gq8 v/zVs7NayllFcsURTQUJ/mjnosF80o9LA3nAsWCtbgFAxjREC+Pyo3bVFhkp1JFLO4+dsR6sEuJ JHyXVkJHHjsNZ/14hOKVGPbNMRxRJ2WYOny2JPJzhOf9/XmTX2QSmDMXRMqrD0Ib/uvGrEIWfPd vFySxiaGDBg9oWJ1eoCWy+FvaATqgniBrKWagCOcW3EENMti44fmomSKpF/XnXSeVCs/DWxBul1 4K+KJ0= X-Received: by 2002:a17:90b:5685:b0:38e:4cb:51f with SMTP id 98e67ed59e1d1-3955a799fcdmr8128358a91.11.1787033458228; Mon, 17 Aug 2026 23:10:58 -0700 (PDT) Received: from kernel.tail6741c6.ts.net ([116.128.244.169]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-3954d38bf94sm4633942a91.9.2026.08.17.23.10.51 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 17 Aug 2026 23:10:57 -0700 (PDT) From: Kunwu Chan X-Google-Original-From: Kunwu Chan To: sj@kernel.org, akpm@linux-foundation.org Cc: damon@lists.linux.dev, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, rostedt@goodmis.org, mhiramat@kernel.org, mathieu.desnoyers@efficios.com, shuah@kernel.org, lianux.mm@gmail.com, Kunwu Chan Subject: [RFC PATCH 2/7] mm/damon/perf: implement observe API and per-CPU statistics engine Date: Tue, 18 Aug 2026 14:10:26 +0800 Message-ID: <20260818061031.827057-3-kunwu.chan@linux.dev> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260818061031.827057-1-kunwu.chan@linux.dev> References: <20260818061031.827057-1-kunwu.chan@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable From: Kunwu Chan Implement the observe_*() core: per-CPU monotonic counters, the event-lifecycle hooks with lazily-allocated per-event per-CPU state alongside the global per-CPU stage, NMI-safe sample classification, ring enqueue/dequeue/peak, match/miss/update, and framework init. The global per-CPU stage advances monotonically so a late create/bind/ enable cannot regress a CPU that is already RUNNING. damon_perf_observe_sample() records the execution context (process/softirq/hardirq/NMI) so overflow callbacks can be verified to fire in the expected context. Co-developed-by: Lian Wang Signed-off-by: Lian Wang Signed-off-by: Kunwu Chan --- mm/damon/perf/Makefile | 2 + mm/damon/perf/stats.c | 295 +++++++++++++++++++++++++++++++++++++++++ 2 files changed, 297 insertions(+) create mode 100644 mm/damon/perf/stats.c diff --git a/mm/damon/perf/Makefile b/mm/damon/perf/Makefile index cc0d4f1d1d28..5c46d3da7ef8 100644 --- a/mm/damon/perf/Makefile +++ b/mm/damon/perf/Makefile @@ -1,3 +1,5 @@ # SPDX-License-Identifier: GPL-2.0 =20 # Observability: per-CPU counters, tracepoints, debugfs perf_stats +obj-$(CONFIG_DAMON_PERF_OBSERVE) +=3D damon-perf.o +damon-perf-objs :=3D stats.o diff --git a/mm/damon/perf/stats.c b/mm/damon/perf/stats.c new file mode 100644 index 000000000000..a869f115bb26 --- /dev/null +++ b/mm/damon/perf/stats.c @@ -0,0 +1,295 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * DAMON Perf Observability =E2=80=94 Per-CPU Statistics & Observe API + * + * All hardware-sampling backends funnel through the observe functions + * defined here. Counters always increment when + * CONFIG_DAMON_PERF_OBSERVE=3Dy; tracepoints are guarded by + * trace_*_enabled() for zero overhead when ftrace is not attached. + * + * Per-CPU counters are best-effort: individual u64 writes are atomic + * on 64-bit platforms, but no cross-field consistency is guaranteed. + * Snapshot reads may race with concurrent writers. This is debug data + * =E2=80=94 do not build policy on it. + */ + +#include +#include + +#include + +#include "perf.h" + +/* + * Per-CPU statistics + */ +static DEFINE_PER_CPU(struct damon_perf_stats, damon_perf_stats); + +/* + * Per-CPU event state for the state machine =E2=80=94 the maximum lifecyc= le + * stage ever reached on each CPU (any event). Set by the NMI + * observe_sample() (ENABLED=E2=86=92RUNNING) and by lifecycle functions a= s a + * monotonic "best CPU so far". Never reset to UNINIT =E2=80=94 destroyin= g one + * event must not hide that another event is still running on the same + * CPU. Per-event per-CPU state (event->cpu_state) tracks individual + * event lifecycle for correctness when multiple events exist. + */ +static DEFINE_PER_CPU(int, damon_perf_cpu_state); + +/* + * Event lifecycle =E2=80=94 process context (kdamond / cpuhp callbacks). + * + * Each event tracks its own per-CPU lifecycle state via + * event->cpu_state (allocated on first CREATED, freed on DESTROYED). + * The global damon_perf_cpu_state is also advanced (monotonic) so that + * debugfs perf_stats shows the "best" stage any event reached on each + * CPU. Diagnostics are emitted via tracepoints, not dmesg. + */ + +/* + * Advance the global per-CPU lifecycle state monotonically. The global + * reflects the best stage any event ever reached on this CPU; a late + * create/bind/enable must not regress a CPU that is already RUNNING. + */ +static void damon_perf_cpu_state_advance(int cpu, int state) +{ + int cur =3D READ_ONCE(per_cpu(damon_perf_cpu_state, cpu)); + + if (state > cur) + WRITE_ONCE(per_cpu(damon_perf_cpu_state, cpu), state); +} + +void damon_perf_observe_event_created(struct damon_perf_event *event, int = cpu) +{ + if (!event->cpu_state) { + event->cpu_state =3D alloc_percpu(int); + if (!event->cpu_state) + return; + } + *per_cpu_ptr(event->cpu_state, cpu) =3D DAMON_PERF_STATE_CREATED; + damon_perf_cpu_state_advance(cpu, DAMON_PERF_STATE_CREATED); +} + +void damon_perf_observe_event_bound(struct damon_perf_event *event, + int cpu, struct perf_event *perf_event) +{ + if (event->cpu_state) + *per_cpu_ptr(event->cpu_state, cpu) =3D DAMON_PERF_STATE_BOUND; + damon_perf_cpu_state_advance(cpu, DAMON_PERF_STATE_BOUND); +} + +void damon_perf_observe_event_enabled(struct damon_perf_event *event, + int cpu, int state, int oncpu) +{ + if (event->cpu_state) + *per_cpu_ptr(event->cpu_state, cpu) =3D DAMON_PERF_STATE_ENABLED; + damon_perf_cpu_state_advance(cpu, DAMON_PERF_STATE_ENABLED); +} + +void damon_perf_observe_event_disabled(struct damon_perf_event *event, + int cpu, int state) +{ + /* State unchanged: the event may be re-enabled later. */ +} + +void damon_perf_observe_event_destroyed(struct damon_perf_event *event, in= t cpu) +{ + /* + * Reset only this CPU's slot. The per-CPU array is owned by the + * event as a whole and must NOT be freed here: destroyed() is + * invoked from the per-CPU CPU-offline callback, so freeing the + * whole array on the first offline CPU would leave every other + * still-online CPU with a dangling pointer. The array is freed + * once by damon_perf_observe_event_free() at event teardown. + */ + if (event->cpu_state) + *per_cpu_ptr(event->cpu_state, cpu) =3D DAMON_PERF_STATE_UNINIT; +} + +void damon_perf_observe_event_free(struct damon_perf_event *event) +{ + if (event->cpu_state) { + free_percpu(event->cpu_state); + event->cpu_state =3D NULL; + } +} + +/* + * Sample observed =E2=80=94 NMI-safe. + * + * The @cpu argument is the CPU the sample fired on; it is passed + * through to the tracepoint but stats are always written on the + * current CPU via this_cpu_ptr(). + */ + +void damon_perf_observe_sample(unsigned long addr, u64 data_src, + u64 period, int cpu, u8 reason, + u64 sample_flags, u64 sample_type) +{ + this_cpu_inc(damon_perf_stats.callback); + + switch (reason) { + case 0: + this_cpu_inc(damon_perf_stats.sample_valid); + break; + case 1: + this_cpu_inc(damon_perf_stats.sample_null); + break; + case 2: + this_cpu_inc(damon_perf_stats.sample_addr_zero); + break; + case 3: + this_cpu_inc(damon_perf_stats.sample_kernel); + break; + case 4: + this_cpu_inc(damon_perf_stats.sample_invalid_phys); + break; + } + + /* First callback advances state to RUNNING */ + if (this_cpu_read(damon_perf_cpu_state) =3D=3D DAMON_PERF_STATE_ENABLED) + this_cpu_write(damon_perf_cpu_state, DAMON_PERF_STATE_RUNNING); + + /* + * Record the execution context so callers can verify whether + * overflow callbacks actually fire in the expected context + * (e.g. NMI for IBS, process for SPE AUX drain). + * + * 0 =3D process, 1 =3D softirq, 2 =3D hardirq, 3 =3D NMI + */ + if (trace_damon_perf_sample_enabled()) + trace_damon_perf_sample(addr, data_src, period, cpu, reason, + sample_flags, sample_type, + in_nmi() ? 3 : in_hardirq() ? 2 : + in_serving_softirq() ? 1 : 0); +} + +/* + * Ring operations =E2=80=94 enqueue / overflow are NMI-safe + */ + +void damon_perf_observe_ring_enqueue(void) +{ + this_cpu_inc(damon_perf_stats.enqueue); +} + +void damon_perf_observe_ring_overflow(int cpu) +{ + this_cpu_inc(damon_perf_stats.overflow); + if (trace_damon_perf_ring_overflow_enabled()) + trace_damon_perf_ring_overflow(cpu); +} + +void damon_perf_observe_ring_dequeue(int cpu) +{ + per_cpu_ptr(&damon_perf_stats, cpu)->dequeue++; +} + +void damon_perf_observe_ring_peak(unsigned int occupancy) +{ + struct damon_perf_stats *st =3D this_cpu_ptr(&damon_perf_stats); + + if (occupancy > READ_ONCE(st->ring_peak)) + WRITE_ONCE(st->ring_peak, occupancy); +} + +/* + * Matching =E2=80=94 process context (kdamond drain loop) + */ + +void damon_perf_observe_match(unsigned long addr, int cpu) +{ + per_cpu_ptr(&damon_perf_stats, cpu)->match++; +} + +void damon_perf_observe_miss(unsigned long addr, int cpu, int reason) +{ + struct damon_perf_stats *st =3D per_cpu_ptr(&damon_perf_stats, cpu); + + switch (reason) { + case DAMON_REPORT_MISS_TGID: + st->miss_tgid++; + break; + case DAMON_REPORT_MISS_NOREGION: + st->miss_region++; + break; + case DAMON_REPORT_MISS_BOUNDARY: + st->miss_boundary++; + break; + } + + if (trace_damon_perf_report_missed_enabled()) + trace_damon_perf_report_missed(addr, cpu, reason); +} + +void damon_perf_observe_update(int cpu) +{ + per_cpu_ptr(&damon_perf_stats, cpu)->update++; +} + +void damon_perf_observe_drain(unsigned int total, unsigned int matched) +{ + if (trace_damon_perf_drain_enabled()) + trace_damon_perf_drain(total, matched); +} + +/* + * Stats accessors (for debugfs) + */ + +/* + * Best-effort per-CPU snapshot. Individual u64 writes are atomic on + * 64-bit platforms; no cross-field consistency is guaranteed. This is + * debug data only =E2=80=94 do not build policy on it. + */ +void damon_perf_stats_snapshot(int cpu, struct damon_perf_stats *dst) +{ + struct damon_perf_stats *st =3D per_cpu_ptr(&damon_perf_stats, cpu); + + *dst =3D *st; + dst->cpu_state =3D per_cpu(damon_perf_cpu_state, cpu); +} + +void damon_perf_stats_aggregate(struct damon_perf_stats *dst) +{ + int cpu; + + memset(dst, 0, sizeof(*dst)); + dst->cpu_state =3D DAMON_PERF_STATE_RUNNING; /* start optimistic, take mi= n */ + cpus_read_lock(); + for_each_online_cpu(cpu) { + struct damon_perf_stats *st =3D per_cpu_ptr(&damon_perf_stats, cpu); + + dst->callback +=3D st->callback; + dst->sample_valid +=3D st->sample_valid; + dst->sample_null +=3D st->sample_null; + dst->sample_addr_zero +=3D st->sample_addr_zero; + dst->sample_kernel +=3D st->sample_kernel; + dst->sample_invalid_phys +=3D st->sample_invalid_phys; + dst->enqueue +=3D st->enqueue; + dst->dequeue +=3D st->dequeue; + dst->overflow +=3D st->overflow; + dst->ring_peak =3D max(dst->ring_peak, st->ring_peak); + dst->match +=3D st->match; + dst->miss_tgid +=3D st->miss_tgid; + dst->miss_region +=3D st->miss_region; + dst->miss_boundary +=3D st->miss_boundary; + dst->update +=3D st->update; + /* + * Aggregate per-CPU state: the pipeline hasn't passed a + * stage until ALL CPUs have passed it, so take the min. + */ + dst->cpu_state =3D min_t(int, dst->cpu_state, + per_cpu(damon_perf_cpu_state, cpu)); + } + cpus_read_unlock(); +} + +/* + * Subsystem init + */ + +int damon_perf_framework_init(void) +{ + return 0; +} --=20 2.43.0 From nobody Mon Sep 28 20:12:34 2026 Received: from mail-pj1-f47.google.com (mail-pj1-f47.google.com [209.85.216.47]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5AC183E4510 for ; Tue, 18 Aug 2026 06:11:05 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.47 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787033466; cv=none; b=F+p+2eN/Y81FNlsay34vqPN4iTs6O73yZqnSUScJvja6UqXThLHHhO5hDMfQm/nFdmfi09S6u4Mzj9Fah7EsOqhmKQWfgHMXpEA3x/JFUsjNmpsKSKlVCVRdJhDwI4xPYl9mzW8dl0hg1ljLXGKpG2RCmUoNn5rg8YI790ug4hM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787033466; c=relaxed/simple; bh=mZKFwQauANDJ1gHu+gg8RedjLenlA7e9YJ8y+FF2WoU=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=k9nbFaom3XMUrwmOGG7zqvPgTBLFm4OOqYXABtGCzttJjIzDm28a+C8ZycGL4xcANtK5Z0zsvUpjmmbh83sFA6krXSaDubLfcFF05TFy40EbDQuWhyJXf1BBrv6zP0ZAU1E0mKx0+UeaDJGwDPLAUQ/r8KuPnGQIfsNlb2i40cE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=M/fDfup6; arc=none smtp.client-ip=209.85.216.47 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="M/fDfup6" Received: by mail-pj1-f47.google.com with SMTP id 98e67ed59e1d1-38a0c7e841fso5492049a91.2 for ; Mon, 17 Aug 2026 23:11:05 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787033465; x=1787638265; darn=vger.kernel.org; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:from:to:cc:subject :date:message-id:reply-to:content-type; bh=PFTkBKzl9DDztkv9Ak+eApuO538I5zXX6YMBdyPiC24=; b=M/fDfup6yb5Ul3kWAVU5gN/854XlbjmxUj/vwLrIZ/S0C1NeLvdSG80xWbO6Vz0oAC SjoR9flU7n/TN2phwHYB3lACBPlM3KS3Kwev4+lfpmlez+vZPg1853ZJryEaNh2qLltf A1+pINquqFwjaLEEBDpmUN5/I5Zb95kZKnf4WYurJwed3qc6tgpBLgTl/bTf3UuBqR+m L1MB8cvYYpS9d5uw8sUBX06r2Ky8Bl0lU8YLaoLgnrxo4cNk4M7dj78BTZaIcOqix4Wu bwh2x6a4VZhD/qVfxBLtV1Kms88fiLPgUqmY0tGoF1tql9JFxo4nykvptKvMXs8i6rry BiIg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787033465; x=1787638265; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=PFTkBKzl9DDztkv9Ak+eApuO538I5zXX6YMBdyPiC24=; b=KEtO4LVrO2C17db4wFDNvBsw2Kfq2L33q2vHJOb0JOAegAmY3WM2XhmC/sUViAu7pI Ehi+/lIDNY+HSdpM7+3sGtgNrR2Tf31M5tLPGo/bTIVOjQhLl7ZmZYFNDLuV25NUKA8W OjnAJd1A81QfQITRzRdJL5pYUBwCDPx8pfJTphY3OoEZV6LA2VHa5GiXI9jo1fCMf1f+ e3qA6WE/sfFrN81K0+HbGfNw5gsLTNvsZck2U6lcgW++DbPJOrnkUF4x59RkMiuT2mYe 4aaj6mtS8KbSnNJPIfN4UeYFSscXMZ44z4YLfov0EWVoIEm2LRd/4vAAJJHyZ58WMKU7 3vkg== X-Forwarded-Encrypted: i=1; AHgh+Rpkh3a/SB4Ss3XnhazIhYpZ4f3O+8DDaUYgBa0MxPT/TtmaykZ1pUR9HoeuCTZYmjxbCDXrIE0TUFdWYtI=@vger.kernel.org X-Gm-Message-State: AOJu0YxZteIJMbK1caFU6CPipxmKACqWZNblyByjzQRJArk8yFsL7KHn LtFud/1A+jVW1wAFARtpufXkTm0hhtQGrw29zVDsCeFABCrXSgtEwllK X-Gm-Gg: AR+sD13Dq2OpK0DeMgG9tC+B5G7XkQ1fPH03vz/w59n7F/YsOHk0QHypIy3pARrzAEW WTwF8y7IXT9ZAiuY7BgjKGkzmwjcFi1kG9gHmAcLRi9545Vub30Ro7U54yRdAKBf7JfaW7Ua55m kDaxBrNiwYyM42EOd5wKtXeRnJIymLTHkHyOt/MWTmmAqozOK+phUS6CaFilUKWlbS0SnwzKOhP pjq83N2Jyr6YNG/ZJLLrY0dBmCRwlFNNFBCzmihzDRwMZAdsofYWwXdeqmgsoH6ntorqz5v8YPT VargYnBad2qVkFI76LBuUzLY2wdVUNbpo/Ka91ioTpe4o2IN5bkp035XxRogy04xKmfT1rI6CZl A+wNfD2UJ8Mv1d5znve+O58gqNywf25S4sj9Pz8FhP8oGwJCDsjPblJD13KDRpMfc08+ptsCPT/ VmroM/aNELpVR/ILTRpVYrTXd00BMWPuKhc+cv3rBh4ivx9ubD518FseDNiT/cRWoSI7wPGZm4T bU/lIE= X-Received: by 2002:a17:90a:ec83:b0:381:e74f:8a6a with SMTP id 98e67ed59e1d1-3933b8c1041mr33609337a91.16.1787033464444; Mon, 17 Aug 2026 23:11:04 -0700 (PDT) Received: from kernel.tail6741c6.ts.net ([116.128.244.169]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-3954d38bf94sm4633942a91.9.2026.08.17.23.10.58 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 17 Aug 2026 23:11:04 -0700 (PDT) From: Kunwu Chan X-Google-Original-From: Kunwu Chan To: sj@kernel.org, akpm@linux-foundation.org Cc: damon@lists.linux.dev, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, rostedt@goodmis.org, mhiramat@kernel.org, mathieu.desnoyers@efficios.com, shuah@kernel.org, lianux.mm@gmail.com, Kunwu Chan Subject: [RFC PATCH 3/7] mm/damon/perf: add debugfs statistics interface Date: Tue, 18 Aug 2026 14:10:27 +0800 Message-ID: <20260818061031.827057-4-kunwu.chan@linux.dev> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260818061031.827057-1-kunwu.chan@linux.dev> References: <20260818061031.827057-1-kunwu.chan@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable From: Lian Wang Expose the aggregated per-CPU counters and the global event-stage in a debug-only debugfs perf_stats file. The format is explicitly unstable; tracepoints are the stable diagnostic interface. Co-developed-by: Kunwu Chan Signed-off-by: Kunwu Chan Signed-off-by: Lian Wang --- mm/damon/perf/Makefile | 2 +- mm/damon/perf/debugfs.c | 142 ++++++++++++++++++++++++++++++++++++++++ mm/damon/perf/stats.c | 2 +- 3 files changed, 144 insertions(+), 2 deletions(-) create mode 100644 mm/damon/perf/debugfs.c diff --git a/mm/damon/perf/Makefile b/mm/damon/perf/Makefile index 5c46d3da7ef8..150cbaa875fa 100644 --- a/mm/damon/perf/Makefile +++ b/mm/damon/perf/Makefile @@ -2,4 +2,4 @@ =20 # Observability: per-CPU counters, tracepoints, debugfs perf_stats obj-$(CONFIG_DAMON_PERF_OBSERVE) +=3D damon-perf.o -damon-perf-objs :=3D stats.o +damon-perf-objs :=3D stats.o debugfs.o diff --git a/mm/damon/perf/debugfs.c b/mm/damon/perf/debugfs.c new file mode 100644 index 000000000000..c54dd7644ac3 --- /dev/null +++ b/mm/damon/perf/debugfs.c @@ -0,0 +1,142 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * DAMON Perf Observability =E2=80=94 debugfs Interface + * + * Exposes one file under /sys/kernel/debug/damon/: + * + * perf_stats =E2=80=94 per-CPU pipeline statistics in tabular form + * + * DEBUG ONLY =E2=80=94 format may change without notice; do not parse in + * scripts. For stable diagnostics, use the tracepoints under + * /sys/kernel/debug/tracing/events/damon/ + */ + +#include +#include +#include +#include + +#include "perf.h" + +static struct dentry *damon_debugfs_dir; + +/* + * perf_stats + */ + +static const char *state_name(int s) +{ + switch (s) { + case DAMON_PERF_STATE_UNINIT: return "UNINIT"; + case DAMON_PERF_STATE_CREATED: return "CREATED"; + case DAMON_PERF_STATE_BOUND: return "BOUND"; + case DAMON_PERF_STATE_ENABLED: return "ENABLED"; + case DAMON_PERF_STATE_RUNNING: return "RUNNING"; + case DAMON_PERF_STATE_ERROR: return "ERROR"; + default: return "?"; + } +} + +static int perf_stats_show(struct seq_file *m, void *v) +{ + struct damon_perf_stats agg, st; + int cpu; + bool first =3D true; + + damon_perf_stats_aggregate(&agg); + + seq_puts(m, " -------------- ----------\n"); + seq_puts(m, " Counter Value\n"); + seq_puts(m, " -------------- ----------\n"); + +#define STAT_ROW(label, field) \ + seq_printf(m, " %-12s %8llu\n", label, agg.field) + + STAT_ROW("callback", callback); + STAT_ROW("valid", sample_valid); + STAT_ROW("null", sample_null); + STAT_ROW("addr_zero", sample_addr_zero); + STAT_ROW("kernel", sample_kernel); + STAT_ROW("inv_phys", sample_invalid_phys); + STAT_ROW("enqueue", enqueue); + STAT_ROW("dequeue", dequeue); + STAT_ROW("overflow", overflow); + STAT_ROW("ring_peak", ring_peak); + STAT_ROW("match", match); + STAT_ROW("miss_tgid", miss_tgid); + STAT_ROW("miss_region", miss_region); + STAT_ROW("miss_bound", miss_boundary); + STAT_ROW("update", update); + +#undef STAT_ROW + + seq_puts(m, " -------------- ----------\n\n"); + + /* Per-CPU breakdown */ + cpus_read_lock(); + for_each_online_cpu(cpu) { + damon_perf_stats_snapshot(cpu, &st); + + /* Skip truly idle CPUs */ + if (st.cpu_state =3D=3D DAMON_PERF_STATE_UNINIT && + !st.callback && !st.enqueue && !st.dequeue) + continue; + + if (first) { + seq_puts(m, " Per-CPU (non-zero / non-UNINIT):\n"); + first =3D false; + } + + seq_printf(m, " CPU%02d: st=3D%-7s cb=3D%llu enq=3D%llu deq=3D%llu ovf= =3D%llu match=3D%llu tgid=3D%llu noreg=3D%llu bound=3D%llu upd=3D%llu\n", + cpu, state_name(st.cpu_state), + st.callback, st.enqueue, st.dequeue, + st.overflow, st.match, + st.miss_tgid, st.miss_region, st.miss_boundary, + st.update); + } + cpus_read_unlock(); + + return 0; +} + +static int perf_stats_open(struct inode *inode, struct file *file) +{ + return single_open(file, perf_stats_show, NULL); +} + +static const struct file_operations perf_stats_fops =3D { + .open =3D perf_stats_open, + .read =3D seq_read, + .llseek =3D seq_lseek, + .release =3D single_release, +}; + +/* + * Init / teardown + */ + +int damon_perf_debugfs_init(void) +{ + if (!debugfs_initialized()) + return -ENODEV; + + /* + * Create the "damon" directory first. debugfs_create_dir() mounts + * debugfs via simple_pin_fs() before touching debugfs_mount, so it + * is safe to call during initcall time -- unlike debugfs_lookup(), + * which dereferences debugfs_mount unconditionally and crashes with + * a NULL mount. If the directory already exists (created by another + * DAMON interface) debugfs_create_dir() returns -EEXIST; fall back + * to debugfs_lookup(), which is now safe because the mount exists. + */ + damon_debugfs_dir =3D debugfs_create_dir("damon", NULL); + if (damon_debugfs_dir =3D=3D ERR_PTR(-EEXIST)) + damon_debugfs_dir =3D debugfs_lookup("damon", NULL); + if (IS_ERR(damon_debugfs_dir)) + return PTR_ERR(damon_debugfs_dir); + + debugfs_create_file("perf_stats", 0400, damon_debugfs_dir, + NULL, &perf_stats_fops); + + return 0; +} diff --git a/mm/damon/perf/stats.c b/mm/damon/perf/stats.c index a869f115bb26..ae5b0037a31d 100644 --- a/mm/damon/perf/stats.c +++ b/mm/damon/perf/stats.c @@ -291,5 +291,5 @@ void damon_perf_stats_aggregate(struct damon_perf_stats= *dst) =20 int damon_perf_framework_init(void) { - return 0; + return damon_perf_debugfs_init(); } --=20 2.43.0 From nobody Mon Sep 28 20:12:34 2026 Received: from mail-pj1-f44.google.com (mail-pj1-f44.google.com [209.85.216.44]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EC1F13E4510 for ; Tue, 18 Aug 2026 06:11:11 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.44 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787033473; cv=none; b=epUnSAd9+bYmeggftNznKdzkM39uy2/H1bKsHQApW+MzxVZ/AJYu6qpqSU6Ak0/p0apiHr/KUPb9HcwOZT4FAWW0YzhDwVnHoFklPpgDylnXIvAeifG66bft4bxvC/NMHrJ0RZ/E2aG8Gk+FZ2Kr81UfixYVtgvrzyRRWkIQRBk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787033473; c=relaxed/simple; bh=GCJfR6dFKIlj1FLeCSmVPQc4zronWrksr3AN1wVaOAU=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=NNpioR0povs9pMIrKrn56kJLY44M+Oauugksr+0VuP/EUxHkIfji8CoFKwJvDGJw9clOLi8lNNk56FfMPlstE0jnspzD4oTFMy/jq2n4yaxpJ6qQ8m4p5FHJOM9vYI0mv3lwl/1vdtnHTT/SQBoW0e79w2D/dXAf5Ol4Lo3cJss= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=E/yIXIZP; arc=none smtp.client-ip=209.85.216.44 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="E/yIXIZP" Received: by mail-pj1-f44.google.com with SMTP id 98e67ed59e1d1-38e7109321dso2818020a91.3 for ; Mon, 17 Aug 2026 23:11:11 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787033471; x=1787638271; darn=vger.kernel.org; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:from:to:cc:subject :date:message-id:reply-to:content-type; bh=7MnRAqJMoLRpsIrwycTnpOzFMIb3xRRTSx1mdIHCeG0=; b=E/yIXIZP++4i7rNJXXZOJlg4ZFktD5ROTnLYPnsQAlwuFlPjEzuMzBtN1+SaEIpGnC Ycqh+VqQWLeaT7wWp5cIc+M3fxqUpm6ULuRhF20UAk4QGBeQHJNTMT7x112xeETwOSwF 6QEmZLqxgrRdDivmVxdurHooaf3QeG8Q4K6nYxbTNM69C8HuywgaKcU4pWsc4iRQBRsI IcR2i8vgMuHGVhdZYxpEPoWa7m+yi6FIqbNVt7af7/0X7y6yIP93cykCewM2q98kDCB/ 9VzvRDkNCO39Dh6+xm3GWlWpjcMd05XD+0vDR7QroB+Mup35NFguTvJ/aBQNyKLBUg8f iBCw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787033471; x=1787638271; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=7MnRAqJMoLRpsIrwycTnpOzFMIb3xRRTSx1mdIHCeG0=; b=SXQqGV1itfzkWuJQ33Jl4C8UdJSwBGle0B5CrmM3JwnCR68hHno4iL68eUm+wU/0Fd uoooVr8fTSrdnr8FDEfqEx9CAM7yE1Hvd5MG+PPYjGl883ebhQqUFftq3PXyIRRg6kzn 43DMoEQuVS8THrtHRvpI8LVzaHKEcyuDBGAXyGufYtI5hqf018l6j4BmBeF+V3j/KTo8 izUhE+9zapVJAA+jcfWWxqzQZc/HlgeSXjf3oYRlEHvyNa+g0YX5yMVufuPijsGgv9S0 6RZjNdpVyA+4YZALWZUHVerowApncdhBZzjIZEcAarCMM+RrpFREC5eWZLlWPWsCHHoV pHuQ== X-Forwarded-Encrypted: i=1; AHgh+Rr5TxXcYKhINXhe8nD3PVlxgmY54u+MDuQBGCFo6yQxgTIHZtsK+NNVxdmvqyW5N2xWRnuzTvLdzwIQw1w=@vger.kernel.org X-Gm-Message-State: AOJu0YxMwA3gKNrDU0LK3ymU+k9KkDS6S9AmHMTIS4ca11LS2h2ammag ls3fnupw0BpFH3Akx5hShqj2ghh/4xbYkKUPposHuTadWhzewcC/RLFn X-Gm-Gg: AR+sD12nPUAZUDuT46tModBE7oD8Qqs+07UZHnj7MEo/XuL3p1f/24XFWhSo+lKgJqK pDd7npPf+lhWWjIAcJXOSMg8j0hvgnc70n1+7fdZ2uyOX28B3IpLHWxi+UVBNRBfVRANzd6CcDg wuC8VUqoXl7ZOJvas8+kza19t6iNHznQud5WZBWWu94Wj1Vq9vkrAgegJBptyhOOX9NHRY5Y8fs pqKUhU+b++F176ZhTrLpnxpxzHFv+L0/TVsWyEwWz+DkrPuJQdnmNUXL8EwlVmTDgZCePloK3dm O7vaiRlOHDd9rcIDmo4BOAA6Sav8O/IwMrBejxTwnlaGhcwi0CcKYXUc96HJiN2W82LZS6SZKvx ldGoirrLeeVVhyIyMTDzCU8XZweX9v5p5UHEnvrjN2uT7iYk/PdFB2MKH0cTr6ELHqAhz6tun9N 0rgyH/QKjfHw01Gv//X9zVNgP++QCMNhnErJ61PkAaWSHuI+jzZ6Za4QhoPoNilrthk8q7kypEo w2WytU= X-Received: by 2002:a17:90b:1d45:b0:366:10f1:3d91 with SMTP id 98e67ed59e1d1-3933b7c3b71mr32119576a91.1.1787033470824; Mon, 17 Aug 2026 23:11:10 -0700 (PDT) Received: from kernel.tail6741c6.ts.net ([116.128.244.169]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-3954d38bf94sm4633942a91.9.2026.08.17.23.11.04 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 17 Aug 2026 23:11:10 -0700 (PDT) From: Kunwu Chan X-Google-Original-From: Kunwu Chan To: sj@kernel.org, akpm@linux-foundation.org Cc: damon@lists.linux.dev, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, rostedt@goodmis.org, mhiramat@kernel.org, mathieu.desnoyers@efficios.com, shuah@kernel.org, lianux.mm@gmail.com, Kunwu Chan Subject: [RFC PATCH 4/7] mm/damon: integrate observe API into vaddr overflow handlers and core drain Date: Tue, 18 Aug 2026 14:10:28 +0800 Message-ID: <20260818061031.827057-5-kunwu.chan@linux.dev> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260818061031.827057-1-kunwu.chan@linux.dev> References: <20260818061031.827057-1-kunwu.chan@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable From: Kunwu Chan Wire the observe_*() calls into the DAMON hot paths: vaddr access-check overflow handlers report into the per-CPU ring, and the kdamond drain matches each report against the target whose tgid owns it. The observe calls are pure side-effect statistics (no-ops under CONFIG_DAMON_PERF_OBSERVE=3Dn), so the switch never changes DAMON matching semantics. The vaddr teardown frees the per-event cpu_state array. Read ring->tail once with READ_ONCE in damon_report_access() and reuse the cached value for the peak-occupancy estimate, avoiding a torn read and a compiler reload on the producer side. Co-developed-by: Lian Wang Signed-off-by: Lian Wang Signed-off-by: Kunwu Chan --- mm/damon/core.c | 129 +++++++++++++++++++++++++++++++++++++---------- mm/damon/vaddr.c | 104 +++++++++++++++++++++++++++++++++++--- 2 files changed, 199 insertions(+), 34 deletions(-) diff --git a/mm/damon/core.c b/mm/damon/core.c index 609d627e2b33..377f07122fb0 100644 --- a/mm/damon/core.c +++ b/mm/damon/core.c @@ -21,6 +21,7 @@ =20 /* for damon_get_folio() used by node eligible memory metrics */ #include "ops-common.h" +#include "perf/perf.h" =20 #define CREATE_TRACE_POINTS #include @@ -2243,24 +2244,41 @@ void damon_report_access(struct damon_access_report= *report) preempt_disable(); if (local_inc_return(this_cpu_ptr(&damon_report_ring_busy)) !=3D 1) { /* NMI nested on a process-context producer; drop. */ - trace_damon_perf_ring_overflow(smp_processor_id()); +#ifdef CONFIG_DAMON_PERF_OBSERVE + damon_perf_observe_ring_overflow(smp_processor_id()); +#endif /* CONFIG_DAMON_PERF_OBSERVE */ goto out; } =20 ring =3D this_cpu_ptr(&damon_report_rings); head =3D ring->head; next =3D (head + 1) & DAMON_REPORT_RING_MASK; + { + unsigned int tail =3D READ_ONCE(ring->tail); =20 - if (next =3D=3D READ_ONCE(ring->tail)) { - trace_damon_perf_ring_overflow(smp_processor_id()); - goto out; - } + if (next =3D=3D tail) { +#ifdef CONFIG_DAMON_PERF_OBSERVE + damon_perf_observe_ring_overflow(smp_processor_id()); +#endif /* CONFIG_DAMON_PERF_OBSERVE */ + goto out; + } =20 - ring->entries[head] =3D *report; - ring->entries[head].report_jiffies =3D jiffies; - smp_wmb(); /* publish entry before head advance */ - WRITE_ONCE(ring->head, next); - WRITE_ONCE(*this_cpu_ptr(&damon_ring_pending), 1); + ring->entries[head] =3D *report; + ring->entries[head].report_jiffies =3D jiffies; + smp_wmb(); /* publish entry before head advance */ + WRITE_ONCE(ring->head, next); + WRITE_ONCE(*this_cpu_ptr(&damon_ring_pending), 1); +#ifdef CONFIG_DAMON_PERF_OBSERVE + damon_perf_observe_ring_enqueue(); + /* + * Track peak occupancy for health evaluation. + * next is the new head; tail was read before enqueue + * (may be slightly stale =E2=80=94 acceptable for a peak estimate). + */ + damon_perf_observe_ring_peak( + (next - tail) & DAMON_REPORT_RING_MASK); +#endif /* CONFIG_DAMON_PERF_OBSERVE */ + } out: local_dec(this_cpu_ptr(&damon_report_ring_busy)); preempt_enable(); @@ -2276,6 +2294,9 @@ void damon_report_page_fault(struct vm_fault *vmf, bo= ol huge_pmd) .tid =3D current->pid, .tgid =3D task_tgid_nr(current), .is_write =3D vmf->flags & FAULT_FLAG_WRITE, +#ifdef CONFIG_DAMON_PERF_OBSERVE + .source =3D DAMON_REPORT_SRC_PAGE_FAULT, +#endif /* CONFIG_DAMON_PERF_OBSERVE */ }; =20 if (huge_pmd) @@ -3917,8 +3938,16 @@ static bool damon_sample_filter_out(struct damon_acc= ess_report *report, return !filter->allow; } =20 -static void kdamond_apply_access_report(struct damon_access_report *report, - struct damon_target *t, +/* + * Try to apply one access report to a target's region snapshot. + * + * Caller has already resolved tgid (for pid-based monitoring), so this + * function only does address-to-region matching. Miss reasons for + * trace_damon_perf_report_missed use enum damon_report_miss_reason. + * + * Return: true if the report fell inside a known region, false otherwise. + */ +static bool kdamond_apply_access_report(struct damon_access_report *report, struct damon_region **regions, unsigned int nr_regions, struct damon_ctx *ctx) { @@ -3926,13 +3955,7 @@ static void kdamond_apply_access_report(struct damon= _access_report *report, unsigned long addr; int left, right, mid; =20 - if (damon_target_has_pid(ctx)) { - if (pid_nr(t->pid) !=3D report->tgid) - return; - addr =3D report->vaddr; - } else { - addr =3D report->paddr; - } + addr =3D damon_target_has_pid(ctx) ? report->vaddr : report->paddr; =20 /* Binary search the snapshot for the region containing addr. */ left =3D 0; @@ -3951,17 +3974,27 @@ static void kdamond_apply_access_report(struct damo= n_access_report *report, } } =20 - if (!r) - return; + if (!r) { + damon_perf_observe_miss(addr, report->cpu, + DAMON_REPORT_MISS_NOREGION); + return false; + } /* Reject reports straddling a region boundary. */ - if (addr + report->size > r->ar.end) - return; + if (addr + report->size > r->ar.end) { + damon_perf_observe_miss(addr, report->cpu, + DAMON_REPORT_MISS_BOUNDARY); + return false; + } if (!r->access_reported) { damon_update_region_access_rate(r, true, &ctx->attrs); r->access_reported =3D true; + damon_perf_observe_update(report->cpu); } + damon_perf_observe_match(addr, report->cpu); + return true; } =20 + static unsigned int kdamond_apply_zero_access_report(struct damon_ctx *ctx) { struct damon_target *t; @@ -4045,6 +4078,7 @@ static unsigned int kdamond_check_reported_accesses(s= truct damon_ctx *ctx) struct damon_target_lookup *tbl; unsigned int nr_targets =3D 0; unsigned int i; + unsigned int total_reports =3D 0, matched_reports =3D 0; =20 tbl =3D damon_build_target_lookup(ctx, &nr_targets); if (!tbl) { @@ -4077,6 +4111,10 @@ static unsigned int kdamond_check_reported_accesses(= struct damon_ctx *ctx) while (tail !=3D head) { struct damon_access_report *report =3D &ring->entries[tail]; + bool applied =3D false; + + /* Count every entry removed from the ring */ + damon_perf_observe_ring_dequeue(report->cpu); =20 if (time_before(report->report_jiffies, jiffies - usecs_to_jiffies( @@ -4085,16 +4123,52 @@ static unsigned int kdamond_check_reported_accesses= (struct damon_ctx *ctx) if (damon_sample_filter_out(report, &ctx->sample_control)) goto next; - for (i =3D 0; i < nr_targets; i++) - kdamond_apply_access_report(report, - tbl[i].t, + /* + * For pid-based monitoring, resolve tgid to the + * single matching target before calling + * kdamond_apply_access_report(), avoiding a + * spurious miss tracepoint for every non-matching + * target. + */ + if (damon_target_has_pid(ctx)) { + for (i =3D 0; i < nr_targets; i++) { + if (pid_nr(tbl[i].t->pid) =3D=3D + report->tgid) { + applied =3D + kdamond_apply_access_report( + report, + tbl[i].regions, + tbl[i].nr_regions, + ctx); + break; + } + } + if (!applied && i =3D=3D nr_targets) + damon_perf_observe_miss( + report->vaddr, + report->cpu, + DAMON_REPORT_MISS_TGID); + } else { + for (i =3D 0; i < nr_targets; i++) + applied |=3D + kdamond_apply_access_report( + report, tbl[i].regions, tbl[i].nr_regions, ctx); + } + total_reports++; + if (applied) + matched_reports++; + next: tail =3D (tail + 1) & DAMON_REPORT_RING_MASK; } WRITE_ONCE(ring->tail, tail); } + + if (total_reports) + damon_perf_observe_drain(total_reports, matched_reports); + /* For nr_accesses_bp, absence of access should also be reported. */ return kdamond_apply_zero_access_report(ctx); } @@ -4158,8 +4232,9 @@ static int kdamond_fn(void *data) ctx->passed_sample_intervals++; =20 if (!list_empty(&ctx->perf_events) || - ctx->sample_control.primitives_enabled.page_fault) + ctx->sample_control.primitives_enabled.page_fault) { max_nr_accesses =3D kdamond_check_reported_accesses(ctx); + } else if (ctx->ops.check_accesses) max_nr_accesses =3D ctx->ops.check_accesses(ctx); if (ctx->ops.apply_probes) diff --git a/mm/damon/vaddr.c b/mm/damon/vaddr.c index 73fcea91afa0..a68c7262d533 100644 --- a/mm/damon/vaddr.c +++ b/mm/damon/vaddr.c @@ -17,6 +17,8 @@ #include #include =20 +#include "perf/perf.h" + #include "../internal.h" #include "ops-common.h" =20 @@ -975,13 +977,49 @@ static void damon_perf_overflow_vaddr(struct perf_eve= nt *perf_event, struct perf_sample_data *data, struct pt_regs *regs) { struct damon_access_report report; + u64 data_src_val; + u64 period_val; + + /* + * Observe every hardware sample through the unified API. + * + * reason encodes why a sample was dropped at the handler level: + * 0 =3D valid, queued to ring + * 1 =3D data =3D=3D NULL + * 2 =3D addr =3D=3D 0 (PMU did not populate data->addr) + * 3 =3D kernel address (addr >=3D TASK_SIZE) + */ + if (!data) { + damon_perf_observe_sample(0, 0, 0, + smp_processor_id(), 1, 0, + perf_event->attr.sample_type); + return; + } =20 - if (!data || !data->addr) + data_src_val =3D data->data_src.val; + period_val =3D data->period; + + if (!data->addr) { + damon_perf_observe_sample(0, data_src_val, period_val, + smp_processor_id(), 2, + data->sample_flags, + perf_event->attr.sample_type); return; + } =20 /* Drop kernel-VA hits -- only user-space VAs land in damon vaddr regions= . */ - if (data->addr >=3D TASK_SIZE) + if (data->addr >=3D TASK_SIZE) { + damon_perf_observe_sample(data->addr, data_src_val, period_val, + smp_processor_id(), 3, + data->sample_flags, + perf_event->attr.sample_type); return; + } + + damon_perf_observe_sample(data->addr, data_src_val, period_val, + smp_processor_id(), 0, + data->sample_flags, + perf_event->attr.sample_type); =20 report =3D (struct damon_access_report){ .vaddr =3D data->addr & PAGE_MASK, @@ -990,6 +1028,9 @@ static void damon_perf_overflow_vaddr(struct perf_even= t *perf_event, .tid =3D current->pid, .tgid =3D current->tgid, .is_write =3D !!(data->data_src.mem_op & PERF_MEM_OP_STORE), +#ifdef CONFIG_DAMON_PERF_OBSERVE + .source =3D DAMON_REPORT_SRC_PERF_OVERFLOW, +#endif /* CONFIG_DAMON_PERF_OBSERVE */ }; damon_report_access(&report); } @@ -998,9 +1039,18 @@ static void damon_perf_overflow_paddr(struct perf_eve= nt *perf_event, struct perf_sample_data *data, struct pt_regs *regs) { struct damon_access_report report; + u64 data_src_val; + u64 period_val; =20 - if (!data) + if (!data) { + damon_perf_observe_sample(0, 0, 0, + smp_processor_id(), 1, 0, + perf_event->attr.sample_type); return; + } + + data_src_val =3D data->data_src.val; + period_val =3D data->period; =20 /* * AMD IBS Op only populates data->phys_addr when @@ -1008,14 +1058,27 @@ static void damon_perf_overflow_paddr(struct perf_e= vent *perf_event, * carries a stale value. Gate on sample_flags rather than testing * phys_addr for zero (which would also drop legitimate page 0). */ - if (!(data->sample_flags & PERF_SAMPLE_PHYS_ADDR)) + if (!(data->sample_flags & PERF_SAMPLE_PHYS_ADDR)) { + damon_perf_observe_sample(0, data_src_val, + period_val, smp_processor_id(), 4, + data->sample_flags, + perf_event->attr.sample_type); return; + } + + damon_perf_observe_sample(data->phys_addr, data_src_val, period_val, + smp_processor_id(), 0, + data->sample_flags, + perf_event->attr.sample_type); =20 report =3D (struct damon_access_report){ .paddr =3D data->phys_addr & PAGE_MASK, .size =3D PAGE_SIZE, .cpu =3D smp_processor_id(), .is_write =3D !!(data->data_src.mem_op & PERF_MEM_OP_STORE), +#ifdef CONFIG_DAMON_PERF_OBSERVE + .source =3D DAMON_REPORT_SRC_PERF_OVERFLOW, +#endif /* CONFIG_DAMON_PERF_OBSERVE */ }; damon_report_access(&report); } @@ -1070,6 +1133,8 @@ static int damon_perf_cpu_online(unsigned int cpu, st= ruct hlist_node *node) if (!perf) return 0; =20 + damon_perf_observe_event_created(event, cpu); + damon_perf_event_init_attr(event, &attr); =20 /* @@ -1092,14 +1157,20 @@ static int damon_perf_cpu_online(unsigned int cpu, = struct hlist_node *node) return 0; /* never block CPU online */ } *per_cpu_ptr(perf->event, cpu) =3D perf_event; + + damon_perf_observe_event_bound(event, cpu, perf_event); + /* * Late-online CPU after the substrate is armed: events are created * with attr.disabled =3D 1 and would otherwise stay quiescent on this * CPU until the next arm walk. Enable here so coverage matches the * already-online CPUs. */ - if (event->ctx && READ_ONCE(event->ctx->perf_events_active)) + if (event->ctx && READ_ONCE(event->ctx->perf_events_active)) { perf_event_enable(perf_event); + damon_perf_observe_event_enabled(event, cpu, + perf_event->state, perf_event->oncpu); + } return 0; } =20 @@ -1115,6 +1186,7 @@ static int damon_perf_cpu_offline(unsigned int cpu, s= truct hlist_node *node) =20 perf_event =3D per_cpu(*perf->event, cpu); if (perf_event) { + damon_perf_observe_event_destroyed(event, cpu); perf_event_disable(perf_event); perf_event_release_kernel(perf_event); *per_cpu_ptr(perf->event, cpu) =3D NULL; @@ -1133,8 +1205,12 @@ void damon_perf_event_arm(struct damon_perf_event *e= vent) =20 for_each_online_cpu(cpu) { perf_event =3D *per_cpu_ptr(perf->event, cpu); - if (perf_event) + if (perf_event) { perf_event_enable(perf_event); + damon_perf_observe_event_enabled(event, cpu, + perf_event->state, + perf_event->oncpu); + } } } =20 @@ -1149,8 +1225,11 @@ void damon_perf_event_disarm(struct damon_perf_event= *event) =20 for_each_online_cpu(cpu) { perf_event =3D *per_cpu_ptr(perf->event, cpu); - if (perf_event) + if (perf_event) { perf_event_disable(perf_event); + damon_perf_observe_event_disabled(event, cpu, + perf_event->state); + } } } =20 @@ -1192,6 +1271,7 @@ int damon_perf_init(struct damon_ctx *ctx, struct dam= on_perf_event *event) return 0; =20 free_event: + damon_perf_observe_event_free(event); free_percpu(perf->event); free_perf: kfree(perf); @@ -1203,6 +1283,8 @@ void damon_perf_cleanup(struct damon_ctx *ctx, struct= damon_perf_event *event) { struct damon_perf *perf =3D event->priv; =20 + damon_perf_observe_event_free(event); + if (!perf) return; =20 @@ -1244,6 +1326,14 @@ static int __init damon_va_initcall(void) if (err < 0) return err; damon_perf_cpuhp_state =3D err; + +#ifdef CONFIG_DAMON_PERF_OBSERVE + err =3D damon_perf_framework_init(); + if (err < 0) + pr_warn("damon-perf: framework init failed, observability unavailable: %= d\n", + err); + /* Non-fatal: vaddr/fvaddr ops still register. */ +#endif /* CONFIG_DAMON_PERF_OBSERVE */ #endif =20 err =3D damon_register_ops(&ops); --=20 2.43.0 From nobody Mon Sep 28 20:12:34 2026 Received: from mail-pj1-f47.google.com (mail-pj1-f47.google.com [209.85.216.47]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CE48E3E3176 for ; Tue, 18 Aug 2026 06:11:18 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.47 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787033481; cv=none; b=LqL8981GR/jYVhWKWTGuN2a8VMSZhdu71sImUrlobnamtB+cFYoKSMFePCa213s0m0ZUFZl+5x3PybhBJ8S9N2SpVCr5/lUGvFzSLN3ZKdS24cUSa2MTTuxfHvsf2+xLeuNDUJ7X9Q08/WL9cifJoBEajsBOGkuoTQUJ2WVQx1A= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787033481; c=relaxed/simple; bh=j7Uj2GO9RyEHH/pDmfQU4PHrULq40dlgjm/Tga2f6BY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=iPDWN3u3aSYMdVq2SoeDfI1MiwrJ/NR6cD4RSk3dn5nd+or93wvRGokCYd6tt3YL5TWhfRkEmy3V/XURt5kDYKdcUEiBYV6D6xz8OdRAU/CQkq9+2Hn8JThcTihRiFlpt+8m8CY+On4nmnjh+fZ2yUvZag2uz/V/A70v379lgBM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=ahTbZzYE; arc=none smtp.client-ip=209.85.216.47 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="ahTbZzYE" Received: by mail-pj1-f47.google.com with SMTP id 98e67ed59e1d1-3900e39d935so3825919a91.0 for ; Mon, 17 Aug 2026 23:11:18 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787033478; x=1787638278; darn=vger.kernel.org; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:from:to:cc:subject :date:message-id:reply-to:content-type; bh=stz2R8m9lkuSZFQdtg7AzG6IQOSAxbyACeZhPcbbimg=; b=ahTbZzYEPyDWB4nzzCaNwWcO9pRtZA5rTj1sGYcm4wSXeEX4fwhFlWIMskte7W1uIK NVY88kEDVq5nOieYmo/teplcx7/dAdhr18o3g5Srvt8A4E3A4aSfc2iKwmWcZ8Zz7ICv 2SVMU27Iigb0QQeY+CQ43M/L3gUQiuuHDw/kEQPUd0xhCRyAlnhlQL44hfTSP2hYbbna bZFLvKYVsKG/HZGUMnuiSjbPetCSG+nBJO2HCdlVE5xs4L6CIQJYpe11jvepuvHSyZLw TAL/E3JJyvuKUH5ncO/v2dLofG7/ILLpbt/76K3S8oYryURHvLoS5mU7rUVKp6IUF5AS oL5w== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787033478; x=1787638278; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=stz2R8m9lkuSZFQdtg7AzG6IQOSAxbyACeZhPcbbimg=; b=kz7rCTuL2Wax1dAHG0FflqRW/6ytVZv1SO7B5nARL3TkXcr7sOWA/WAyV3vIkPCNBH 2S7nsCLZgewe0mseV+TS3JATx7AOGUDVQyVjfnrn2yGrIcynXwQdkICcCtsppCaT44Gq xP4xVxHQr5QhGI+ffkLdrCssajnbLY/F5Q60avGCrfqzP1qEArwkMBA1fdsS9VrJQ8K8 RsMlwb1lQUmzfofQ6NCFOzB0fQ6l0Wm/StpXsS83ETE7Piz9MRWNtlg3NWCKrP9t7O0D SbYqkGlg2xpZ+gEyufoyz1lmsRa/gmzew+K1jM2fDJL2CoEsgk+Mve5qGo+0FCjVfUgr 4tYg== X-Forwarded-Encrypted: i=1; AHgh+RrSgpk2DHqCjwfCg8XsJb1dkZKye4Qk96y/52EH/6oHhv3+6zgxjzgYXs3Q/6dQ7OhbCgZyRFsZs7mMtR8=@vger.kernel.org X-Gm-Message-State: AOJu0YxaE0uosCYKDicPcX1V+Uuv2oFrxmje/VDljTi42P5cvKhizZbo Dbm57EsBgC9L1EcR06szEWo8bs6c6PZWD1PebmaEGhDcggGfrQWM3bv/ X-Gm-Gg: AR+sD10yUxxUa/7dCfL6M9Owe7rlSSxA7+cLE2fvkN30KJBsRpeIpWULZxml5KIT1gR AJFBL4404B4V/WTAI7egp9f9CqqpMFyOpizHln+l3Yt6aV8OcGeIWhTFs8Ecm06PipxQlO5V98J 5+w5sdj+KjMjxW6QHb2GBhqfcVdb1t4THWUrhWaZ5Xn1k/sjDZ5Ym53ct5UhwuFsXWKiKX3499r 9ex6RtSl5VbjhXlHqO2nHnD8lKHv2zPzZKcaT3yTFFpJ/4ipjSFPbFv939TJhiLkGN0rju0pBPl lTSbvEUGA8FkKbPnHWHrQdd9qFEN9NG+LYnPlrgDgykMZggrTNb5mQe7mqeA768jephBEpALc4R E7AN8bvxX9RJ3JxiI+V/EYg1+d4akUFwdgg3yo2liPaFAKebNhXRlQL8REYsIn+sJR4SHngBPWk 8E1gshpUAGifJhvwr7VlKWOUMpbH69TwnQh51fHaR7C69UEutpNwxpPXGduQlgDNM6U/9dG7PPz zV7gzo= X-Received: by 2002:a17:90b:58e5:b0:381:a766:efc9 with SMTP id 98e67ed59e1d1-3933caf3a05mr33753321a91.7.1787033477673; Mon, 17 Aug 2026 23:11:17 -0700 (PDT) Received: from kernel.tail6741c6.ts.net ([116.128.244.169]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-3954d38bf94sm4633942a91.9.2026.08.17.23.11.11 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 17 Aug 2026 23:11:17 -0700 (PDT) From: Kunwu Chan X-Google-Original-From: Kunwu Chan To: sj@kernel.org, akpm@linux-foundation.org Cc: damon@lists.linux.dev, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, rostedt@goodmis.org, mhiramat@kernel.org, mathieu.desnoyers@efficios.com, shuah@kernel.org, lianux.mm@gmail.com, Kunwu Chan Subject: [RFC PATCH 5/7] selftests/damon: add automated layer-by-layer observability test Date: Tue, 18 Aug 2026 14:10:29 +0800 Message-ID: <20260818061031.827057-6-kunwu.chan@linux.dev> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260818061031.827057-1-kunwu.chan@linux.dev> References: <20260818061031.827057-1-kunwu.chan@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable From: Kunwu Chan Drive the observe framework end-to-end with a software page-fault PMU positive control (no hardware PMU required): create a kdamond with a perf event via the DAMON sysfs interface, run a memory-pressure workload, capture a bounded trace window, and assert each pipeline layer from per-run snapshot/delta counter deltas (counters are cumulative since boot, so the snapshot is taken before the workload window): callbacks, valid data addresses, ring enqueue/dequeue/ overflow, drain match/update, the four DAMON perf tracepoints, and the per-CPU state column advancing CREATED/BOUND/ENABLED to RUNNING after the first callback. A clean-session guard refuses to run while a kdamond already exists. Ends with a PMU support verdict (FULLY INTEGRATED / PLUMBING-ONLY / UNUSABLE). Cleanup is ownership safe: it tears down only the kdamond, workload and debugfs mount created by this invocation, and retains the raw evidence directory by default. Co-developed-by: Lian Wang Signed-off-by: Lian Wang Signed-off-by: Kunwu Chan --- tools/testing/selftests/damon/Makefile | 1 + .../selftests/damon/damon_perf_obs_test.sh | 562 ++++++++++++++++++ 2 files changed, 563 insertions(+) create mode 100755 tools/testing/selftests/damon/damon_perf_obs_test.sh diff --git a/tools/testing/selftests/damon/Makefile b/tools/testing/selftes= ts/damon/Makefile index 2180c328a825..1db8fa95ba2d 100644 --- a/tools/testing/selftests/damon/Makefile +++ b/tools/testing/selftests/damon/Makefile @@ -23,4 +23,5 @@ TEST_PROGS +=3D sysfs_no_op_commit_break.py =20 EXTRA_CLEAN =3D __pycache__ =20 +TEST_PROGS +=3D damon_perf_obs_test.sh include ../lib.mk diff --git a/tools/testing/selftests/damon/damon_perf_obs_test.sh b/tools/t= esting/selftests/damon/damon_perf_obs_test.sh new file mode 100755 index 000000000000..4c4074cdd191 --- /dev/null +++ b/tools/testing/selftests/damon/damon_perf_obs_test.sh @@ -0,0 +1,562 @@ +#!/bin/bash +# SPDX-License-Identifier: GPL-2.0 +# +# DAMON Perf Observability Framework =E2=80=94 Automated Layer-by-Layer Te= st +# +# Validates all 7 pipeline stages: +# Layer 1: Event Create Layer 2: Event Bind +# Layer 3: Event Enable Layer 4: Sampling (callback) +# Layer 5: Ring Layer 6: Drain +# Layer 7: Match & Update +# +# The framework counters are cumulative since boot, so this script +# snapshots them before the workload and reports per-run deltas. It +# also clears the trace buffer before the sampling window so trace.txt +# carries only records produced by this run. +# +# Usage: +# # Software page-fault positive control (exercises the FULL pipeline, +# # works on any machine, no HW PMU required). Defaults to sampling +# # every page fault (period 1); override with --freq/--period: +# sudo ./damon_perf_obs_test.sh --pmu software +# sudo ./damon_perf_obs_test.sh --pmu software --freq 1 --sample-freq 100 +# +# # With ARM SPE: +# sudo ./damon_perf_obs_test.sh --pmu arm_spe_0 --freq 0 --period 256 +# +# # With any PMU type number: +# sudo ./damon_perf_obs_test.sh --pmu-type 38 --freq 0 --period 256 +# +# Notes: +# - `--config N` sets perf_event_attr.config. For PERF_TYPE_SOFTWARE, +# config 2 (PERF_COUNT_SW_PAGE_FAULTS) is the only software event +# that populates data->addr; cpu-clock (config 0) delivers callbacks +# with addr always 0, so it can only validate Layers 1-4. +# - ARM SPE cannot sample through perf_event_create_kernel_counter() +# (it requires an AUX ring buffer, see arm_spe_pmu.c), so an SPE run +# is expected to report zero callbacks until an AUX backend exists. +# A zero-callback delta is the correct "PMU not usable" verdict. +# +# Requirements: +# - CONFIG_DAMON_PERF_OBSERVE=3Dy (fatal if missing) +# - Root privileges +# - debugfs mounted + +set -e + +# ---- defaults ---- +PMU_NAME=3D"" +PMU_TYPE=3D"" +PMU_CONFIG=3D0 +CONFIG_EXPLICIT=3D0 +FREQ=3D0 +PERIOD=3D256 +FREQ_EXPLICIT=3D0 +PERIOD_EXPLICIT=3D0 +SAMPLE_FREQ=3D100 +TIMEOUT=3D10 +TRACE_WINDOW=3D3 +TARGET_PID=3D"" +RESULTS_DIR=3D"/tmp/damon_perf_test_$$" +PASSED=3D0 +FAILED=3D0 +SKIPPED=3D0 +STRESS_PID=3D"" +CREATED_KDAMOND=3D0 +MOUNTED_DEBUGFS=3D0 +KEEP_RESULTS=3D${KEEP_RESULTS:-1} + +# ---- helpers ---- +pass() { echo " [PASS] $1"; PASSED=3D$((PASSED + 1)); } +fail() { echo " [FAIL] $1 =E2=80=94 $2"; FAILED=3D$((FAILED + 1)); } +skip() { echo " [SKIP] $1 =E2=80=94 $2"; SKIPPED=3D$((SKIPPED + 1)); } +die() { echo "FATAL: $1"; exit 1; } + +# ---- sysfs roots (kept under 100 columns) ---- +KD=3D/sys/kernel/mm/damon/admin/kdamonds +ADMIN=3D$KD/0 +TRACE=3D/sys/kernel/debug/tracing +TPD=3D$TRACE/events/damon +PE=3D$ADMIN/contexts/0/monitoring_attrs/sample/perf_events + +# ---- saved pre-test state (restored in cleanup so the test is +# ---- side-effect free: tracepoints, tracing_on) +ORIG_TRACING_ON=3D$(cat $TRACE/tracing_on 2>/dev/null || echo 0) +ORIG_TP_SAMPLE=3D$(cat $TRACE/events/damon/damon_perf_sample/enable 2>/dev= /null || echo 0) +ORIG_TP_OVERFLOW=3D$(cat $TRACE/events/damon/damon_perf_ring_overflow/enab= le 2>/dev/null || echo 0) +ORIG_TP_MISSED=3D$(cat $TRACE/events/damon/damon_perf_report_missed/enable= 2>/dev/null || echo 0) +ORIG_TP_DRAIN=3D$(cat $TRACE/events/damon/damon_perf_drain/enable 2>/dev/n= ull || echo 0) + +cleanup() { + echo "" + echo "=3D=3D=3D Cleaning up =3D=3D=3D" + if [[ -n "$STRESS_PID" ]]; then + kill "$STRESS_PID" 2>/dev/null || true + wait "$STRESS_PID" 2>/dev/null || true + STRESS_PID=3D"" + fi + # Tear down only the kdamond instance created by this test. In + # particular, the early "existing kdamonds" guard must be read-only. + if [[ "$CREATED_KDAMOND" =3D=3D "1" ]]; then + echo off > $ADMIN/state 2>/dev/null || true + echo 0 > $KD/nr_kdamonds 2>/dev/null || true + CREATED_KDAMOND=3D0 + fi + # Restore tracepoint and tracing state + echo 0 > $TRACE/tracing_on 2>/dev/null || true + echo "$ORIG_TP_SAMPLE" > $TRACE/events/damon/damon_perf_sample/enable 2>/= dev/null || true + echo "$ORIG_TP_OVERFLOW" > $TPD/damon_perf_ring_overflow/enable 2>/dev/nu= ll || true + echo "$ORIG_TP_MISSED" > $TPD/damon_perf_report_missed/enable 2>/dev/null= || true + echo "$ORIG_TP_DRAIN" > $TRACE/events/damon/damon_perf_drain/enable 2>/de= v/null || true + echo "$ORIG_TRACING_ON" > $TRACE/tracing_on 2>/dev/null || true + if [[ "$MOUNTED_DEBUGFS" =3D=3D "1" ]]; then + umount /sys/kernel/debug 2>/dev/null || true + MOUNTED_DEBUGFS=3D0 + fi + [ "$KEEP_RESULTS" !=3D "1" ] && rm -rf "$RESULTS_DIR" +} +trap cleanup EXIT + +# ---- argument parsing ---- +while [[ $# -gt 0 ]]; do + case "$1" in + --pmu) PMU_NAME=3D"$2"; shift 2 ;; + --pmu-type) PMU_TYPE=3D"$2"; shift 2 ;; + --config) PMU_CONFIG=3D"$2"; CONFIG_EXPLICIT=3D1; shift 2 ;; + --freq) FREQ=3D"$2"; FREQ_EXPLICIT=3D1; shift 2 ;; + --period) PERIOD=3D"$2"; PERIOD_EXPLICIT=3D1; shift 2 ;; + --sample-freq) SAMPLE_FREQ=3D"$2"; shift 2 ;; + --trace-window) TRACE_WINDOW=3D"$2"; shift 2 ;; + --timeout) TIMEOUT=3D"$2"; shift 2 ;; + --pid) TARGET_PID=3D"$2"; shift 2 ;; + *) echo "Unknown: $1"; exit 1 ;; + esac +done + +# Resolve PMU type +if [[ -n "$PMU_NAME" && -z "$PMU_TYPE" ]]; then + if [[ "$PMU_NAME" =3D=3D "software" ]]; then + PMU_TYPE=3D1 + else + PMU_TYPE=3D$(cat /sys/bus/event_source/devices/$PMU_NAME/type 2>/dev/nul= l) || + die "Cannot find PMU: $PMU_NAME" + fi +fi +[[ -z "$PMU_TYPE" ]] && die "Specify --pmu or --pmu-type " + +# For PERF_TYPE_SOFTWARE default to PERF_COUNT_SW_PAGE_FAULTS (config 2): +# the only software event that carries a data address, i.e. the only one +# that can exercise Layers 5-7. Override with --config 0 for a pure +# plumbing (cpu-clock) smoke test. +if [[ "$PMU_TYPE" =3D=3D "1" && "$CONFIG_EXPLICIT" =3D=3D "0" ]]; then + PMU_CONFIG=3D2 +fi + +# For the page-fault positive control, sample every fault (period 1) by +# default so enough reports flow for the ring/drain/match checks to be +# meaningful on a short run. Explicit --freq/--period override this. +if [[ "$PMU_TYPE" =3D=3D "1" && "$PMU_CONFIG" =3D=3D "2" && + "$FREQ_EXPLICIT" =3D=3D "0" && "$PERIOD_EXPLICIT" =3D=3D "0" ]]; then + FREQ=3D0 + PERIOD=3D1 +fi + +# For cpu-clock (config 0), sample_period is a TIME in ns, so the +# default period 256 would mean 4 MHz of callbacks per CPU. Never +# let an unguarded default hit that: fall back to a gentle 100 Hz. +if [[ "$PMU_TYPE" =3D=3D "1" && "$PMU_CONFIG" =3D=3D "0" && + "$FREQ_EXPLICIT" =3D=3D "0" && "$PERIOD_EXPLICIT" =3D=3D "0" ]]; then + FREQ=3D1 + SAMPLE_FREQ=3D100 +fi + +# ---- Layer 0: Environment ---- +echo "=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D" +echo " DAMON Perf Observability =E2=80=94 Layer-by-Layer Test" +echo "=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D" +echo "PMU type: $PMU_TYPE config: $PMU_CONFIG freq: $FREQ period: $PERI= OD timeout: ${TIMEOUT}s" +if [[ -n "$TARGET_PID" ]]; then + echo "Target PID: $TARGET_PID (explicit)" +else + echo "Target PID: workload process (started below)" +fi +echo "" + +mkdir -p "$RESULTS_DIR" + +echo "--- Layer 0: Environment ---" + +# Check kernel config +CONFIG=3D"" +if [[ -f /proc/config.gz ]]; then + CONFIG=3D$(zcat /proc/config.gz) +elif [[ -f /boot/config-$(uname -r) ]]; then + CONFIG=3D$(cat /boot/config-$(uname -r)) +else + die "Cannot read /proc/config.gz or /boot/config-$(uname -r)" +fi + +for opt in DAMON DAMON_SYSFS DAMON_VADDR PERF_EVENTS DEBUG_FS TRACING \ + TRACEPOINTS; do + if echo "$CONFIG" | grep -q "CONFIG_${opt}=3Dy"; then + pass "CONFIG_${opt}=3Dy" + else + fail "CONFIG_${opt}" "not enabled" + fi +done + +# DAMON_PERF_OBSERVE is fatal =E2=80=94 the test cannot run without it +if echo "$CONFIG" | grep -q "CONFIG_DAMON_PERF_OBSERVE=3Dy"; then + pass "CONFIG_DAMON_PERF_OBSERVE=3Dy" +else + die "kernel not built with CONFIG_DAMON_PERF_OBSERVE=3Dy" +fi + +# Check root +[[ $(id -u) -eq 0 ]] || die "Must run as root" + +# Mount debugfs only when this test owns the mount, and undo it on exit. +if ! mountpoint -q /sys/kernel/debug; then + mount -t debugfs none /sys/kernel/debug || die "Cannot mount debugfs" + MOUNTED_DEBUGFS=3D1 +fi +[[ -d /sys/kernel/debug/damon ]] || die "debugfs damon/ not found" +pass "debugfs mounted" + +# Check tracepoints +for tp in damon_perf_sample damon_perf_ring_overflow damon_perf_report_mis= sed damon_perf_drain; do + if [[ -d /sys/kernel/debug/tracing/events/damon/$tp ]]; then + pass "tracepoint $tp exists" + else + fail "tracepoint $tp" "not found" + fi +done + +# Check debugfs file (perf_stats only; format is debug-only, not an ABI) +if [[ -f /sys/kernel/debug/damon/perf_stats ]]; then + pass "debugfs perf_stats exists" +else + fail "debugfs perf_stats" "not found" +fi + +# ---- Guard: refuse to run if kdamonds already exist ---- +NR_KDAMONDS=3D$(cat $KD/nr_kdamonds 2>/dev/null || echo 0) +if [[ "$NR_KDAMONDS" -gt 0 ]]; then + skip "runtime" "kdamonds exist (nr_kdamonds=3D$NR_KDAMONDS) =E2=80=94 ref= using" + KEEP_RESULTS=3D1 + exit 0 +fi + +# ---- Generate memory pressure workload first: the default DAMON +# ---- target must be the workload process itself, so the workload +# ---- must be running before the target PID is written. +echo "" +echo "Starting memory workload for ${TIMEOUT}s..." +if command -v stress-ng &>/dev/null; then + stress-ng --vm 2 --vm-bytes 256M --timeout "${TIMEOUT}s" & + STRESS_PID=3D$! +elif command -v stress &>/dev/null; then + stress --vm 2 --vm-bytes 256M --timeout "${TIMEOUT}s" & + STRESS_PID=3D$! +else + # Fallback: dd-based memory pressure + dd if=3D/dev/zero of=3D/dev/null bs=3D1M count=3D1024 & + STRESS_PID=3D$! +fi + +# Resolve target PID: explicit --pid wins, otherwise the workload +if [[ -z "$TARGET_PID" ]]; then + TARGET_PID=3D$STRESS_PID +fi + +echo "" +echo "--- Layer 1: Event Create ---" + +# Configure kdamond +echo 1 > /sys/kernel/mm/damon/admin/kdamonds/nr_kdamonds +CREATED_KDAMOND=3D1 +echo 1 > /sys/kernel/mm/damon/admin/kdamonds/0/contexts/nr_contexts +echo 1 > /sys/kernel/mm/damon/admin/kdamonds/0/contexts/0/targets/nr_targe= ts +echo "$TARGET_PID" > /sys/kernel/mm/damon/admin/kdamonds/0/contexts/0/targ= ets/0/pid_target + +# Disable page-fault-based access check +echo 0 > $PE/nr_perf_events 2>/dev/null || true + +# Configure perf event +echo 1 > $PE/nr_perf_events +echo "$PMU_TYPE" > $PE/0/type +echo "$PMU_CONFIG" > $PE/0/config + +if [[ "$FREQ" -eq 1 ]]; then + echo "$SAMPLE_FREQ" > $PE/0/sample_freq +else + echo "$PERIOD" > $PE/0/sample_period +fi +echo "$FREQ" > $PE/0/freq + +# Take a dmesg snapshot before enabling the kdamond, so we can +# detect error messages that appear during the test. +DMESG_BEFORE=3D"$RESULTS_DIR/dmesg_before.txt" +dmesg > "$DMESG_BEFORE" 2>/dev/null || true +# Snapshot counters BEFORE the workload window. All counters are +# cumulative since boot; every Analysis number below is a delta against +# this snapshot. Read before echo on so the window starts at zero. +STATS_BASE=3D"$RESULTS_DIR/perf_stats_base.txt" +cat /sys/kernel/debug/damon/perf_stats > "$STATS_BASE" 2>/dev/null || true + +echo on > /sys/kernel/mm/damon/admin/kdamonds/0/state +sleep 2 + +DMESG_OUT=3D"$RESULTS_DIR/dmesg_create.txt" +DMESG_DELTA=3D"$RESULTS_DIR/dmesg_delta.txt" +dmesg > "$DMESG_OUT" 2>/dev/null || true +# Delta: lines in $DMESG_OUT not already present in $DMESG_BEFORE +awk 'NR=3D=3DFNR { seen[$0]++ } + NR>FNR { if (seen[$0] > 0) seen[$0]--; else print }' \ + "$DMESG_BEFORE" "$DMESG_OUT" > "$DMESG_DELTA" || true + +# Check state is on +STATE_VAL=3D$(cat /sys/kernel/mm/damon/admin/kdamonds/0/state 2>/dev/null) +if [[ "$STATE_VAL" =3D=3D "on" ]]; then + pass "Kdamond state is on" +else + fail "Kdamond state" "expected 'on', got '$STATE_VAL'" +fi + +echo "" +echo "--- Layer 2-3: Enable & Run (via per-CPU state) ---" + +# Parse the maximum per-CPU state from the debugfs output. +# The per-CPU line format is: +# CPU%02d: st=3D cb=3D... enq=3D... ... +max_cpu_state() { + awk -F'st=3D' '/^ CPU/ { + split($2, a, " ") + s =3D a[1] + # Map state name to numeric rank + if (s =3D=3D "ERROR") v =3D 5 + else if (s =3D=3D "RUNNING") v =3D 4 + else if (s =3D=3D "ENABLED") v =3D 3 + else if (s =3D=3D "BOUND") v =3D 2 + else if (s =3D=3D "CREATED") v =3D 1 + else v =3D 0 + if (v > max) max =3D v + } END { print max+0 }' "$1" 2>/dev/null +} + +CPU_ST_BASE=3D$(max_cpu_state "$STATS_BASE") + +if [[ "$CPU_ST_BASE" -ge 1 ]]; then + pass "Event Created (max per-CPU state >=3D CREATED)" +else + fail "Event Created" "max per-CPU state is $CPU_ST_BASE" +fi + +if [[ "$CPU_ST_BASE" -ge 2 ]]; then + pass "Event Bound (max per-CPU state >=3D BOUND)" +else + fail "Event Bound" "max per-CPU state is $CPU_ST_BASE" +fi + +if [[ "$CPU_ST_BASE" -ge 3 ]]; then + pass "Event Enabled (max per-CPU state >=3D ENABLED)" +else + fail "Event Enabled" "max per-CPU state is $CPU_ST_BASE" +fi + +# Check dmesg delta for errors (vaddr.c pr_warn_ratelimited paths) +if grep -q 'damon-perf.*failed\|event create failed' "$DMESG_DELTA" 2>/dev= /null; then + fail "Perf event" "perf event creation failed (see dmesg delta)" +else + pass "Perf event creation (no errors in dmesg delta)" +fi + +echo "" +echo "--- Layer 4: Sampling (Callback) ---" + +# Clear the trace buffer so trace.txt only contains records produced by +# this run, then enable tracepoints for a bounded sampling window. +echo 0 > $TRACE/tracing_on 2>/dev/null || true +echo > $TRACE/trace 2>/dev/null || true +echo 1 > $TRACE/events/damon/damon_perf_sample/enable +echo 1 > $TRACE/events/damon/damon_perf_ring_overflow/enable +echo 1 > $TRACE/events/damon/damon_perf_report_missed/enable +echo 1 > $TRACE/events/damon/damon_perf_drain/enable +echo 1 > $TRACE/tracing_on + +# Sampling window: keeps trace.txt bounded even on PMUs that sample at +# tens of kHz. Counters keep accumulating for the full TIMEOUT. +sleep "$TRACE_WINDOW" +echo 0 > $TRACE/tracing_on 2>/dev/null || true +[[ "$TIMEOUT" -gt "$TRACE_WINDOW" ]] && sleep $((TIMEOUT - TRACE_WINDOW)) + +kill $STRESS_PID 2>/dev/null || true +wait $STRESS_PID 2>/dev/null || true +STRESS_PID=3D"" +echo "Workload done." + +sleep 2 # let kdamond drain + +# Collect trace +TRACE_OUT=3D"$RESULTS_DIR/trace.txt" +cat /sys/kernel/debug/tracing/trace > "$TRACE_OUT" 2>/dev/null || true +TRACE_LINES=3D$(wc -l < "$TRACE_OUT" 2>/dev/null || echo 0) + +# Collect stats (re-read after workload) +STATS_OUT=3D"$RESULTS_DIR/perf_stats.txt" +cat /sys/kernel/debug/damon/perf_stats > "$STATS_OUT" 2>/dev/null || true + +# ---- Analysis ---- +echo "" +echo "--- Analysis ---" + + +# Aggregate counter reads +stat_of() { + awk -v k=3D"$1" '$1=3D=3Dk {print $2}' "$2" 2>/dev/null +} + +# Per-run delta between baseline and post-workload snapshots +delta() { + local b a + b=3D$(stat_of "$1" "$STATS_BASE") + a=3D$(stat_of "$1" "$STATS_OUT") + [[ -z "$b" ]] && b=3D0 + [[ -z "$a" ]] && a=3D0 + echo $((a-b)) +} + +CALLBACK=3D$(delta callback) +VALID=3D$(delta valid) +ADDR_ZERO=3D$(delta addr_zero) +KERNEL=3D$(delta kernel) +ENQUEUE=3D$(delta enqueue) +DEQUEUE=3D$(delta dequeue) +OVERFLOW=3D$(delta overflow) +MATCH=3D$(delta match) +UPDATE=3D$(delta update) + +echo " Callback delta this run: ${CALLBACK} (cumulative totals in perf_st= ats.txt)" + +if [[ "$CALLBACK" -gt 0 ]]; then + pass "Sampling: ${CALLBACK} callbacks received" +else + fail "Sampling" "0 callbacks =E2=80=94 PMU is not delivering samples to D= AMON" +fi + +echo " Callback breakdown (delta): valid=3D${VALID} addr_zero=3D${ADDR_ZE= RO} kernel=3D${KERNEL}" + +# Verify RUNNING state: first callback advances per-CPU state to +# RUNNING, which persists until the kdamond is stopped. +CPU_ST_FINAL=3D$(max_cpu_state "$STATS_OUT") +if [[ "$CALLBACK" -gt 0 && "$CPU_ST_FINAL" -ge 4 ]]; then + pass "Event Running (max per-CPU state >=3D RUNNING)" +elif [[ "$CALLBACK" -eq 0 ]]; then + skip "Event Running" "no callbacks =E2=80=94 state cannot advance past EN= ABLED" +else + fail "Event Running" "callbacks > 0 but max state is $CPU_ST_FINAL" +fi + +echo " Ring: enqueue=3D${ENQUEUE} dequeue=3D${DEQUEUE} overflow=3D${OVERF= LOW}" + +if [[ "$ENQUEUE" -gt 0 ]]; then + pass "Ring: enqueue > 0" + if [[ "$DEQUEUE" -gt 0 ]]; then + pass "Ring: dequeue > 0" + else + fail "Ring: dequeue" "enqueued but never dequeued" + fi +else + skip "Ring" "no enqueues (no valid samples: addr=3D0 or no callbacks)" +fi + +echo " Match: match=3D${MATCH} update=3D${UPDATE}" + +if [[ "$MATCH" -gt 0 ]]; then + pass "Drain & Match: ${MATCH} matched" + if [[ "$UPDATE" -gt 0 ]]; then + pass "Update: ${UPDATE} region updates" + else + fail "Update" "matched but never updated" + fi +else + skip "Match/Update" "no matches (no valid samples reached region matching= )" +fi + +# Drain tracepoint: kdamond fires damon_perf_drain whenever it drained +# at least one report. The drain tracepoint is only expected when the +# ring actually produced entries. +DRAIN_COUNT=3D$(grep -c "damon_perf_drain" "$TRACE_OUT" 2>/dev/null || tru= e) +DRAIN_COUNT=3D${DRAIN_COUNT:-0} +echo " Drain tracepoint: ${DRAIN_COUNT} records" +if [[ "$MATCH" -gt 0 ]]; then + if [[ "$DRAIN_COUNT" -gt 0 ]]; then + pass "Drain tracepoint fired (${DRAIN_COUNT} records)" + else + fail "Drain tracepoint" "matches occurred but damon_perf_drain never fir= ed" + fi +else + skip "Drain tracepoint" "no drained reports to summarize" +fi + +# Context verification +if grep -q 'context=3D' "$TRACE_OUT" 2>/dev/null; then + NMI_COUNT=3D$(grep -c 'context=3D3' "$TRACE_OUT" 2>/dev/null || true) + PROC_COUNT=3D$(grep -c 'context=3D0' "$TRACE_OUT" 2>/dev/null || true) + NMI_COUNT=3D${NMI_COUNT:-0} + PROC_COUNT=3D${PROC_COUNT:-0} + echo " Context: NMI=3D${NMI_COUNT} process=3D${PROC_COUNT}" + pass "Context field in trace output" +else + skip "Context" "no trace output to analyze" +fi + +# Show a few sample records as raw evidence (they are the per-sample +# view of the pipeline; useful for PMU support evaluation). +echo "" +echo " First damon_perf_sample record(s) this run:" +if grep -q "damon_perf_sample" "$TRACE_OUT" 2>/dev/null; then + grep -m 3 "damon_perf_sample" "$TRACE_OUT" | sed 's/^/ /' +else + echo " (none =E2=80=94 no samples captured in the ${TRACE_WINDOW}s win= dow)" +fi +echo " First damon_perf_drain record(s) this run:" +if grep -q "damon_perf_drain" "$TRACE_OUT" 2>/dev/null; then + grep -m 3 "damon_perf_drain" "$TRACE_OUT" | sed 's/^/ /' +else + echo " (none)" +fi + +# ---- PMU support verdict ---- +echo "" +echo "--- PMU support verdict ---" +if [[ "$CALLBACK" -eq 0 ]]; then + echo " UNUSABLE: the PMU never delivered a sample to DAMON" + echo " (e.g. ARM SPE requires an AUX ring buffer that kernel" + echo " counters do not provide; see arm_spe_pmu.c)" +elif [[ "$VALID" -eq 0 ]]; then + echo " PLUMBING-ONLY: callbacks flow but no data addresses" + echo " (address-less PMU, e.g. cpu-clock / task-clock)" +elif [[ "$ENQUEUE" -gt 0 && "$MATCH" -gt 0 ]]; then + echo " FULLY INTEGRATED: samples carry addresses and reach" + echo " DAMON region matching/update" +else + echo " PARTIAL: callbacks with addresses, but the drain/match" + echo " pipeline did not complete (see counters above)" +fi + +# ---- Summary ---- +echo "" +echo "=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D" +echo " SUMMARY: $PASSED passed, $FAILED failed, $SKIPPED skipped" +echo "=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D" +echo "" +echo "Results retained in: $RESULTS_DIR (set KEEP_RESULTS=3D0 to remove)" + +if [[ "$FAILED" -gt 0 ]]; then + echo "Overall: FAIL ($FAILED checks failed)" + exit 1 +else + echo "Overall: PASS" + exit 0 +fi --=20 2.43.0 From nobody Mon Sep 28 20:12:34 2026 Received: from mail-pl1-f171.google.com (mail-pl1-f171.google.com [209.85.214.171]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 140D13E95AB for ; Tue, 18 Aug 2026 06:11:24 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.171 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787033487; cv=none; b=X75DGJxRgShZwClu7No0AlHj4Pqjn+S4EauTunE4SlEps2p3JxRZsR8vRj0he37213iwXz8uNdmC23tcJsxNj1C4WjmW81BeGTT5fcyFPb0wE4BVizvoP+C8VltZyn3TKwqmSlmvf2nzBSR2DZtCwKhRMbB8z0t1+DUoULfZHS8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787033487; c=relaxed/simple; bh=gXkE5z9olmERVRUNaNaZmccPuEwdj1m2bfUVPwANH4I=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=kCO64FY6lwwdRIYuxlPcKkF6ukVnAFBdgku+YhwxeBH5tmE9rLadZWcqGIo8biJeRwKaqxz2jMjEHl8+qcP5B5xIHHr7OXppA6icLr0ONnwsUT8nuZxnhExlyfK/zla5WDwqKTkVLmrRlGyUgN9y72mKAM7kW9i3i4BoIt9neLc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=LXqhVtPE; arc=none smtp.client-ip=209.85.214.171 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="LXqhVtPE" Received: by mail-pl1-f171.google.com with SMTP id d9443c01a7336-2d560775ca2so20652625ad.1 for ; Mon, 17 Aug 2026 23:11:24 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787033484; x=1787638284; darn=vger.kernel.org; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:from:to:cc:subject :date:message-id:reply-to:content-type; bh=OkufMBhFwqowO3sZfc9l3+J3Km3N7KYxt/zmtGdjLcc=; b=LXqhVtPEXSB1hNwTk+lK+OgM2xooLTIhZ1UIlKmpaugG15DUSxdsUvgn5TJeoLcnbu r/I1W0pLCKBTP1qyKpOUawfmeeSJCCXN2+UKmbWU7/PghMFJtFUmHfpUjThiG5GcpctM XBtUU5Yl2dQrIqmqHoGipAkwF4SR/PBbb8O2QwlcMNxK7lIHDL4NgF9lni2SxtS/qJ13 WUX4bgz5Et+mysLZywxPcDphd5qW1kIZyI5n9MIta0H3K+pbP9kRV6p1WDEng2QvSuAi 2bZIBLFfRM4rOOiG6VaFUTU+SADDHF54yFvMYlGlMyiiA2e9NQKos3UrtJP4jOX3C4wE vUeg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787033484; x=1787638284; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=OkufMBhFwqowO3sZfc9l3+J3Km3N7KYxt/zmtGdjLcc=; b=byy3CScCrV+FCSP5e+OUe2HXam/AP0HCN0GWlOGixb8WR9+Lku7XnwFsGv/Feq6gqu vdPWilOBn0UiZuLvh7jVaQSCkVW+yaur0izze5lvPVPvWTa/EpA9E/mtt4zXaFO+3Cwb TrgqwwrfKeRDfmEVUsbdQWFQrWhgz5otM+e9prd0msmCyrVV0w/IcgWoW5xqWnb6O7bt 6RpTZwshSpd4t9DtDk8o9cpVM+TTwzNmZLv/Iwhf4MrVHZ3tj5DMwkv8UrnGGGcZp2Gr w1E1ArSJieoKT368r42Z8nx/HpoDo1Udct543c71q70ysHm2NS7zVqjB4yhliDMn1O7m IxqQ== X-Forwarded-Encrypted: i=1; AHgh+RpbaD2LnOKmaX5dRgTPIiwFXtfmcqRdCw61cXKzO071vAlL7xqDRQRi+opXrj3WGSz9vX33eCmRLAJ81vM=@vger.kernel.org X-Gm-Message-State: AOJu0YxiNgrIG9GogJnQEmkb3tko4edhilEZQu2i73sQ3QNfvpz63KRa HBa+Cr7Y5GZo2yZJpSmVV+9g+usPwgE9dBad5AEIVB3pomFrUuWUJjul X-Gm-Gg: AR+sD12C5Mfo5wBkXSaDX4P3dr++V9A25mvcsIzWdEaiZloCxnSOjc4ByVtXdpdyzqr idP1jDQ/zv5gZnExaJ4z2YmPfKkqirs9pgfpI/8mYHyPcFzWoWQCwN3ABLJfVKVg3/RSljPE0Ty FTZhWBkvtiWfgNioQpniUNqzxgmjRx9VvJclJ4ucOqDYQ3yI5ISUdQglaD9syJeZo4zL2d4rDxi /sggd02AZ61AuXUc6G8+LHQtJRnP97Y6dK9Jx3NEvXqxhuEEHmliTFsX/Qj/D8ccZuNs5GeTyvP JehAh6cTsQzWr6J1icbYecsfYGdZkfj5D5zpaFMNc1AgMftF1xtIH+sI6yjlJWgTTxOhFy/lqHa KeWwN9o902a3FvM4gmqfK1F8sPLNlUS/OoMuz5k3LAVOruScgw1ylC+5N7sVNRFgaZhttin/IMr lBV8TPLFWnpMTzBx7ETsoJHxrsCHRUZW8s5pf2z3MsWz/rXAGs3p2f71qgUBqhnl6Zvb43WS/3J G/q8FGQ7VW/ X-Received: by 2002:a17:90b:2f10:b0:38e:f6eb:2b38 with SMTP id 98e67ed59e1d1-3933e63ba0bmr31731771a91.17.1787033483851; Mon, 17 Aug 2026 23:11:23 -0700 (PDT) Received: from kernel.tail6741c6.ts.net ([116.128.244.169]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-3954d38bf94sm4633942a91.9.2026.08.17.23.11.18 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 17 Aug 2026 23:11:23 -0700 (PDT) From: Kunwu Chan X-Google-Original-From: Kunwu Chan To: sj@kernel.org, akpm@linux-foundation.org Cc: damon@lists.linux.dev, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, rostedt@goodmis.org, mhiramat@kernel.org, mathieu.desnoyers@efficios.com, shuah@kernel.org, lianux.mm@gmail.com, Kunwu Chan Subject: [RFC PATCH 6/7] Docs/mm/damon: document the perf observability framework Date: Tue, 18 Aug 2026 14:10:30 +0800 Message-ID: <20260818061031.827057-7-kunwu.chan@linux.dev> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260818061031.827057-1-kunwu.chan@linux.dev> References: <20260818061031.827057-1-kunwu.chan@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable From: Lian Wang Add documentation for the DAMON perf observability framework, covering the CONFIG_DAMON_PERF_OBSERVE Kconfig option, the debugfs perf_stats interface, the tracepoints, and the per-CPU pipeline counter model. The debugfs format is explicitly marked as unstable and must not be used by scripts. Co-developed-by: Kunwu Chan Signed-off-by: Kunwu Chan Signed-off-by: Lian Wang --- Documentation/admin-guide/mm/damon/index.rst | 1 + .../mm/damon/perf-observability.rst | 210 ++++++++++++++++++ 2 files changed, 211 insertions(+) create mode 100644 Documentation/admin-guide/mm/damon/perf-observability.r= st diff --git a/Documentation/admin-guide/mm/damon/index.rst b/Documentation/a= dmin-guide/mm/damon/index.rst index 3ce3164480c7..623a5c312b69 100644 --- a/Documentation/admin-guide/mm/damon/index.rst +++ b/Documentation/admin-guide/mm/damon/index.rst @@ -15,3 +15,4 @@ access monitoring and access-aware system operations. reclaim lru_sort stat + perf-observability diff --git a/Documentation/admin-guide/mm/damon/perf-observability.rst b/Do= cumentation/admin-guide/mm/damon/perf-observability.rst new file mode 100644 index 000000000000..3aa8185de314 --- /dev/null +++ b/Documentation/admin-guide/mm/damon/perf-observability.rst @@ -0,0 +1,210 @@ +.. SPDX-License-Identifier: GPL-2.0 + +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D +DAMON Perf Event Observability Framework +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +The DAMON perf event observability framework provides per-CPU counters and +tracepoints for hardware-sampled access reports. When DAMON is configured= to +use a hardware PMU (e.g. AMD IBS, Intel PEBS, or ARM SPE) instead of page-= table +walks, this framework exposes raw pipeline diagnostics so that every stage= of +the PMU-to-DAMON pipeline can be inspected. + +Counters are best-effort: individual ``u64`` writes are atomic on 64-bit +platforms, but no cross-field consistency is guaranteed. Do not build +policy on snapshot reads. For stable, structured diagnostics, use the +tracepoints under ``/sys/kernel/debug/tracing/events/damon/``. + +Pipeline Stages +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +:: + + PMU hardware =E2=86=92 overflow_handler / AUX drain + =E2=86=92 damon_report_access() =E2=86=92 per-CPU SPSC ring + =E2=86=92 kdamond drain =E2=86=92 target match =E2=86=92 region update + + Layer 1: Event Create perf_event_create_kernel_counter() + Layer 2: Event Bind per-CPU PMU attachment + Layer 3: Event Enable perf_event_enable() + Layer 4: Sampling callback / AUX record received + Layer 5: Ring SPSC enqueue / dequeue / overflow + Layer 6: Drain kdamond consumes entries from ring + Layer 7: Match & Update region access-rate update + +Each layer has a dedicated counter, and most layers have corresponding +tracepoints. Per-CPU event state (UNINIT =E2=86=92 CREATED =E2=86=92 BOUN= D =E2=86=92 ENABLED =E2=86=92 +RUNNING) is recorded unconditionally and exposed via the debugfs +perf_stats file. + +Overhead Control +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +Two levels of overhead control are provided: + +1. **Compile-time** =E2=80=94 ``CONFIG_DAMON_PERF_OBSERVE`` + When set to ``n``, all observe functions are compiled to static-inline + no-ops. No code is generated and no runtime overhead exists. + +2. **Per-tracepoint on/off** =E2=80=94 standard ftrace ``enable`` files + Individual tracepoints (``damon_perf_sample``, etc.) can be + enabled or disabled independently via + ``/sys/kernel/debug/tracing/events/damon/``. Counter increments are + unconditional (cheap per-CPU ``inc``); tracepoint decisions are + guarded by the ftrace static key and are zero-overhead when disabled. + +When ``CONFIG_DAMON_PERF_OBSERVE=3Dy``, per-CPU counters always increment. +There is no runtime toggle for counters; compile-time is the sole gate. + +Debugfs Interface +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +Mount debugfs:: + + # mount -t debugfs none /sys/kernel/debug + +One file is created under ``/sys/kernel/debug/damon/``: + +perf_stats +---------- + +**DEBUG ONLY =E2=80=94 format may change without notice.** Do not parse in +scripts or tools. For stable diagnostics, use the tracepoints. + +Read-only. Aggregated counter table with all pipeline counters plus +per-CPU breakdown:: + + # cat /sys/kernel/debug/damon/perf_stats + =E2=94=8C=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =AC=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=90 + =E2=94=82 Counter =E2=94=82 Value =E2=94=82 + =E2=94=9C=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =BC=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=A4 + =E2=94=82 callback =E2=94=82 233 =E2=94=82 + =E2=94=82 valid =E2=94=82 0 =E2=94=82 + =E2=94=82 null =E2=94=82 0 =E2=94=82 + =E2=94=82 addr_zero =E2=94=82 233 =E2=94=82 + =E2=94=82 kernel =E2=94=82 0 =E2=94=82 + =E2=94=82 inv_phys =E2=94=82 0 =E2=94=82 + =E2=94=82 enqueue =E2=94=82 0 =E2=94=82 + =E2=94=82 dequeue =E2=94=82 0 =E2=94=82 + =E2=94=82 overflow =E2=94=82 0 =E2=94=82 + =E2=94=82 ring_peak =E2=94=82 0 =E2=94=82 + =E2=94=82 match =E2=94=82 0 =E2=94=82 + =E2=94=82 miss_tgid =E2=94=82 0 =E2=94=82 + =E2=94=82 miss_region =E2=94=82 0 =E2=94=82 + =E2=94=82 miss_bound =E2=94=82 0 =E2=94=82 + =E2=94=82 update =E2=94=82 0 =E2=94=82 + =E2=94=94=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2= =94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94= =B4=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80=E2=94=80= =E2=94=80=E2=94=80=E2=94=98 + + Per-CPU (non-zero / non-UNINIT): + CPU00: st=3DBOUND cb=3D44 enq=3D0 deq=3D0 ovf=3D0 match=3D0 miss= =3D0 upd=3D0 + ... + +The ``st=3D`` column shows the per-CPU event state machine position +(UNINIT, CREATED, BOUND, ENABLED, RUNNING, ERROR), derived from the +lifecycle observe calls. This allows verifying lifecycle progression +without parsing dmesg. + +All counters are monotonic (cumulative since boot); userspace computes +deltas between snapshots. + +Tracepoints +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +Four tracepoints are defined:: + + damon_perf_sample + damon_perf_ring_overflow + damon_perf_report_missed + damon_perf_drain + +Enable via ftrace:: + + # echo 1 > /sys/kernel/debug/tracing/events/damon/damon_perf_sample/en= able + # cat /sys/kernel/debug/tracing/trace_pipe + +Each ``damon_perf_sample`` record includes: + + - ``addr``: the accessed virtual address (0 if the PMU did not populate) + - ``data_src``: PERF_MEM_* encoding (PMU-dependent) + - ``period``: sample period or frequency count + - ``cpu``: CPU that generated the sample + - ``reason``: 0=3Dvalid, 1=3Dnull-data, 2=3Daddr-zero, 3=3Dkernel-addr, = 4=3Dinvalid-phys + - ``sample_flags``: what the PMU actually populated + - ``sample_type``: what DAMON requested + - ``context``: 0=3Dprocess, 1=3Dsoftirq, 2=3Dhardirq, 3=3DNMI + +The ``context`` field is particularly useful for cross-PMU validation. +For example, AMD IBS samples arrive in NMI context (context=3D3), while +ARM SPE data from an AUX backend would arrive in process context (context= =3D0). +A mismatch between the expected and actual context is immediately visible. + +Selftest +=3D=3D=3D=3D=3D=3D=3D=3D + +A comprehensive automated test script is provided:: + + # cd tools/testing/selftests/damon + # sudo ./damon_perf_obs_test.sh --pmu arm_spe_0 --freq 0 --period 256 + +The script performs a layer-by-layer validation: + +1. Checks kernel configuration (CONFIG_DAMON, CONFIG_DAMON_PERF_OBSERVE, e= tc.) +2. Verifies PMU availability, tracepoints, and the debugfs perf_stats file +3. Refuses to run if existing kdamonds are present (side-effect guard) +4. Configures DAMON with the specified PMU via sysfs +5. Runs a memory workload (stress-ng, stress, or dd fallback) +6. Collects dmesg delta, trace output, and perf_stats +7. Verifies per-CPU state progression and counter values + +Example output:: + + --- Layer 0: Environment --- + [PASS] CONFIG_DAMON_PERF_OBSERVE=3Dy + [PASS] debugfs perf_stats exists + + --- Layer 2-3: Enable & Run (via per-CPU state) --- + [PASS] Event Created (max per-CPU state >=3D CREATED) + [PASS] Event Bound (max per-CPU state >=3D BOUND) + [PASS] Event Enabled (max per-CPU state >=3D ENABLED) + + --- Layer 4: Sampling (Callback) --- + [PASS] Sampling: 84532 callbacks received + Callback breakdown: valid=3D82103 addr_zero=3D0 kernel=3D2429 + + --- Layer 5: Ring --- + [PASS] Ring: enqueue > 0 + [PASS] Ring: dequeue > 0 + Ring: enqueue=3D82100 dequeue=3D81987 overflow=3D0 + + --- Layer 6: Drain & Match --- + [PASS] Drain & Match: 81987 matched + [PASS] Update: 81987 region updates + +Additional PMU examples:: + + # Software page-fault event (positive control): + sudo ./damon_perf_obs_test.sh --pmu software --freq 1 --sample-freq 100 + + # Any PMU by type number: + sudo ./damon_perf_obs_test.sh --pmu-type 38 --freq 0 --period 256 + +Kernel Configuration +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +Required for observability:: + + CONFIG_DAMON=3Dy + CONFIG_DAMON_SYSFS=3Dy + CONFIG_DAMON_VADDR=3Dy + CONFIG_PERF_EVENTS=3Dy + CONFIG_DEBUG_FS=3Dy + CONFIG_TRACING=3Dy + CONFIG_TRACEPOINTS=3Dy + +Optional (enables observability framework):: + + CONFIG_DAMON_PERF_OBSERVE=3Dy + +When ``CONFIG_DAMON_PERF_OBSERVE=3Dn``, ``/sys/kernel/debug/damon/perf_sta= ts`` +is not created, tracepoints are not registered, and all observe functions +are compiled to empty static inlines with zero overhead. --=20 2.43.0 From nobody Mon Sep 28 20:12:34 2026 Received: from mail-pj1-f44.google.com (mail-pj1-f44.google.com [209.85.216.44]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B26E53EB80E for ; Tue, 18 Aug 2026 06:11:31 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.44 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787033493; cv=none; b=FgHPL6xK024kS0UXa3FhdtmIg9foQRVzsZbJyYD3RFW+v5EIuVZkEJtVaRfSXtJNyJz4euTWesSLHu6iwm+VQLuZ/0Yxer6gkYOdg1H5V6Jt0KwkWDzKd8n6WWWOjYKWGFzg/JippKYUDMPHV1WjPz77HNCe9DohzyFp0anyb2o= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787033493; c=relaxed/simple; bh=LIvxhaVqnrIODtLHUsR+wAXp0bwosE7wdYeBrcUBUT8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=UnFlplCvsDxm4tTOyi0p+l/cIoSn4GNJv/iBbH9LdLn4/de9aFDCylGeO3Xd4uIcuBmH3pDhJ98IStw5k89inl6gsuPH2VxxzwhTH2wH1C6shcd3Uo+eeb01NwTkHpdNMwDnYSUDZXKr9STE/VnuW8OSMR4I9+RXQia20W6RzVk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=E7SFwnHF; arc=none smtp.client-ip=209.85.216.44 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="E7SFwnHF" Received: by mail-pj1-f44.google.com with SMTP id 98e67ed59e1d1-38f620399a0so3696236a91.2 for ; Mon, 17 Aug 2026 23:11:31 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787033491; x=1787638291; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=qyQvFdJaou+co7R0CJUYpuqHbBnGV73Fl2ebn6kSllM=; b=E7SFwnHFzgJX2zvcZSWeyK7qzXjcd8x90LGSvAFRaREhRji1c3PFG48YqHCqNszH9I J7CxWaVy40vIgfRXpzoCWbtINrlfyJN3OoRnTwijxyGpSjD1KYbSF6Z3acQoTvIKnsvP Q38vl5X74SVL/SqsLxZprEmbVsBxCUkRT8bJrnoXagXBYNke+jTEi5XWJbHn1SQk/rTM g7Tx0vq3LRWmLJpMHpjJmhgjOwbYYtxFiwVhJ2TYOdjLI21r/9CesLA6IP4LvBFvC7+C SOWKNBiAixXmya/0/uasdXjODKZfzMy0aJLipx9shexJjIWF1nexHo6zDOZyCjX+jiTt U2Yw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787033491; x=1787638291; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=qyQvFdJaou+co7R0CJUYpuqHbBnGV73Fl2ebn6kSllM=; b=n+tU7CNotHerHsUXJ4UD1y5PzgxxkU+7r163mOO4vHTQpdc0yQ4AHAv7MxOeyBKwaR 0JZl7nof2g2bXni6fww3Cblee0dYd/fZU6e4Cl206ZE26aSDCYT0dpOd1TpyrAs8esPy +Z39ZEZDTp+25EFIA4+iQPm/2IxMZaTFx31rUGz9O1RE1xkP9T0eCdZ2KYjM5lvWC7+Z q0i1bBs9yVVx9ca0N56aVmafssJfiNMIpm+gKfNUhHt+pyakl44B7C9a0YtEPdQTaJz3 f7zn6ixNFWju00s9WAh0ZUPqaAw5kAWQhdwAS+kBUbh/drnISrFtKI4gb9GqQfZ7z+sZ l/eg== X-Forwarded-Encrypted: i=1; AHgh+RqOXutXgVJ2NsSY1op72/SaAwqYu/3FBTbd9RAWquBpy/dlUylkBaAsReXMlwWq42fHbtcvpi9IGnClW6I=@vger.kernel.org X-Gm-Message-State: AOJu0YzclZZLqt4HBPvL5PfqT0b5kvbaNgivXEkUM2tYci+8/7NeBWkU FjKaX81H3QK75raC1NiNg8lvDxh3KRRpI2pfrkaIJ3FqKPHfKxXuqTxD X-Gm-Gg: AR+sD13PocfXaj3iXFJ2gEG7sAIr+xiZGrZVJtL6ppYD9BcRne2y2ggV7TBg2UFTtsH afgn3u7D1VKE+8klZ6Hzn2YJmhcDFXGZDY+l/1C71SaBmDLTBoZe2F/XUJ9I9PwDAXleqP3vsK1 5qKP3+TTlL7rqvDdkVNGHndmuZnOrJKpyhM4aPyBX2+q1VnRiiT/HfqoQq7Xa3s20Luvv0ZdG/r oZTWdQPC+OiMQZ7cu5+MAWfBaim4fGsvoe4TYY52pkdULHWU7NF5vJ7MuqZLdDsS0N0PuVHLmi3 vTSTi+DS/sgs3ljnGP4tDPsYURE9H48QJ9v/v29CtB3tGlVeWha02uvabiiJ+FAQjpRYNU7qni3 /kqRC3EdBf1pOTDnJaZP2do6CPvzpxUNo9IZDg13SRwv99EQNLFZFKvJCJ3Rg6b1XwdMgNZwIiJ 9sOYsEM8FBznPsXWnaILPg274uQcWex4O23SaN9oCJSpqozRphzwGIYsN8OrkY9DLUFFYtLs9VB vtos1o= X-Received: by 2002:a17:90b:39a6:b0:37c:6130:7a5b with SMTP id 98e67ed59e1d1-3955a7467f8mr6401620a91.8.1787033490697; Mon, 17 Aug 2026 23:11:30 -0700 (PDT) Received: from kernel.tail6741c6.ts.net ([116.128.244.169]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-3954d38bf94sm4633942a91.9.2026.08.17.23.11.24 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 17 Aug 2026 23:11:30 -0700 (PDT) From: Kunwu Chan X-Google-Original-From: Kunwu Chan To: sj@kernel.org, akpm@linux-foundation.org Cc: damon@lists.linux.dev, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, rostedt@goodmis.org, mhiramat@kernel.org, mathieu.desnoyers@efficios.com, shuah@kernel.org, lianux.mm@gmail.com, Kunwu Chan Subject: [RFC PATCH 7/7] mm/damon/perf: add CONFIG_DAMON_PERF_DEBUG and pipeline health check Date: Tue, 18 Aug 2026 14:10:31 +0800 Message-ID: <20260818061031.827057-8-kunwu.chan@linux.dev> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260818061031.827057-1-kunwu.chan@linux.dev> References: <20260818061031.827057-1-kunwu.chan@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Lian Wang Add CONFIG_DAMON_PERF_DEBUG as an optional Kconfig option that enables pr_debug() output for the observability pipeline via the damon_perf_dbg() macro. Default off, zero overhead when disabled. Add a pipeline health check to the observability selftest that diagnoses which stage is broken when callbacks are zero, using the existing per-CPU state and counter deltas. Co-developed-by: Kunwu Chan Signed-off-by: Kunwu Chan Signed-off-by: Lian Wang --- mm/damon/Kconfig | 16 ++++++++++++++++ mm/damon/perf/debugfs.c | 5 ++++- mm/damon/perf/perf.h | 7 +++++++ mm/damon/perf/stats.c | 15 ++++++++++++++- .../selftests/damon/damon_perf_obs_test.sh | 18 +++++++++++++++++- 5 files changed, 58 insertions(+), 3 deletions(-) diff --git a/mm/damon/Kconfig b/mm/damon/Kconfig index 9f811510760f..35ec4d12c5b8 100644 --- a/mm/damon/Kconfig +++ b/mm/damon/Kconfig @@ -148,4 +148,20 @@ config DAMON_PERF_OBSERVE static-inline no-ops with zero runtime overhead. =20 If unsure, say N. + +config DAMON_PERF_DEBUG + bool "DAMON Perf verbose debugging output" + depends on DAMON_PERF_OBSERVE + help + Enable verbose per-event and per-drain pr_debug() output + for the DAMON perf observability pipeline. When enabled, + key lifecycle transitions and sampling events are logged + at KERN_DEBUG level, visible via dynamic_debug or when + DEBUG is defined at compile time. + + This adds dmesg noise and should only be enabled for + development or troubleshooting. + + If unsure, say N. + endmenu diff --git a/mm/damon/perf/debugfs.c b/mm/damon/perf/debugfs.c index c54dd7644ac3..48f3d23c2dbc 100644 --- a/mm/damon/perf/debugfs.c +++ b/mm/damon/perf/debugfs.c @@ -132,11 +132,14 @@ int damon_perf_debugfs_init(void) damon_debugfs_dir =3D debugfs_create_dir("damon", NULL); if (damon_debugfs_dir =3D=3D ERR_PTR(-EEXIST)) damon_debugfs_dir =3D debugfs_lookup("damon", NULL); - if (IS_ERR(damon_debugfs_dir)) + if (IS_ERR(damon_debugfs_dir)) { + damon_perf_dbg("debugfs init failed: %ld\n", PTR_ERR(damon_debugfs_dir)); return PTR_ERR(damon_debugfs_dir); + } =20 debugfs_create_file("perf_stats", 0400, damon_debugfs_dir, NULL, &perf_stats_fops); =20 + damon_perf_dbg("debugfs init ok\n"); return 0; } diff --git a/mm/damon/perf/perf.h b/mm/damon/perf/perf.h index 78e23d436336..908c06f2e3db 100644 --- a/mm/damon/perf/perf.h +++ b/mm/damon/perf/perf.h @@ -18,6 +18,13 @@ struct perf_event; #include =20 +#ifdef CONFIG_DAMON_PERF_DEBUG +#define damon_perf_dbg(fmt, ...) \ + pr_debug("damon-perf: " fmt, ##__VA_ARGS__) +#else +#define damon_perf_dbg(fmt, ...) no_printk(fmt, ##__VA_ARGS__) +#endif + struct damon_perf_event; =20 /* diff --git a/mm/damon/perf/stats.c b/mm/damon/perf/stats.c index ae5b0037a31d..e2c1e5764d41 100644 --- a/mm/damon/perf/stats.c +++ b/mm/damon/perf/stats.c @@ -67,6 +67,7 @@ void damon_perf_observe_event_created(struct damon_perf_e= vent *event, int cpu) return; } *per_cpu_ptr(event->cpu_state, cpu) =3D DAMON_PERF_STATE_CREATED; + damon_perf_dbg("cpu %d: event created\n", cpu); damon_perf_cpu_state_advance(cpu, DAMON_PERF_STATE_CREATED); } =20 @@ -76,6 +77,7 @@ void damon_perf_observe_event_bound(struct damon_perf_eve= nt *event, if (event->cpu_state) *per_cpu_ptr(event->cpu_state, cpu) =3D DAMON_PERF_STATE_BOUND; damon_perf_cpu_state_advance(cpu, DAMON_PERF_STATE_BOUND); + damon_perf_dbg("cpu %d: event bound\n", cpu); } =20 void damon_perf_observe_event_enabled(struct damon_perf_event *event, @@ -84,12 +86,14 @@ void damon_perf_observe_event_enabled(struct damon_perf= _event *event, if (event->cpu_state) *per_cpu_ptr(event->cpu_state, cpu) =3D DAMON_PERF_STATE_ENABLED; damon_perf_cpu_state_advance(cpu, DAMON_PERF_STATE_ENABLED); + damon_perf_dbg("cpu %d: event enabled\n", cpu); } =20 void damon_perf_observe_event_disabled(struct damon_perf_event *event, int cpu, int state) { /* State unchanged: the event may be re-enabled later. */ + damon_perf_dbg("cpu %d: event disabled\n", cpu); } =20 void damon_perf_observe_event_destroyed(struct damon_perf_event *event, in= t cpu) @@ -104,6 +108,7 @@ void damon_perf_observe_event_destroyed(struct damon_pe= rf_event *event, int cpu) */ if (event->cpu_state) *per_cpu_ptr(event->cpu_state, cpu) =3D DAMON_PERF_STATE_UNINIT; + damon_perf_dbg("cpu %d: event destroyed\n", cpu); } =20 void damon_perf_observe_event_free(struct damon_perf_event *event) @@ -111,6 +116,7 @@ void damon_perf_observe_event_free(struct damon_perf_ev= ent *event) if (event->cpu_state) { free_percpu(event->cpu_state); event->cpu_state =3D NULL; + damon_perf_dbg("event freed\n"); } } =20 @@ -231,6 +237,7 @@ void damon_perf_observe_drain(unsigned int total, unsig= ned int matched) { if (trace_damon_perf_drain_enabled()) trace_damon_perf_drain(total, matched); + damon_perf_dbg("drain: total=3D%u matched=3D%u\n", total, matched); } =20 /* @@ -291,5 +298,11 @@ void damon_perf_stats_aggregate(struct damon_perf_stat= s *dst) =20 int damon_perf_framework_init(void) { - return damon_perf_debugfs_init(); + int ret =3D damon_perf_debugfs_init(); + + if (ret) + damon_perf_dbg("framework init failed: %d\n", ret); + else + damon_perf_dbg("framework init ok\n"); + return ret; } diff --git a/tools/testing/selftests/damon/damon_perf_obs_test.sh b/tools/t= esting/selftests/damon/damon_perf_obs_test.sh index 4c4074cdd191..cd567c151ae6 100755 --- a/tools/testing/selftests/damon/damon_perf_obs_test.sh +++ b/tools/testing/selftests/damon/damon_perf_obs_test.sh @@ -344,7 +344,7 @@ max_cpu_state() { } END { print max+0 }' "$1" 2>/dev/null } =20 -CPU_ST_BASE=3D$(max_cpu_state "$STATS_BASE") +CPU_ST_BASE=3D$(max_cpu_state /sys/kernel/debug/damon/perf_stats) =20 if [[ "$CPU_ST_BASE" -ge 1 ]]; then pass "Event Created (max per-CPU state >=3D CREATED)" @@ -431,6 +431,22 @@ VALID=3D$(delta valid) ADDR_ZERO=3D$(delta addr_zero) KERNEL=3D$(delta kernel) ENQUEUE=3D$(delta enqueue) +# Pipeline health check: diagnose which stage is broken when +# callbacks are zero, using the existing per-CPU state and +# counter deltas. This is a best-effort diagnostic, not a +# substitute for detailed per-backend debugging. +if [[ "$CALLBACK" -eq 0 ]]; then + CPU_ST_BASE_VAL=3D$(max_cpu_state /sys/kernel/debug/damon/perf_stats) + if [[ "$CPU_ST_BASE_VAL" -le 1 ]]; then + echo " Pipeline diagnosis: event not created or bound (state=3D$CPU_ST_= BASE_VAL)" + elif [[ "$CPU_ST_BASE_VAL" -eq 2 ]]; then + echo " Pipeline diagnosis: event bound but not enabled (state=3DBOUND)" + elif [[ "$ENQUEUE" -eq 0 ]]; then + echo " Pipeline diagnosis: PMU not producing data or AUX pipeline broke= n" + else + echo " Pipeline diagnosis: samples enqueued but none valid" + fi +fi DEQUEUE=3D$(delta dequeue) OVERFLOW=3D$(delta overflow) MATCH=3D$(delta match) --=20 2.43.0