From nobody Fri Sep 25 03:16:03 2026 Received: from mail-pj1-f71.google.com (mail-pj1-f71.google.com [209.85.216.71]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C44F843B6F5 for ; Thu, 17 Sep 2026 06:42:50 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.71 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789627372; cv=none; b=W7L+Le6h7Kvba37ewwT+nZSf5XPv0tnxoWEr/4u5oM2uqcX86rtoMT7wMiJ8qyvD3Mm6R8nQpauBiH5+Jnc67vWnlk7B1OQHYbAq32QnRcKvCnO4W5wUJmLJqPgctJUMj/ffTsoCfl/6lDk8xZEivkNyectz0R8sAO497FD7oBA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789627372; c=relaxed/simple; bh=AIhDuxYXm6TPBGEUGmeWMYRdVi5UgiMgDR7MkYv1egM=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=D1RYx5UHXC8k7i3EZMISH6ZvO9Y4bBwcJ0Gt84NCQGynwcriu+CjjvRCIsQ4xrYsdWdDbCSVgSMXJqHjXnrzEeyBjWeWraYmUPu+nd922IiPQvbdbnTjmmt8yLwx7wqG2wcVQYK8iG2cHRHxoaggis3hj8rL2pn9pvoBRBRbV/w= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--irogers.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=N6X5waKt; arc=none smtp.client-ip=209.85.216.71 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--irogers.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="N6X5waKt" Received: by mail-pj1-f71.google.com with SMTP id 98e67ed59e1d1-39af92138f9so1040773a91.0 for ; Wed, 16 Sep 2026 23:42:50 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789627370; x=1790232170; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=XZ4suTwVR71QEmnzgMO/8pQXqTLVqgY6KmKt7zHzuV0=; b=N6X5waKtBrJ3DFkBvIxotV62jCMLeBTvHVZEA7vl/y3xuU6aoeA76Pbms3qPYLRMMP BbuFqC5XNAhEBUgduPeaqfSXffpD0tWQWufDwxvghPpBc/6eF3o6un5pFfHBg1dBm3yb ZWX1piL2aojnbC3Hb3fQx60+ieAphoAQH+fKd1Rvvwv8NH0YVO95MR+aMTydawuwKbqS dHKpwhPH6a3f5BDvJlbzVmli1Tai11cTPys2fSURNr4Lu394QzhuMzvvVp9ydLfSoEiE 0u9vy9I0WNd4xasa3ua8t/ixbWjB+WoPyNAnpiKQSKNSDELKZHbBl8oR6PmRWlgFv5+e 4u6g== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789627370; x=1790232170; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=XZ4suTwVR71QEmnzgMO/8pQXqTLVqgY6KmKt7zHzuV0=; b=BlSLXRcho9ItdwbzQI9GXe5czzs33hjcXlayod8ujZOg4e103QukzrqkIXdKZaDyq0 pBGqAWy2SMXQee5ADMVXoCJYt8xCzgdWOs2ffnIbkNZvTrhh+XfU4SzOs6RzJp9Aob8T /jqFPgWdOsUg8D2Jt2hqVg8vyN1JUcy/O/aaBXfIlC9YiXYshor0/PvquoaJ1IxtmbI6 Iic6idNlFXr+IMG7GYZaRO69w0hCGowTD/wn/098sRjG05K+icT0RDNipnI5Fhjm7gdj 0SKiLrDu3rQ9RE3oGqHOQOdR9v5t20ieI96dtzfAfdU3ytpIx8QKn0kEoNJzFZzB3kuW jiuQ== X-Forwarded-Encrypted: i=1; AKwUvByb41KCo/+0OIsggWeu6BRslLqr/95VzpEDwA+q74YfiI/lf35ywoPwdPIwrZvKoZNBil8Sivye4nJGF8E=@vger.kernel.org X-Gm-Message-State: AFuF++nwNCLz9cSqFNAZC0r7EMbQlwNY9eoleG1fLGXxBTfJ67qjZ0Qt z5fBA4XUB/nHXwFDmwp1du/FsxJqMVxR5eRge+fXzqImSv4Pnv9Dpza2HMb6nnFItzrv7pubXdO rfnEZifu3xQ== X-Received: from dleb1-n2.prod.google.com ([2002:a05:701b:4241:20b0:144:4896:d41]) (user=irogers job=prod-delivery.src-stubby-dispatcher) by 2002:a17:90b:50cc:b0:39e:fd1:bd46 with SMTP id 98e67ed59e1d1-39e35da22efmr4205823a91.7.1789627369960; Wed, 16 Sep 2026 23:42:49 -0700 (PDT) Date: Wed, 16 Sep 2026 23:42:23 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <91623303ca281e878444068063fba4bec464bc60.1789626978.git.irogers@google.com> Subject: [PATCH v1 01/13] perf trace: Start BPF summary before starting workload From: Ian Rogers To: Arnaldo Carvalho de Melo , Namhyung Kim Cc: Peter Zijlstra , Ingo Molnar , Jiri Olsa , Adrian Hunter , James Clark , linux-perf-users@vger.kernel.org, linux-kernel@vger.kernel.org, Ian Rogers Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" When using --bpf-summary, trace_start_bpf_summary() sets skel->bss->enabled =3D 1. In trace__run(), trace_start_bpf_summary() was previously invoked after evlist__start_workload(). Because evlist__start_workload() immediately unblocks the child process by writing to its go_pipe, short-lived workloads (such as `cat /dev/null`) can execute and complete their initial system calls before trace_start_bpf_summary() is reached by the parent process. Furthermore, under high system load, the child process may finish before the BPF summary tracking is enabled in the kernel at all, causing syscall summary tests to fail. Additionally, if initial_delay was configured, the workload was started before sleeping. Move trace_start_bpf_summary() to be invoked before evlist__start_workload(), matching evlist__enable(), and ensure it respects target.initial_delay. Assisted-by: Antigravity:gemini-3.1-pro Signed-off-by: Ian Rogers --- tools/perf/builtin-trace.c | 8 +++++--- 1 file changed, 5 insertions(+), 3 deletions(-) diff --git a/tools/perf/builtin-trace.c b/tools/perf/builtin-trace.c index 20fffc24507b..8da0c51ec380 100644 --- a/tools/perf/builtin-trace.c +++ b/tools/perf/builtin-trace.c @@ -4838,17 +4838,19 @@ static int trace__run(struct trace *trace, int argc= , const char **argv) if (!target__none(&trace->opts.target) && !trace->opts.target.initial_del= ay) evlist__enable(evlist); =20 + if (trace->summary_bpf && !trace->opts.target.initial_delay) + trace_start_bpf_summary(); + if (forks) evlist__start_workload(evlist); =20 if (trace->opts.target.initial_delay) { usleep(trace->opts.target.initial_delay * 1000); evlist__enable(evlist); + if (trace->summary_bpf) + trace_start_bpf_summary(); } =20 - if (trace->summary_bpf) - trace_start_bpf_summary(); - trace->multiple_threads =3D perf_thread_map__pid(evlist__core(evlist)->th= reads, 0) =3D=3D -1 || perf_thread_map__nr(evlist__core(evlist)->threads) > 1 || evlist__first(evlist)->core.attr.inherit; --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Fri Sep 25 03:16:03 2026 Received: from mail-pj1-f69.google.com (mail-pj1-f69.google.com [209.85.216.69]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DA18A43DA2C for ; Thu, 17 Sep 2026 06:42:52 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.69 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789627374; cv=none; b=Jag2R4eZUfifqy0QPzci+nCxitn/GkQ8xDE3WKiU59AIk+YsYfoRnE8YUV98Lz4GCIbsyMmnl10Xa7Vf7LMHw8MVJOgZf145TOP3Xb89pbFvRe7AJ1nH5lNR3cssBSEOFIkwBsZUP6NB/LMm5DYTwot5CmAvWwsM3eyWeJRrSuE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789627374; c=relaxed/simple; bh=dcUOLow5aChmhw1kMzamxREcpQrCRtFEy2QrDEznGAQ=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=spwrO+hkN7dbZxSdWp9WOTzWFqsBvIeEKNErJEMPZJ3uNly2fvCj0t65AgA9yB4ylaidAaiHwn0yA1b+a5cdZVIM6bE2+pYxytar0im+L5VRrOlYuUeJlqm2g5JfBwPr9PmvBGNoQFbgrn4PSg+lcsnl7gdGuRkuOCxp5vk/33Y= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--irogers.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=uMml9T3f; arc=none smtp.client-ip=209.85.216.69 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--irogers.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="uMml9T3f" Received: by mail-pj1-f69.google.com with SMTP id 98e67ed59e1d1-396901263b6so1009431a91.2 for ; Wed, 16 Sep 2026 23:42:52 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789627372; x=1790232172; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=RXFxG2hpxuzjtIT76zgvB57Mi03KVQvKbpMxQ+Gclps=; b=uMml9T3fcBzMGrixJGbQfiKYtkH/WN3u43VDpJWRMbJeaz7OeBNHzGNSJt4jvg/BaB EZOep3t9uM9JjLcHj9axkGNejrXWgmQ5m2+VhzJM9TYZDmx6WdZi/4FjojdLwk6WqJsi RbIBz/tsWxbq8HWYKQM4OvOKza7F16p/ecpvFo9zP4p5xyNnFzzJh9RG0xknUqNkMZX0 l+cgrBr4WXU4qjRVdsMBgrT7QdPQNAMphzZ06H39p+kzT5lRi2rvdvmAgY1Ypj0VGRkj vPwQjweWzDzYuGjrn/nrL4M2Vstzk0wQgF0IjonWRydkO96e1oKmmknaJkgKe8xV6I/Q u+dQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789627372; x=1790232172; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=RXFxG2hpxuzjtIT76zgvB57Mi03KVQvKbpMxQ+Gclps=; b=Efmyee/SfF3gYV8MBI4zoEMkBvBnOHEAFfMPQqrdHRYVV0sanYG05RwXxisMswW7Qk Tdiwv8W3wLwG6d12ZMcIMOT+iJYwUg3rhx5Y8TULpzz0p4DEF7nGPYYquoX1MGXq9UX7 86ZoLY5IQW1VYFVzr0dYkMzZHfcqK0OUL3cLCe4846dT8QZZd1UzhzwA4NB98TR4sP0C nzQhJVBg0x8aYtIL0b5Iy7XjDFQMKirE1F3KSdogZ9g9pEjXpsgyG7RCd9zTDIVrKw9Z ygqAaSW+F7fy1TAYro+ctmGmMlE8+SMtbzzV8ePLH5Sr8w6m+JmQenQcYayE/iDq3Br9 ZT+w== X-Forwarded-Encrypted: i=1; AKwUvBxNKEjkveHXF7LNF16X8RPEkOm5fCD43U8SndTZP5DoHXex/ME8vow/TKvRtHJGMHr6/OxPy0AiRq59wKE=@vger.kernel.org X-Gm-Message-State: AFuF++l/ozg5ksPl35UNgAkIW8KjoI1cPdyUXJidHEqIX1XQMuj/DpkZ waUR3PaYkzITtUPUX52Q1oZJPxGraULhM1CGA/fFyq2yEqd7uxlHjIXpWbKxFoQPCCWWgzOV8f6 Kau98frwM0w== X-Received: from dlev4.prod.google.com ([2002:a05:701b:4644:b0:143:8502:417a]) (user=irogers job=prod-delivery.src-stubby-dispatcher) by 2002:a17:90b:582c:b0:39d:f18c:1ad1 with SMTP id 98e67ed59e1d1-39e1e2501e6mr19387622a91.3.1789627371882; Wed, 16 Sep 2026 23:42:51 -0700 (PDT) Date: Wed, 16 Sep 2026 23:42:24 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <79b1a85a5a40701e57f72f6a205148f2c1155058.1789626978.git.irogers@google.com> Subject: [PATCH v1 02/13] perf trace: Skip internal tracepoint fields in formatting and beauty map From: Ian Rogers To: Arnaldo Carvalho de Melo , Namhyung Kim Cc: Peter Zijlstra , Ingo Molnar , Jiri Olsa , Adrian Hunter , James Clark , linux-perf-users@vger.kernel.org, linux-kernel@vger.kernel.org, Ian Rogers Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Linux 6.19+ added __data_loc char[] internal fields for string arguments in syscalls:sys_enter_ tracepoints (e.g., __data_loc_oldname in sys_enter_renameat2). While is_internal_field() was added to detect them, several places did not properly account for them: 1. In syscall_arg_fmt__init_array(), when an internal field was skipped, the arg pointer was still incremented, causing the subsequent argument formatters to be mismatched. 2. In syscall__scnprintf_args(), internal fields were not skipped, causing spurious trailing arguments like ", 0, 16" to be formatted and printed. 3. In trace__bpf_sys_enter_beauty_map(), internal fields were not skipped, offsetting beauty array argument indices and breaking string and buffer augmentation. 4. In trace__find_usable_bpf_prog_entry(), candidate pointer checks matched on internal pointer fields, breaking signature compatibility matching between syscalls for augmenter sharing. Introduce next_user_arg() and advance both cursors with it, so that the two argument lists are always compared at a real argument and the walk ends when one syscall runs out of arguments rather than when one happens to have trailing internal fields. 5. syscall__augmented_args() computed the augmented payload as sample->raw_size - sc->args_size for any sys_enter style sample. sc->args_size deliberately stops at the last non-internal field, so on 6.19+ a native syscalls:sys_enter_ record leaves the __data_loc words and their string payloads in the remainder. Those bytes are not a struct augmented_arg, so syscall_arg__scnprintf_augmented_string() read a bogus length and walked arg->augmented.args out of bounds. This is reachable from trace__event_handler(), which calls trace__fprintf_sys_enter() for any evsel whose tracepoint name starts with "sys_enter_". 6. In syscall__read_info(), syscall__alloc_arg_fmts() was called before checking and dropping the leading __syscall_nr (or nr) field, using nr_fields - 1 unconditionally. If a tracepoint format lacks that leading field, the allocated arg_fmt array is one entry too small and syscall_arg_fmt__init_array() writes one entry past the end of the heap buffer. Drop __syscall_nr/nr first and size the allocation from the remaining fields. Update these functions to check and skip is_internal_field() so that arguments are correctly formatted and beauty map entries match the expected syscall signatures, restrict syscall__augmented_args() to the __augmented_syscalls__ bpf-output evsel, and size arg_fmt after dropping the syscall number field. Assisted-by: Antigravity:gemini-3.1-pro Signed-off-by: Ian Rogers --- tools/perf/builtin-trace.c | 153 ++++++++++++++++++++++++++++--------- 1 file changed, 115 insertions(+), 38 deletions(-) diff --git a/tools/perf/builtin-trace.c b/tools/perf/builtin-trace.c index 8da0c51ec380..af9696aadaec 100644 --- a/tools/perf/builtin-trace.c +++ b/tools/perf/builtin-trace.c @@ -2277,15 +2277,20 @@ syscall_arg_fmt__init_array(struct syscall_arg_fmt = *arg, struct tep_format_field struct tep_format_field *last_field =3D NULL; int len; =20 - for (; field; field =3D field->next, ++arg) { - /* assume it's the last argument */ + for (; field; field =3D field->next) { + /* + * Skip internal tracepoint fields (e.g., __data_loc strings in + * Linux 6.19+) so they do not advance the syscall arg array index. + */ if (is_internal_field(field)) continue; =20 last_field =3D field; =20 - if (arg->scnprintf) + if (arg->scnprintf) { + ++arg; continue; + } =20 len =3D strlen(field->name); =20 @@ -2342,6 +2347,7 @@ syscall_arg_fmt__init_array(struct syscall_arg_fmt *a= rg, struct tep_format_field } } } + ++arg; } =20 return last_field; @@ -2363,6 +2369,7 @@ static int syscall__read_info(struct syscall *sc, str= uct trace *trace) char tp_name[128]; const char *name; struct tep_format_field *field; + int nr_args; int err; =20 if (sc->nonexistent) @@ -2401,24 +2408,25 @@ static int syscall__read_info(struct syscall *sc, s= truct trace *trace) return err; } =20 - /* - * The tracepoint format contains __syscall_nr field, so it's one more - * than the actual number of syscall arguments. - */ - if (syscall__alloc_arg_fmts(sc, sc->tp_format->format.nr_fields - 1)) - return -ENOMEM; - sc->args =3D sc->tp_format->format.fields; + nr_args =3D sc->tp_format->format.nr_fields; /* * We need to check and discard the first variable '__syscall_nr' * or 'nr' that mean the syscall number. It is needless here. * So drop '__syscall_nr' or 'nr' field but does not exist on older kerne= ls. + * + * Do this before allocating, and size the array from what is left, so + * that a format without the field does not leave + * syscall_arg_fmt__init_array() walking one entry past the end. */ if (sc->args && (!strcmp(sc->args->name, "__syscall_nr") || !strcmp(sc->a= rgs->name, "nr"))) { sc->args =3D sc->args->next; - --sc->nr_args; + --nr_args; } =20 + if (syscall__alloc_arg_fmts(sc, nr_args)) + return -ENOMEM; + field =3D sc->args; while (field) { if (is_internal_field(field)) @@ -2636,11 +2644,17 @@ static size_t syscall__scnprintf_args(struct syscal= l *sc, char *bf, size_t size, if (sc->args !=3D NULL) { struct tep_format_field *field; =20 - for (field =3D sc->args; field; - field =3D field->next, ++arg.idx, bit <<=3D 1) { - if (arg.mask & bit) + for (field =3D sc->args; field; field =3D field->next) { + /* Skip internal fields so they are not printed as spurious arguments */ + if (is_internal_field(field)) continue; =20 + if (arg.mask & bit) { + ++arg.idx; + bit <<=3D 1; + continue; + } + arg.fmt =3D &sc->arg_fmt[arg.idx]; val =3D syscall_arg__val(&arg, arg.idx); /* @@ -2658,8 +2672,11 @@ static size_t syscall__scnprintf_args(struct syscall= *sc, char *bf, size_t size, */ if (val =3D=3D 0 && !trace->show_zeros && !(sc->arg_fmt && sc->arg_fmt[arg.idx].show_zero) && - !(sc->arg_fmt && sc->arg_fmt[arg.idx].strtoul =3D=3D STUL_BTF_TYPE)) + !(sc->arg_fmt && sc->arg_fmt[arg.idx].strtoul =3D=3D STUL_BTF_TYPE)= ) { + ++arg.idx; + bit <<=3D 1; continue; + } =20 printed +=3D scnprintf(bf + printed, size - printed, "%s", printed ? ",= " : ""); =20 @@ -2674,12 +2691,16 @@ static size_t syscall__scnprintf_args(struct syscal= l *sc, char *bf, size_t size, size - printed, val, field->type); if (btf_printed) { printed +=3D btf_printed; + ++arg.idx; + bit <<=3D 1; continue; } } =20 printed +=3D syscall_arg_fmt__scnprintf_val(&sc->arg_fmt[arg.idx], bf + printed, size - printed, &arg, val); + ++arg.idx; + bit <<=3D 1; } } else if (IS_ERR(sc->tp_format)) { /* @@ -2940,7 +2961,9 @@ static int trace__fprintf_sample(struct trace *trace,= struct perf_sample *sample return printed; } =20 -static void *syscall__augmented_args(struct syscall *sc, struct perf_sampl= e *sample, int *augmented_args_size, int raw_augmented_args_size) +static void *syscall__augmented_args(struct trace *trace, struct syscall *= sc, + struct perf_sample *sample, + int *augmented_args_size, int raw_augmented_args_size) { /* * For now with BPF raw_augmented we hook into raw_syscalls:sys_enter @@ -2958,6 +2981,24 @@ static void *syscall__augmented_args(struct syscall = *sc, struct perf_sample *sam */ int args_size =3D raw_augmented_args_size ?: sc->args_size; =20 + /* + * Augmented arguments are a perf trace specific payload, they are only + * ever appended to samples emitted by the BPF __augmented_syscalls__ + * bpf-output event. + * + * Native syscalls:sys_enter_NAME tracepoints may also carry trailing + * data of their own: since Linux 6.19 they append __data_loc char[] + * fields plus the string payloads they point at. Those bytes are not a + * struct augmented_arg, so treating them as one would make + * syscall_arg__scnprintf_augmented_string() read a bogus length and + * walk arg->augmented.args far out of bounds. + * + * So only look for augmented arguments on the event that can actually + * produce them. + */ + if (sample->evsel !=3D trace->syscalls.events.bpf_output) + return NULL; + *augmented_args_size =3D sample->raw_size - args_size; if (*augmented_args_size > 0) { static uintptr_t argbuf[1024]; /* assuming single-threaded */ @@ -3016,17 +3057,13 @@ static int trace__sys_enter(struct trace *trace, if (!(trace->duration_filter || trace->summary_only || trace->min_stack)) trace__printf_interrupted_entry(trace); /* - * If this is raw_syscalls.sys_enter, then it always comes with the 6 pos= sible - * arguments, even if the syscall being handled, say "openat", uses only = 4 arguments - * this breaks syscall__augmented_args() check for augmented args, as we = calculate - * syscall->args_size using each syscalls:sys_enter_NAME tracefs format f= ile, - * so when handling, say the openat syscall, we end up getting 6 args for= the - * raw_syscalls:sys_enter event, when we expected just 4, we end up mista= kenly - * thinking that the extra 2 u64 args are the augmented filename, so just= check - * here and avoid using augmented syscalls when the evsel is the raw_sysc= alls one. + * syscall__augmented_args() only returns a payload for the BPF + * __augmented_syscalls__ event, so raw_syscalls:sys_enter (which always + * carries all 6 possible arguments rather than sc->args_size worth) and + * the native syscalls:sys_enter_NAME tracepoints are both handled there. */ - if (evsel !=3D trace->syscalls.events.sys_enter) - augmented_args =3D syscall__augmented_args(sc, sample, &augmented_args_s= ize, trace->raw_augmented_syscalls_args_size); + augmented_args =3D syscall__augmented_args(trace, sc, sample, &augmented_= args_size, + trace->raw_augmented_syscalls_args_size); ttrace->entry_time =3D sample->time; ttrace->entry_cpu =3D sample->cpu; msg =3D ttrace->entry_str; @@ -3071,7 +3108,7 @@ static int trace__fprintf_sys_enter(struct trace *tra= ce, struct perf_sample *sam struct syscall *sc; char msg[1024]; void *args, *augmented_args =3D NULL; - int augmented_args_size, e_machine; + int augmented_args_size =3D 0, e_machine; size_t printed =3D 0; =20 =20 @@ -3089,7 +3126,8 @@ static int trace__fprintf_sys_enter(struct trace *tra= ce, struct perf_sample *sam goto out_put; =20 args =3D perf_evsel__sc_tp_ptr(args, sample); - augmented_args =3D syscall__augmented_args(sc, sample, &augmented_args_si= ze, trace->raw_augmented_syscalls_args_size); + augmented_args =3D syscall__augmented_args(trace, sc, sample, &augmented_= args_size, + trace->raw_augmented_syscalls_args_size); printed +=3D syscall__scnprintf_args(sc, msg, sizeof(msg), args, augmente= d_args, augmented_args_size, trace, thread); fprintf(trace->output, "%.*s", (int)printed, msg); err =3D 0; @@ -4121,10 +4159,16 @@ static int trace__bpf_sys_enter_beauty_map(struct t= race *trace, int e_machine, i if (trace->btf =3D=3D NULL) return -1; =20 - for (i =3D 0, field =3D sc->args; field; ++i, field =3D field->next) { + for (i =3D 0, field =3D sc->args; field; field =3D field->next) { + /* Skip internal fields to keep beauty array index aligned with syscall = arguments */ + if (is_internal_field(field)) + continue; + // XXX We're only collecting pointer payloads _from_ user space - if (!sc->arg_fmt[i].from_user) + if (!sc->arg_fmt[i].from_user) { + ++i; continue; + } =20 struct_offset =3D strstr(field->type, "struct "); if (struct_offset =3D=3D NULL) @@ -4143,8 +4187,10 @@ static int trace__bpf_sys_enter_beauty_map(struct tr= ace *trace, int e_machine, i name[cnt] =3D '\0'; =20 /* cache struct's btf_type and type_id */ - if (syscall_arg_fmt__cache_btf_struct(&sc->arg_fmt[i], trace->btf, name= )) + if (syscall_arg_fmt__cache_btf_struct(&sc->arg_fmt[i], trace->btf, name= )) { + ++i; continue; + } =20 bt =3D sc->arg_fmt[i].type; beauty_array[i] =3D bt->size; @@ -4170,7 +4216,9 @@ static int trace__bpf_sys_enter_beauty_map(struct tra= ce *trace, int e_machine, i struct tep_format_field *field_tmp; =20 /* find the size of the buffer that appears in pairs with buf */ - for (j =3D 0, field_tmp =3D sc->args; field_tmp; ++j, field_tmp =3D fie= ld_tmp->next) { + for (j =3D 0, field_tmp =3D sc->args; field_tmp; field_tmp =3D field_tm= p->next) { + if (is_internal_field(field_tmp)) + continue; if (!(field_tmp->flags & TEP_FIELD_IS_POINTER) && /* only integers */ (strstr(field_tmp->name, "count") || strstr(field_tmp->name, "siz") || /* size, bufsiz */ @@ -4180,8 +4228,10 @@ static int trace__bpf_sys_enter_beauty_map(struct tr= ace *trace, int e_machine, i can_augment =3D true; break; } + ++j; } } + ++i; } =20 if (can_augment) @@ -4190,6 +4240,19 @@ static int trace__bpf_sys_enter_beauty_map(struct tr= ace *trace, int e_machine, i return -1; } =20 +/* + * Advance to the first field that is a real syscall argument, so that cal= lers + * walking two argument lists in step never have to reason about internal + * fields appearing in one list but not the other. + */ +static struct tep_format_field *next_user_arg(struct tep_format_field *fie= ld) +{ + while (field && is_internal_field(field)) + field =3D field->next; + + return field; +} + static struct bpf_program *trace__find_usable_bpf_prog_entry(struct trace = *trace, struct syscall *sc) { @@ -4197,7 +4260,7 @@ static struct bpf_program *trace__find_usable_bpf_pro= g_entry(struct trace *trace /* * We're only interested in syscalls that have a pointer: */ - for (field =3D sc->args; field; field =3D field->next) { + for (field =3D next_user_arg(sc->args); field; field =3D next_user_arg(fi= eld->next)) { if (field->flags & TEP_FIELD_IS_POINTER) goto try_to_find_pair; } @@ -4215,21 +4278,31 @@ static struct bpf_program *trace__find_usable_bpf_p= rog_entry(struct trace *trace pair->bpf_prog.sys_enter =3D=3D unaugmented_prog) continue; =20 - for (field =3D sc->args, candidate_field =3D pair->args; - field && candidate_field; field =3D field->next, candidate_field = =3D candidate_field->next) { + /* + * Both cursors only ever point at real arguments, so the loop + * ends when one of the two syscalls runs out of them, rather + * than when one happens to have trailing internal fields. + */ + field =3D next_user_arg(sc->args); + candidate_field =3D next_user_arg(pair->args); + while (field && candidate_field) { bool is_pointer =3D field->flags & TEP_FIELD_IS_POINTER, candidate_is_pointer =3D candidate_field->flags & TEP_FIELD_IS_POI= NTER; =20 if (is_pointer) { - if (!candidate_is_pointer) { + if (!candidate_is_pointer) { // The candidate just doesn't copies our pointer arg, might copy othe= r pointers we want. + field =3D next_user_arg(field->next); + candidate_field =3D next_user_arg(candidate_field->next); continue; - } + } } else { if (candidate_is_pointer) { // The candidate might copy a pointer we don't have, skip it. goto next_candidate; } + field =3D next_user_arg(field->next); + candidate_field =3D next_user_arg(candidate_field->next); continue; } =20 @@ -4250,6 +4323,8 @@ static struct bpf_program *trace__find_usable_bpf_pro= g_entry(struct trace *trace goto next_candidate; =20 is_candidate =3D true; + field =3D next_user_arg(field->next); + candidate_field =3D next_user_arg(candidate_field->next); } =20 if (!is_candidate) @@ -4261,7 +4336,9 @@ static struct bpf_program *trace__find_usable_bpf_pro= g_entry(struct trace *trace * more than what is common to the two syscalls. */ if (candidate_field) { - for (candidate_field =3D candidate_field->next; candidate_field; candid= ate_field =3D candidate_field->next) + candidate_field =3D next_user_arg(candidate_field->next); + for (; candidate_field; + candidate_field =3D next_user_arg(candidate_field->next)) if (candidate_field->flags & TEP_FIELD_IS_POINTER) goto next_candidate; } --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Fri Sep 25 03:16:03 2026 Received: from mail-pl1-f199.google.com (mail-pl1-f199.google.com [209.85.214.199]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 98D3843C05B for ; Thu, 17 Sep 2026 06:42:54 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.199 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789627376; cv=none; b=St4DwEd1A3hzcHmv1V+XSf88fmAh1z4x3eZnMe6p80JEBnXlIQ2PQ9Cdm5HoJS0978q9F3Dgxh0n6gGoASMtV5SkyoUi2uiBWNf/Wi8JLOQ05gwpnsVRQflX/d08ooKed5j2X9kcyWAcsXTYrbECRzgnMr8qL6G+mTFHXlXulXo= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789627376; c=relaxed/simple; bh=lx7wyF3Qu/HpKTFDnfOd4iCQJKE+aEvsCq9kKuv30HU=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=Kuk2O5vjuvdsbp1LF+0hssSpyoWKxek/q3luRRBXweH3g9t8Cb11WZFB++EmN5G79cr7s/m3HQvfcv9CMItV1Rat+ETot/zR1E7tZr22yW9GfRRA8TrJ0P0q//Jmz7QBj5x5trAOCUjIhzWFq1O+Ng9UkqdLgqVfZMra6B2mtWM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--irogers.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=TiqMasi4; arc=none smtp.client-ip=209.85.214.199 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--irogers.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="TiqMasi4" Received: by mail-pl1-f199.google.com with SMTP id d9443c01a7336-2dcfed75f77so8613295ad.2 for ; Wed, 16 Sep 2026 23:42:54 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789627374; x=1790232174; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=iDa4jCWG1HgPZsp/vqH9L8ol0+fyo/nirzG/z+7DOWE=; b=TiqMasi4FEpnpLyPFNAufz8q/Lc5FNSsQU/6rN4q3vtL7jika3+1/MLKGlNCo/135L tiMF7XgDJfv45e99eDW1oG8lIrsQAA0LmYpKgm7500xP5lSGQh4bk5S5ddnr/ZbKXZ99 sRN13UoUOvCJ342NgXc2gOCQIo5sDtzZkMW6ci2nOPJSAg5V6//cnJy5/SuTVMTqbFP6 ezg4wzKoaRqN29MihJGWslSDyOtXfjqdrxuWpqmQrJ9uBVJQWvf+z2dLfuxeJ9a9R0Vl ki3nSG4aIk0Spu2UTSMZyp1OTfmGLOPCppZnkSeJPlVx3eo8YkUepgUdQsXyOdU1N5l+ eqCg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789627374; x=1790232174; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=iDa4jCWG1HgPZsp/vqH9L8ol0+fyo/nirzG/z+7DOWE=; b=ToiNSANJ0BYV3USNrgwYcWHU0mSgUgqFHf1Br5WlIQmNsnJQc2frKuKEHvWdn/agoF 4jBsqOGtGRL1tPrSsKx21KFpQY5todaE/TU1bfwQO5d4Fw/BiBzs1ygSENek8rjpS/XE SxVdb4G96lHdkAXFjhwPYX6DsPNE6Pk7C/enJmqwBOESL9qPYDZgo2kkk+qtp/AZpPkd /P1W6N8Jq7pZ2ZJTX4ahz6ntbnhgjaQCcML3iArEcaslHhsHIj5OYhxu0l4fSMUfVsPT hwjwcKy+Rnpu4I2Sc4Pv6E/pZYyrpcpnyOrYINIup8NP6vxTPgphWxcRj5rCnW9H/Bnc aBRA== X-Forwarded-Encrypted: i=1; AKwUvByrIdGGVtfSUOm3CGJyWJxp/LB3U5rlAQyc59HEAbV0BFkIpPA+HTIqvfX/2HWHRugJtdfpgfQP3HiS/Es=@vger.kernel.org X-Gm-Message-State: AFuF++lBFu04UGfYE7eijwiVA8DBOQPW3quPUtyn1sh1JnAsLfizVL8w FrpH2fKZKZ5/q46W/JpM4+zgvDG0B2z/6bt68WHkrPC+G4wCxWMgXQeMaZwAdKQEcoQEG86/SP0 djtwBuywpKw== X-Received: from dlbrl7.prod.google.com ([2002:a05:7022:f507:b0:143:94b1:ba45]) (user=irogers job=prod-delivery.src-stubby-dispatcher) by 2002:a17:90b:2803:b0:39d:ef3e:9035 with SMTP id 98e67ed59e1d1-39e1e324463mr18457457a91.10.1789627373850; Wed, 16 Sep 2026 23:42:53 -0700 (PDT) Date: Wed, 16 Sep 2026 23:42:25 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: Subject: [PATCH v1 03/13] perf trace: Do not set unaugmented BPF program on sys_exit map From: Ian Rogers To: Arnaldo Carvalho de Melo , Namhyung Kim Cc: Peter Zijlstra , Ingo Molnar , Jiri Olsa , Adrian Hunter , James Clark , linux-perf-users@vger.kernel.org, linux-kernel@vger.kernel.org, Ian Rogers Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" In trace__init_syscalls_bpf_prog_array_maps(), the BPF program array map for sys_exit (syscalls_sys_exit) was populated with the result of trace__bpf_prog_sys_exit_fd(). When a syscall had no specific exit augmenter, trace__find_syscall_bpf_prog() fell back to unaugmented_prog (syscall_unaugmented). However, syscall_unaugmented is a sys_enter program that outputs enter arguments to __augmented_syscalls__. As a consequence, when an unaugmented syscall exited, sys_exit tail-called syscall_unaugmented, which interpreted the exit arguments as enter arguments and emitted a duplicate, corrupt sys_enter event into __augmented_syscalls__ right as the syscall completed. Fix this by: 1. Returning NULL from trace__find_syscall_bpf_prog() when looking up exit augmenters and none is found. 2. Returning -1 from trace__bpf_prog_sys_exit_fd() when no exit program is present. 3. Only updating map_exit_fd when prog_fd >=3D 0. 4. Clearing err =3D 0 when trace__bpf_sys_enter_beauty_map() returns non-zero (indicating the syscall has no augmentable pointer arguments) before continuing the loop, so a trailing run of such syscalls (e.g. 'perf trace -e close') does not leave err non-zero on return and abort the session. Assisted-by: Antigravity:gemini-3.1-pro Signed-off-by: Ian Rogers --- tools/perf/builtin-trace.c | 28 ++++++++++++++++++++++------ 1 file changed, 22 insertions(+), 6 deletions(-) diff --git a/tools/perf/builtin-trace.c b/tools/perf/builtin-trace.c index af9696aadaec..a68d34256996 100644 --- a/tools/perf/builtin-trace.c +++ b/tools/perf/builtin-trace.c @@ -4117,7 +4117,12 @@ static struct bpf_program *trace__find_syscall_bpf_p= rog(struct trace *trace __ma pr_debug("Couldn't find BPF prog \"%s\" to associate with syscalls:sys_%s= _%s, not augmenting it\n", prog_name, type, sc->name); out_unaugmented: - return unaugmented_prog; + /* + * Do not set unaugmented_prog for exit: syscall_unaugmented is a + * sys_enter program that outputs enter arguments. Exit without a + * specialized return augmenter returns 1 directly from sys_exit. + */ + return !strcmp(type, "exit") ? NULL : unaugmented_prog; } =20 static void trace__init_syscall_bpf_progs(struct trace *trace, int e_machi= ne, int id) @@ -4140,7 +4145,7 @@ static int trace__bpf_prog_sys_enter_fd(struct trace = *trace, int e_machine, int static int trace__bpf_prog_sys_exit_fd(struct trace *trace, int e_machine,= int id) { struct syscall *sc =3D trace__syscall_info(trace, NULL, e_machine, id); - return sc ? bpf_program__fd(sc->bpf_prog.sys_exit) : bpf_program__fd(unau= gmented_prog); + return sc && sc->bpf_prog.sys_exit ? bpf_program__fd(sc->bpf_prog.sys_exi= t) : -1; } =20 static int trace__bpf_sys_enter_beauty_map(struct trace *trace, int e_mach= ine, int key, unsigned int *beauty_array) @@ -4395,16 +4400,27 @@ static int trace__init_syscalls_bpf_prog_array_maps= (struct trace *trace, int e_m err =3D bpf_map_update_elem(map_enter_fd, &key, &prog_fd, BPF_ANY); if (err) break; + /* Only update the exit prog array map if an exit augmenter exists */ prog_fd =3D trace__bpf_prog_sys_exit_fd(trace, e_machine, key); - err =3D bpf_map_update_elem(map_exit_fd, &key, &prog_fd, BPF_ANY); - if (err) - break; + if (prog_fd >=3D 0) { + err =3D bpf_map_update_elem(map_exit_fd, &key, &prog_fd, BPF_ANY); + if (err) + break; + } =20 /* use beauty_map to tell BPF how many bytes to collect, set beauty_map'= s value here */ memset(beauty_array, 0, sizeof(beauty_array)); err =3D trace__bpf_sys_enter_beauty_map(trace, e_machine, key, (unsigned= int *)beauty_array); - if (err) + if (err) { + /* + * Not a failure: the syscall just has no augmentable + * arguments. Clear err, or a trailing run of such + * syscalls, e.g. all of them for 'perf trace -e close', + * would leave it set on return and abort the session. + */ + err =3D 0; continue; + } err =3D bpf_map_update_elem(beauty_map_fd, &key, beauty_array, BPF_ANY); if (err) break; --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Fri Sep 25 03:16:03 2026 Received: from mail-pg1-f199.google.com (mail-pg1-f199.google.com [209.85.215.199]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D0A7D43E9C6 for ; Thu, 17 Sep 2026 06:42:56 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.199 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789627379; cv=none; b=uwvJcMQ73GHSeguvccYSBGcdHlWyIoM0ozDioaylkqAm3S0sRbl8Dc1Ve3ycMnf07ldY6rpo3z3C/CN+ihvmt3hbCYsR5U2sMd8M/zmOPFfzStS5CW0EAzvyr6lt10eDb7n7qXe7FUYT0ah2UN3idenh6qPWjAow6Ifl3M8AkHI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789627379; c=relaxed/simple; bh=0UaIZcn5Up9teKUUH0KGsftFAImOdwDrtfA/UmQ4Ox0=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=bgFI3HkgdnlCHqOo2A6oJO8x/Brr7SFxPKMfNEdi2WPU6RYQxjYW0kPWWVXG+c/5d3iMrisKZ4NChTTqQFRcAJjljNkRwZDOF6UMxVA47+hAPoe1fHmXmbTlZJ4e69fGte+HPYNyu3MgYAxiy5wPsM9g8SRTWLny1p5HOi/0C3o= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--irogers.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=eyWp3Vt1; arc=none smtp.client-ip=209.85.215.199 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--irogers.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="eyWp3Vt1" Received: by mail-pg1-f199.google.com with SMTP id 41be03b00d2f7-c89704da8c7so755932a12.0 for ; Wed, 16 Sep 2026 23:42:56 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789627376; x=1790232176; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=HG/rJfvlu+ZVWRRlWPs7VziMtPNTB20fH3aBvPtdRcI=; b=eyWp3Vt1X2+6afjQspbZeHA92TU8aCDzR6F01CQezYh/58O8g+4NIXIOeRp8bxw+tm 3p02+2wa0RpzO+MBUYdoQ7GA8EWmk1xJUxlDPAaNj7zKCPwucIT6wLP4J97K00srgLHF /ni87twdn49CRelKDMnyRtAIwJyF82clF2i6emyLo7GXDK8jsXbSGFf7H9buAPoqbWik 5NFV9llYxx586A1jZuMKDLb7nOuwtyvY0Y4CpEgh6GdIWJY6Fxilg9eUpmZjYD7bZn4R WT6cLyt2UDInWxfFaYl6BzJRAFXlZGODXuB6e4pt8/qtmn4PfjaMPmgfCu+C96gkLIc5 +oYQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789627376; x=1790232176; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=HG/rJfvlu+ZVWRRlWPs7VziMtPNTB20fH3aBvPtdRcI=; b=vXdZiwWn4mjjAmYaEyTPJqWSJy1g3UtETATb0+ThT3M7uplzOggGytGta7TXpZNjDS Dq6ySekFgD5NdZ18DOe3aKB0nCDu4KpzK29deg2JBFiPWFm46SQq0h0kfSdWx1MPt9+k xej12MnmOsTaQjN3lX5fR5VrEfSXHkdo+EM6dyLJrGmE66QpgYA1jU4RwkPzo7KTujqD J39AAoclNzk2yI5uzn0k4xmCVbdjBf75fLscrZ6zW/ufgND/OZnbqvBfMk6UOypQ+SI/ a6U5KSrYUXauSXRB0333v2QyUgIN7gieQQIp0RHhFqcSQZ6xIxAqG7/Fw4yjGuLdQutC Re6g== X-Forwarded-Encrypted: i=1; AKwUvBxU3Hf5zHut3ESTVN84P8r/areAv7Jp0V+003mlikjudCpvJZMYaTrUT1wVE1wAoR2BmDhjxZuUoP84kYw=@vger.kernel.org X-Gm-Message-State: AFuF++nHbong70nMXI+HWIapzLeXedpGHHHTsQ7HPsJjhtYu+m5Rs48I JlqwqFTgTFHXCiUrKzP06CBTOkoTPsdvLHVe1HHZ41/mi1T2tA9GO3Xngcz0GSAQb3jTbYQwjae 00bYhh+Zhzg== X-Received: from dldnz11.prod.google.com ([2002:a05:701a:ca0b:b0:143:8cae:a783]) (user=irogers job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a21:4c81:b0:3b4:61f:1fec with SMTP id adf61e73a8af0-3dd5f3c1abcmr13732655637.2.1789627375781; Wed, 16 Sep 2026 23:42:55 -0700 (PDT) Date: Wed, 16 Sep 2026 23:42:26 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <36cdb33f86d7c1d349fb0554f6ee6c3bc621b051.1789626978.git.irogers@google.com> Subject: [PATCH v1 04/13] perf trace: Filter events in BPF and avoid tracepoint vetoes From: Ian Rogers To: Arnaldo Carvalho de Melo , Namhyung Kim Cc: Peter Zijlstra , Ingo Molnar , Jiri Olsa , Adrian Hunter , James Clark , linux-perf-users@vger.kernel.org, linux-kernel@vger.kernel.org, Ian Rogers Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" The BPF augmented_raw_syscalls sys_enter and sys_exit programs returned 0 for syscalls that were not of interest. Returning 0 from a tracepoint BPF program vetoes the event for the whole system, so an unrelated concurrent perf trace, perf record or ftrace session listening to raw_syscalls would silently lose events. This is a cross-session side effect and shows up as flaky failures when perf tests run in parallel. Furthermore, syscall_unaugmented previously returned 1 without writing anything to the __augmented_syscalls__ ring buffer. This forced userspace perf trace to listen to both raw_syscalls:sys_enter and __augmented_syscalls__ in its evlist, requiring userspace event deduplication. Address these issues: 1. In augmented_raw_syscalls.bpf.c, never return 0 from tracepoint handlers: return 1 so non-traced syscalls pass through without vetoing other concurrent listeners. 2. Introduce pids_to_trace and syscalls_to_trace BPF hash maps to perform targeted filtering directly in BPF. Unselected syscalls or PIDs return 1 immediately without writing to the buffer. 3. In syscall_unaugmented, output the unaugmented enter payload into __augmented_syscalls__ and return 1. Change its section from SEC("tp/raw_syscalls/sys_enter") to SEC("tp/syscalls/sys_enter_unaugmented") so libbpf does not attempt to auto-attach it to raw_syscalls:sys_enter. 4. In bpf_trace_augment.c, add helpers to configure target PIDs and syscalls in the BPF maps, setting the activation flags (has_pids_to_trace, has_syscalls_to_trace) only after the maps are fully populated so already-attached BPF programs do not filter against a half-filled map. Explicitly attach only sys_enter and sys_exit via an attach_prog() helper that saves -errno before calling pr_debug() or bpf_link__destroy(). Destroy the skeleton on every failure path. Leaving a loaded but unusable skeleton behind is not inert: the setters called later from trace__run() would program its maps, and a partial attach would leave a BPF program live on raw_syscalls for a session that never starts. Since augmented_syscalls__{prepare,create_bpf_output}() failures fall back to unaugmented tracing rather than aborting, those setters have to become no-ops, which they only do once skel is NULL again. errno is used directly here, so include rather than relying on it arriving via another header, which it does not under musl. 5. In builtin-trace.c, hook trace__set_ev_qualifier_filter() and PID filtering into the BPF maps. When __augmented_syscalls__ is active, remove raw_syscalls:sys_enter from trace.evlist since all traced enter events (both augmented and unaugmented) are now emitted by BPF into __augmented_syscalls__. Restore tracking on the remaining evsel via evlist__set_tracking_event() so PERF_RECORD_COMM and tracking events continue to be recorded. Errors from augmented_syscalls__set_target_syscalls() are reported and propagated, the tracepoint filter string is freed on every exit path, and an allocation failure in trace__set_filter_pids() now returns -ENOMEM instead of being silently ignored. Note that in trace__set_filter_pids() the target pids and the filtered pids are two independent axes and both have to be programmed. Naming pids to leave out with --filter-pids does not widen -p/-t or a workload to the whole system, and a BPF tracepoint program is attached system wide rather than to the target's file descriptors, so pids_to_trace is the only thing keeping other tasks out. 6. Add --syscall-augment option (defaulting to true) to allow users to explicitly use --no-syscall-augment to run perf trace in the classic unaugmented tracepoint mode without BPF. When BPF is unavailable or disabled, ensure the non-augmented tracepoint path cleanly configures sys_enter and sys_exit without duplicate entries. Assisted-by: Antigravity:gemini-3.1-pro Signed-off-by: Ian Rogers --- tools/perf/Documentation/perf-trace.txt | 5 + tools/perf/builtin-trace.c | 146 +++++++++++++++--- .../bpf_skel/augmented_raw_syscalls.bpf.c | 133 ++++++++++++++-- tools/perf/util/bpf_trace_augment.c | 135 +++++++++++++++- tools/perf/util/trace_augment.h | 33 ++++ 5 files changed, 414 insertions(+), 38 deletions(-) diff --git a/tools/perf/Documentation/perf-trace.txt b/tools/perf/Documenta= tion/perf-trace.txt index d20b43ea3d37..4680c69160d7 100644 --- a/tools/perf/Documentation/perf-trace.txt +++ b/tools/perf/Documentation/perf-trace.txt @@ -260,6 +260,11 @@ the thread executes on the designated CPUs. Default is= to monitor all CPUs. Maximum number of lines in the summary mode. Note that this applies to each entry (thread or cgroup). =20 +--syscall-augment:: + Augment syscalls with BPF. Enabled by default when BPF support is availab= le. + Use --no-syscall-augment to disable BPF augmentation and fall back to the + unaugmented tracepoint approach. + =20 PAGEFAULTS ---------- diff --git a/tools/perf/builtin-trace.c b/tools/perf/builtin-trace.c index a68d34256996..e21b2b4a8794 100644 --- a/tools/perf/builtin-trace.c +++ b/tools/perf/builtin-trace.c @@ -200,6 +200,7 @@ struct trace { int max_summary; int raw_augmented_syscalls_args_size; bool raw_augmented_syscalls; + bool syscall_augment; bool fd_path_disabled; bool sort_events; bool not_ev_qualifier; @@ -2054,6 +2055,23 @@ static int trace__process_event(struct trace *trace,= struct machine *machine, "LOST %" PRIu64 " events!\n", (u64)event->lost.lost); ret =3D machine__process_lost_event(machine, event, sample); break; + case PERF_RECORD_FORK: + if (trace->raw_augmented_syscalls && + (augmented_syscalls__has_target_pid(event->fork.ppid) || + augmented_syscalls__has_target_pid(event->fork.ptid))) { + augmented_syscalls__add_target_pid(event->fork.pid); + } + ret =3D machine__process_fork_event(machine, event, sample); + break; + case PERF_RECORD_EXIT: + if (trace->raw_augmented_syscalls) { + if (event->fork.pid =3D=3D event->fork.tid) + augmented_syscalls__del_target_pid(event->fork.pid); + else + augmented_syscalls__del_target_pid(event->fork.tid); + } + ret =3D machine__process_exit_event(machine, event, sample); + break; default: ret =3D machine__process_event(machine, event, sample); break; @@ -4043,7 +4061,7 @@ static int trace__add_syscall_newtp(struct trace *tra= ce) =20 static int trace__set_ev_qualifier_tp_filter(struct trace *trace) { - int err =3D -1; + int err =3D 0; struct evsel *sys_exit; char *filter =3D asprintf_expr_inout_ints("id", !trace->not_ev_qualifier, trace->ev_qualifier_ids.nr, @@ -4052,10 +4070,17 @@ static int trace__set_ev_qualifier_tp_filter(struct= trace *trace) if (filter =3D=3D NULL) goto out_enomem; =20 - if (!evsel__append_tp_filter(trace->syscalls.events.sys_enter, filter)) { - sys_exit =3D trace->syscalls.events.sys_exit; + /* + * With BPF augmentation sys_enter is filtered in BPF and removed from + * the evlist, so only apply the tracepoint filter to the events that + * are actually present. + */ + if (trace->syscalls.events.sys_enter) + err =3D evsel__append_tp_filter(trace->syscalls.events.sys_enter, filter= ); + + sys_exit =3D trace->syscalls.events.sys_exit; + if (!err && sys_exit) err =3D evsel__append_tp_filter(sys_exit, filter); - } =20 free(filter); out: @@ -4502,7 +4527,27 @@ static int trace__init_syscalls_bpf_prog_array_maps(= struct trace *trace __maybe_ =20 static int trace__set_ev_qualifier_filter(struct trace *trace) { - if (trace->syscalls.events.sys_enter) + /* + * Synchronize syscall filter with BPF augmenter map: + * Pass trace->not_ev_qualifier to indicate blacklist mode ('!' prefix, + * e.g., -e !open,close) vs whitelist mode (-e open,close). + * + * A failure here would leave the BPF program filtering on a partially + * populated map, silently dropping or emitting the wrong syscalls, so + * propagate the error rather than continuing. + */ + if (trace->ev_qualifier_ids.nr > 0) { + int err =3D augmented_syscalls__set_target_syscalls(trace->ev_qualifier_= ids.nr, + trace->ev_qualifier_ids.entries, + trace->not_ev_qualifier); + + if (err) { + pr_err("Failed to set the syscalls to trace in the BPF map: %d\n", err); + return err; + } + } + + if (trace->syscalls.events.sys_enter || trace->syscalls.events.sys_exit) return trace__set_ev_qualifier_tp_filter(trace); return 0; } @@ -4543,13 +4588,21 @@ static int trace__set_filter_loop_pids(struct trace= *trace) =20 static int trace__set_filter_pids(struct trace *trace) { - int err =3D 0; + struct perf_thread_map *threads =3D evlist__core(trace->evlist)->threads; /* * Better not use !target__has_task() here because we need to cover the * case where no threads were specified in the command line, but a * workload was, and in that case we will fill in the thread_map when * we fork the workload in evlist__prepare_workload. */ + bool has_target =3D perf_thread_map__pid(threads, 0) !=3D -1; + int err =3D 0; + + /* + * The exclusion list: --filter-pids names tasks to never report, and + * with no target at all we instead exclude perf itself so that tracing + * does not feed back into itself. + */ if (trace->filter_pids.nr > 0) { err =3D evlist__append_tp_filter_pids(trace->evlist, trace->filter_pids.= nr, trace->filter_pids.entries); @@ -4557,10 +4610,37 @@ static int trace__set_filter_pids(struct trace *tra= ce) err =3D augmented_syscalls__set_filter_pids(trace->filter_pids.nr, trace->filter_pids.entries); } - } else if (perf_thread_map__pid(evlist__core(trace->evlist)->threads, 0) = =3D=3D -1) { + } else if (!has_target) { err =3D trace__set_filter_loop_pids(trace); } =20 + if (err) + return err; + + /* + * The inclusion list, which is a separate axis from the exclusion list + * above and so must be programmed even when --filter-pids was given: + * naming tasks to leave out does not widen -p/-t or a workload to the + * whole system. + * + * This matters more than it does on the tracepoint only path. A BPF + * tracepoint program is attached system wide rather than to the + * target's file descriptors, so pids_to_trace is the only thing + * keeping other tasks out. + */ + if (has_target) { + int nr =3D perf_thread_map__nr(threads); + pid_t *pids =3D malloc(nr * sizeof(pid_t)); + + if (pids =3D=3D NULL) + return -ENOMEM; + + for (int i =3D 0; i < nr; i++) + pids[i] =3D perf_thread_map__pid(threads, i); + err =3D augmented_syscalls__set_target_pids(nr, pids); + free(pids); + } + return err; } =20 @@ -4787,7 +4867,8 @@ static int trace__run(struct trace *trace, int argc, = const char **argv) } =20 if (!trace->raw_augmented_syscalls) { - if (trace->trace_syscalls && trace__add_syscall_newtp(trace)) + if (trace->trace_syscalls && !trace->syscalls.events.sys_enter && + trace__add_syscall_newtp(trace)) goto out_error_raw_syscalls; =20 if (trace->trace_syscalls) @@ -5795,6 +5876,7 @@ int cmd_trace(int argc, const char **argv) .show_arg_names =3D true, .args_alignment =3D 70, .trace_syscalls =3D false, + .syscall_augment =3D true, .kernel_syscallchains =3D false, .max_stack =3D UINT_MAX, .max_events =3D ULONG_MAX, @@ -5850,6 +5932,8 @@ int cmd_trace(int argc, const char **argv) OPT_CALLBACK_DEFAULT('F', "pf", &trace.trace_pgfaults, "all|maj|min", "Trace pagefaults", parse_pagefaults, "maj"), OPT_BOOLEAN(0, "syscalls", &trace.trace_syscalls, "Trace syscalls"), + OPT_BOOLEAN(0, "syscall-augment", &trace.syscall_augment, + "Augment syscalls with BPF"), OPT_BOOLEAN('f', "force", &trace.force, "don't complain, do it"), OPT_CALLBACK(0, "call-graph", &trace.opts, "record_mode[,record_size]", record_callchain_help, @@ -5972,7 +6056,7 @@ int cmd_trace(int argc, const char **argv) "cgroup monitoring only available in system-wide mode"); } =20 - if (!trace.trace_syscalls) + if (!trace.trace_syscalls || !trace.syscall_augment) goto skip_augmentation; =20 if ((argc >=3D 1) && (strcmp(argv[0], "record") =3D=3D 0)) { @@ -5997,8 +6081,19 @@ int cmd_trace(int argc, const char **argv) trace__add_syscall_newtp(&trace); =20 err =3D augmented_syscalls__create_bpf_output(trace.evlist); - if (err =3D=3D 0) + if (err =3D=3D 0) { trace.syscalls.events.bpf_output =3D evlist__last(trace.evlist); + } else { + /* + * augmented_syscalls__prepare() already attached sys_enter and + * sys_exit, which are system wide. Falling through to + * skip_augmentation without undoing that would run a BPF + * program for every syscall on the machine, for the whole + * session, with nothing consuming the output. + */ + pr_debug("Failed to create the augmented syscalls bpf-output event, disa= bling augmentation\n"); + augmented_syscalls__cleanup(); + } =20 skip_augmentation: err =3D -1; @@ -6054,7 +6149,9 @@ int cmd_trace(int argc, const char **argv) * syscall. */ if (trace.syscalls.events.bpf_output) { - evlist__for_each_entry(trace.evlist, evsel) { + struct evsel *n; + + evlist__for_each_entry_safe(trace.evlist, n, evsel) { bool raw_syscalls_sys_exit =3D evsel__name_is(evsel, "raw_syscalls:sys_= exit"); =20 if (raw_syscalls_sys_exit) { @@ -6069,21 +6166,26 @@ int cmd_trace(int argc, const char **argv) evsel__init_augmented_syscall_tp_args(augmented)) goto out; /* - * Augmented is __augmented_syscalls__ BPF_OUTPUT event + * Augmented is __augmented_syscalls__ BPF_OUTPUT event. * Above we made sure we can get from the payload the tp fields * that we get from syscalls:sys_enter tracefs format file. + * Since BPF outputs all enter events (both augmented and + * unaugmented) into __augmented_syscalls__, we remove the raw + * sys_enter evsel from evlist so that perf trace only listens + * to __augmented_syscalls__, avoiding duplicate events and + * avoiding kernel tracepoint vetoes. + * + * Because evlist__remove() removes the first evsel (which had + * tracking=3Dtrue by default), re-designate the tracking event + * so PERF_RECORD_COMM and fork tracking continue to be enabled. */ augmented->handler =3D trace__sys_enter; - /* - * Now we do the same for the *syscalls:sys_enter event so that - * if we handle it directly, i.e. if the BPF prog returns 0 so - * as not to filter it, then we'll handle it just like we would - * for the BPF_OUTPUT one: - */ - if (evsel__init_augmented_syscall_tp(evsel, evsel) || - evsel__init_augmented_syscall_tp_args(evsel)) - goto out; - evsel->handler =3D trace__sys_enter; + evlist__remove(trace.evlist, evsel); + evsel__put_and_free_priv(evsel); + trace.syscalls.events.sys_enter =3D NULL; + evlist__set_tracking_event(trace.evlist, + trace.syscalls.events.sys_exit ?: augmented); + continue; } =20 if (strstarts(evsel__name(evsel), "syscalls:sys_exit_")) { diff --git a/tools/perf/util/bpf_skel/augmented_raw_syscalls.bpf.c b/tools/= perf/util/bpf_skel/augmented_raw_syscalls.bpf.c index 3bc9e28a9b8a..6ca9507ecc02 100644 --- a/tools/perf/util/bpf_skel/augmented_raw_syscalls.bpf.c +++ b/tools/perf/util/bpf_skel/augmented_raw_syscalls.bpf.c @@ -114,6 +114,41 @@ struct pids_filtered { __uint(max_entries, 64); } pids_filtered SEC(".maps"); =20 +/* + * Optional hash map containing specific PIDs/TGIDs to trace (e.g., when + * attached to a process with -p or tracing a specific command workload). + * + * has_pids_to_trace: Set to true if target PID filtering is active. + * When false, all processes are eligible for tracing. + */ +struct pids_to_trace { + __uint(type, BPF_MAP_TYPE_HASH); + __type(key, pid_t); + __type(value, bool); + __uint(max_entries, 1024); +} pids_to_trace SEC(".maps"); + +bool has_pids_to_trace; + +/* + * Hash map storing syscall IDs for filtering (via 'perf trace -e ...'). + * + * has_syscalls_to_trace: Set to true if any syscall filter is active. + * not_syscalls_to_trace: Inverts matching when '!' prefix is used in -e + * (e.g., -e !open,close means trace everything EXC= EPT + * open and close; an exclusion blacklist rather th= an + * an inclusion whitelist). + */ +struct syscalls_to_trace { + __uint(type, BPF_MAP_TYPE_HASH); + __type(key, int); + __type(value, bool); + __uint(max_entries, 1024); +} syscalls_to_trace SEC(".maps"); + +bool has_syscalls_to_trace; +bool not_syscalls_to_trace; + struct augmented_args_payload { struct syscall_enter_args args; struct augmented_arg arg, arg2; // We have to reserve space for two argum= ents (rename, etc) @@ -154,8 +189,8 @@ static inline struct augmented_args_payload *augmented_= args_payload(void) =20 static inline int augmented__output(void *ctx, struct augmented_args_paylo= ad *args, int len) { - /* If perf_event_output fails, return non-zero so that it gets recorded u= naugmented */ - return bpf_perf_event_output(ctx, &__augmented_syscalls__, BPF_F_CURRENT_= CPU, args, len); + bpf_perf_event_output(ctx, &__augmented_syscalls__, BPF_F_CURRENT_CPU, ar= gs, len); + return 1; } =20 static inline int augmented__beauty_output(void *ctx, void *data, int len) @@ -191,10 +226,21 @@ unsigned int augmented_arg__read_str(struct augmented= _arg *augmented_arg, const return augmented_len; } =20 -SEC("tp/raw_syscalls/sys_enter") +/* + * Default sys_enter program for syscalls without pointer argument augment= ation. + * Writes the raw struct syscall_enter_args payload into __augmented_sysca= lls__ + * and returns 1 so the tracepoint is never vetoed in the kernel. + */ +SEC("tp/syscalls/sys_enter_unaugmented") int syscall_unaugmented(struct syscall_enter_args *args) { - return 1; + struct augmented_args_payload *augmented_args =3D augmented_args_payload(= ); + + if (augmented_args =3D=3D NULL) + return 1; + + bpf_probe_read_kernel(&augmented_args->args, sizeof(augmented_args->args)= , args); + return augmented__output(args, augmented_args, sizeof(augmented_args->arg= s)); } =20 /* @@ -424,11 +470,41 @@ static pid_t getpid(void) return bpf_get_current_pid_tgid(); } =20 +/* + * Returns true if a PID is explicitly excluded/filtered out (e.g., via --= filter-pids). + */ static bool pid_filter__has(struct pids_filtered *pids, pid_t pid) { return bpf_map_lookup_elem(pids, &pid) !=3D NULL; } =20 +/* + * Checks if the current task (thread PID or process TGID) is targeted for= tracing. + * Checks both PID (thread ID) and TGID (process ID) so that all threads o= f a + * target process match. + */ +static inline bool pid_to_trace__has(pid_t pid) +{ + pid_t tgid =3D bpf_get_current_pid_tgid() >> 32; + + return bpf_map_lookup_elem(&pids_to_trace, &pid) !=3D NULL || + bpf_map_lookup_elem(&pids_to_trace, &tgid) !=3D NULL; +} + +/* + * Determines if a syscall should be traced based on the filter map: + * - When not_syscalls_to_trace is true: blacklist mode (trace if NOT in m= ap). + * - When not_syscalls_to_trace is false: whitelist mode (trace ONLY if IN= map). + */ +static inline bool syscall_to_trace__enabled(int id) +{ + bool in_map =3D bpf_map_lookup_elem(&syscalls_to_trace, &id) !=3D NULL; + + if (not_syscalls_to_trace) + return !in_map; + return in_map; +} + u64 ZERO =3D 0; =20 /* @@ -562,6 +638,11 @@ static int augment_sys_enter(void *ctx, struct syscall= _enter_args *args) return augmented__beauty_output(ctx, payload, sizeof(struct syscall_enter= _args) + output); } =20 +/* + * Main raw_syscalls:sys_enter tracepoint handler. + * Always returns 1 so the tracepoint is never vetoed in the kernel for + * other concurrent listeners. Filtered events simply do not output to the= ring buffer. + */ SEC("tp/raw_syscalls/sys_enter") int sys_enter(struct syscall_enter_args *args) { @@ -576,8 +657,11 @@ int sys_enter(struct syscall_enter_args *args) * initial, non-augmented raw_syscalls:sys_enter payload. */ =20 + if (has_pids_to_trace && !pid_to_trace__has(getpid())) + return 1; + if (pid_filter__has(&pids_filtered, getpid())) - return 0; + return 1; =20 augmented_args =3D augmented_args_payload(); if (augmented_args =3D=3D NULL) @@ -585,25 +669,41 @@ int sys_enter(struct syscall_enter_args *args) =20 bpf_probe_read_kernel(&augmented_args->args, sizeof(augmented_args->args)= , args); =20 + if (has_syscalls_to_trace && !syscall_to_trace__enabled(augmented_args->a= rgs.syscall_nr)) + return 1; + /* - * Jump to syscall specific augmenter, even if the default one, - * "!raw_syscalls:unaugmented" that will just return 1 to return the - * unaugmented tracepoint payload. + * Jump to syscall specific augmenter. If augmented, augment_sys_enter() + * outputs the payload to __augmented_syscalls__ and returns 0. + * Return 1 so we never veto the kernel tracepoint for other listeners. */ - if (augment_sys_enter(args, &augmented_args->args)) - bpf_tail_call(args, &syscalls_sys_enter, augmented_args->args.syscall_nr= ); + if (augment_sys_enter(args, &augmented_args->args) =3D=3D 0) + return 1; =20 - // If not found on the PROG_ARRAY syscalls map, then we're filtering it: - return 0; + bpf_tail_call(args, &syscalls_sys_enter, augmented_args->args.syscall_nr); + + /* + * If not found on the PROG_ARRAY syscalls map, return 1 so we + * don't veto the tracepoint event system-wide for other concurrent + * listeners. + */ + return 1; } =20 +/* + * Main raw_syscalls:sys_exit tracepoint handler. + * Always returns 1 so the tracepoint is never vetoed in the kernel. + */ SEC("tp/raw_syscalls/sys_exit") int sys_exit(struct syscall_exit_args *args) { struct syscall_exit_args exit_args; =20 + if (has_pids_to_trace && !pid_to_trace__has(getpid())) + return 1; + if (pid_filter__has(&pids_filtered, getpid())) - return 0; + return 1; =20 bpf_probe_read_kernel(&exit_args, sizeof(exit_args), args); /* @@ -613,9 +713,12 @@ int sys_exit(struct syscall_exit_args *args) */ bpf_tail_call(args, &syscalls_sys_exit, exit_args.syscall_nr); /* - * If not found on the PROG_ARRAY syscalls map, then we're filtering it: + * If not found on the PROG_ARRAY syscalls map, return 1 so we + * don't veto the tracepoint event system-wide for other concurrent + * listeners. perf trace's own evsel filter will discard non-matching + * syscalls. */ - return 0; + return 1; } =20 char _license[] SEC("license") =3D "GPL"; diff --git a/tools/perf/util/bpf_trace_augment.c b/tools/perf/util/bpf_trac= e_augment.c index a9cf2a77ded1..b9d1208e50f2 100644 --- a/tools/perf/util/bpf_trace_augment.c +++ b/tools/perf/util/bpf_trace_augment.c @@ -1,4 +1,5 @@ #include +#include #include =20 #include "bpf_skel/augmented_raw_syscalls.skel.h" @@ -10,6 +11,23 @@ static struct augmented_raw_syscalls_bpf *skel; static struct evsel *bpf_output; =20 +/* Set by attach_prog() so the first failure is what gets reported. */ +static int attach_err; + +static int attach_prog(struct bpf_link **link, struct bpf_program *prog, c= onst char *name) +{ + *link =3D bpf_program__attach(prog); + if (*link) + return 0; + /* + * Save errno before pr_debug(), which formats and writes output and so + * can overwrite it. + */ + attach_err =3D -errno; + pr_debug("Failed to attach %s BPF program\n", name); + return attach_err; +} + int augmented_syscalls__prepare(void) { struct bpf_program *prog; @@ -35,11 +53,35 @@ int augmented_syscalls__prepare(void) if (err < 0) { libbpf_strerror(err, buf, sizeof(buf)); pr_debug("Failed to load augmented syscalls BPF skeleton: %s\n", buf); + /* + * Tear the skeleton down rather than leaving a half initialized + * one behind. The caller falls back to unaugmented tracing and + * still calls the setters below, which must then do nothing + * instead of failing against a skeleton with no maps. + */ + augmented_syscalls__cleanup(); return err; } =20 - augmented_raw_syscalls_bpf__attach(skel); + /* + * Only sys_enter and sys_exit are attached, the remaining programs are + * reached by tail calls. Attach them explicitly and, on failure, undo + * any partial attachment: leaving sys_enter live on + * raw_syscalls:sys_enter would keep running a BPF program for every + * syscall on the system for a perf trace session that never starts. + */ + if (attach_prog(&skel->links.sys_enter, skel->progs.sys_enter, "sys_enter= ")) + goto out_cleanup; + if (attach_prog(&skel->links.sys_exit, skel->progs.sys_exit, "sys_exit")) + goto out_cleanup; + return 0; + +out_cleanup: + err =3D attach_err; + /* Destroys every link attached above along with the skeleton. */ + augmented_syscalls__cleanup(); + return err; } =20 int augmented_syscalls__create_bpf_output(struct evlist *evlist) @@ -98,6 +140,96 @@ int augmented_syscalls__set_filter_pids(unsigned int nr= , pid_t *pids) return err; } =20 +/* + * Populate target PIDs in the BPF pids_to_trace map (e.g., for -p or + * when tracing a specified command workload). + */ +int augmented_syscalls__set_target_pids(unsigned int nr, pid_t *pids) +{ + bool value =3D true; + int err =3D 0; + + if (skel =3D=3D NULL || nr =3D=3D 0) + return 0; + + for (size_t i =3D 0; i < nr; ++i) { + err =3D bpf_map__update_elem(skel->maps.pids_to_trace, &pids[i], + sizeof(*pids), &value, sizeof(value), + BPF_ANY); + if (err) + return err; + } + /* + * Set the flag only once every target is in the map. The BPF programs + * are attached by this point, so flipping it first would have them + * filter against a partially populated map and drop syscalls made by + * the targets that had not been added yet. + */ + skel->bss->has_pids_to_trace =3D true; + return 0; +} + +int augmented_syscalls__add_target_pid(pid_t pid) +{ + bool value =3D true; + + if (skel =3D=3D NULL || !skel->bss->has_pids_to_trace || skel->maps.pids_= to_trace =3D=3D NULL) + return 0; + + return bpf_map__update_elem(skel->maps.pids_to_trace, &pid, sizeof(pid), + &value, sizeof(value), BPF_ANY); +} + +int augmented_syscalls__del_target_pid(pid_t pid) +{ + if (skel =3D=3D NULL || !skel->bss->has_pids_to_trace || skel->maps.pids_= to_trace =3D=3D NULL) + return 0; + + return bpf_map__delete_elem(skel->maps.pids_to_trace, &pid, sizeof(pid), = 0); +} + +bool augmented_syscalls__has_target_pid(pid_t pid) +{ + bool value; + + if (skel =3D=3D NULL || !skel->bss->has_pids_to_trace || skel->maps.pids_= to_trace =3D=3D NULL) + return false; + + return bpf_map__lookup_elem(skel->maps.pids_to_trace, &pid, sizeof(pid), + &value, sizeof(value), 0) =3D=3D 0; +} + +/* + * Populate syscalls in the BPF syscalls_to_trace map: + * - not_syscalls: true if '!' prefix was specified (blacklist mode: trace + * all syscalls EXCEPT these). + * false if whitelist mode (trace ONLY these syscalls). + */ +int augmented_syscalls__set_target_syscalls(unsigned int nr, int *syscall_= ids, bool not_syscalls) +{ + bool value =3D true; + int err =3D 0; + + if (skel =3D=3D NULL || nr =3D=3D 0) + return 0; + + skel->bss->not_syscalls_to_trace =3D not_syscalls; + for (size_t i =3D 0; i < nr; ++i) { + err =3D bpf_map__update_elem(skel->maps.syscalls_to_trace, &syscall_ids[= i], + sizeof(int), &value, sizeof(value), + BPF_ANY); + if (err) + return err; + } + /* + * As for the pid maps, publish the filter only once it is complete: + * in whitelist mode a half filled map would drop syscalls that were + * asked for but not added yet. + */ + skel->bss->has_syscalls_to_trace =3D true; + return 0; +} + int augmented_syscalls__get_map_fds(int *enter_fd, int *exit_fd, int *beau= ty_fd) { if (skel =3D=3D NULL) @@ -140,4 +272,5 @@ struct bpf_program *augmented_syscalls__find_by_title(c= onst char *name) void augmented_syscalls__cleanup(void) { augmented_raw_syscalls_bpf__destroy(skel); + skel =3D NULL; } diff --git a/tools/perf/util/trace_augment.h b/tools/perf/util/trace_augmen= t.h index 4f729bc67753..56f4dbab7d4b 100644 --- a/tools/perf/util/trace_augment.h +++ b/tools/perf/util/trace_augment.h @@ -12,6 +12,11 @@ int augmented_syscalls__prepare(void); int augmented_syscalls__create_bpf_output(struct evlist *evlist); void augmented_syscalls__setup_bpf_output(void); int augmented_syscalls__set_filter_pids(unsigned int nr, pid_t *pids); +int augmented_syscalls__set_target_pids(unsigned int nr, pid_t *pids); +int augmented_syscalls__add_target_pid(pid_t pid); +int augmented_syscalls__del_target_pid(pid_t pid); +bool augmented_syscalls__has_target_pid(pid_t pid); +int augmented_syscalls__set_target_syscalls(unsigned int nr, int *syscall_= ids, bool not_syscalls); int augmented_syscalls__get_map_fds(int *enter_fd, int *exit_fd, int *beau= ty_fd); struct bpf_program *augmented_syscalls__find_by_title(const char *name); struct bpf_program *augmented_syscalls__unaugmented(void); @@ -39,6 +44,34 @@ static inline int augmented_syscalls__set_filter_pids(un= signed int nr __maybe_un return 0; } =20 +static inline int augmented_syscalls__set_target_pids(unsigned int nr __ma= ybe_unused, + pid_t *pids __maybe_unused) +{ + return 0; +} + +static inline int augmented_syscalls__add_target_pid(pid_t pid __maybe_unu= sed) +{ + return 0; +} + +static inline int augmented_syscalls__del_target_pid(pid_t pid __maybe_unu= sed) +{ + return 0; +} + +static inline bool augmented_syscalls__has_target_pid(pid_t pid __maybe_un= used) +{ + return false; +} + +static inline int augmented_syscalls__set_target_syscalls(unsigned int nr = __maybe_unused, + int *syscall_ids __maybe_unused, + bool not_syscalls __maybe_unused) +{ + return 0; +} + static inline int augmented_syscalls__get_map_fds(int *enter_fd __maybe_un= used, int *exit_fd __maybe_unused, int *beauty_fd __maybe_unused) --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Fri Sep 25 03:16:03 2026 Received: from mail-pl1-f198.google.com (mail-pl1-f198.google.com [209.85.214.198]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B909343CE5E for ; Thu, 17 Sep 2026 06:42:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.198 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789627381; cv=none; b=iUKJMsk5g2+G51KG8YJO2VGAG/X4a3syCzSUJ58EAThavQAJl0cU+B1j9pWlZXtTFRlvzXTl/ht9/9yrpZbJKmlf9l3z4psTFmqxdIkkLnnOR/C2i9Ml0lDS6xvyMGDal3MImlJLxIbWvSOJ5MDJbjh+i+V3gW1MSYrI3471G3g= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789627381; c=relaxed/simple; bh=pCYthVdyoV3UBP/kkzGKJxv/qNQW2xgcg594NKZqa5c=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=sSVGWLDYProDMkM7yJW/k1J+P/Wixo0kXv5MqMZP6iMxFxXvY4Yod75oBuZeYr3FvTDRhTII8zL/9xVmYm+hREqA7uY0jnaG9UPWIUccDzcgupX/PnaIQKfyALr1I+as0gIdPaGf4euKGUI2fw1UK2T1uJdhihFPIUm/f96y3JM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--irogers.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=GN01iobb; arc=none smtp.client-ip=209.85.214.198 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--irogers.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="GN01iobb" Received: by mail-pl1-f198.google.com with SMTP id d9443c01a7336-2dd53f2b27cso8785765ad.0 for ; Wed, 16 Sep 2026 23:42:58 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789627378; x=1790232178; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=Ejyoj3BGFDS1MG1F9xsynJcx/QiWcHBrmC+P3sRSEnQ=; b=GN01iobbRWYrBUnZycaD7KWbKv7F9hVD1TjNyzXPk3gul2rCmrM3mCB4FjcmXZbWck LISFBMtHlEQhxC7ALMYyKvhEhtQZfzcFBox63Dm8AmExMH3ZaXttEwzgxrD9SqPDryyL MCaYFm+3uNx1vXdknZvxosA3+tap4vbWuWczwQr9Arw3nKERegFYxRuT4xUZ2xKSEia+ vYjrMGjJBd00HZChIGH7XKR6TolKvLVSwR5kSbzH1eaJhbk3NwpLOxSsWWGIjTTIUe+g tkE29FlGxInmml5GoKnZ+vlqLgAIlK3ZChhhuaexlnB9vdeYqVdlT5sTAwfcrOD7cYJf KNPg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789627378; x=1790232178; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=Ejyoj3BGFDS1MG1F9xsynJcx/QiWcHBrmC+P3sRSEnQ=; b=LYS1VJVg2rOgGrRxRRe1xIQsjWjXbvAKIGScLH8tCdBekLL+7K0irDGfuk3OlOOhAy Xy90mA4xnPGWmkytmWattPXrwsxMh0BQXLQWcv5ZQRjAnv47My4Y+k756MwNLQo7J8BE sVqW0meUojXfxG4nzgIWEfN0DT7P2vVctYvCTNbdzW37QZSFZEOBzAes7FHPHEjyU80n hfXizgI5+t/PtZo8GGasG2SfqTEIbkkSKnG5x+wmXzUJOUOl4eXmOha6VNNl8T8yzv9W E0MhkalzzBFY1RQ1O4RDlfxDRCXQCohuD9EkecNAx/IPRWEMVHtuig8eBg2mQHNkzzZV oyLQ== X-Forwarded-Encrypted: i=1; AKwUvBz4EZRfFOzt8ifBSZuMhLLhp7t+ZDJRgM6am00tz/zygwPNZxtZ7RZ36NKiXGbrpkgB5Xm6ZXVQK0uNSAw=@vger.kernel.org X-Gm-Message-State: AFuF++mqxr3Bs73mjSK+6KD/c+M+xxf0ltbXbBkDSYpSvymyK/mM/D6a y1qn4OP9rZDXeDBs6DdkH/8YF7JQamUxKnj4n/CccCybvI/02HQKVZXWFfOFssHF4gjgf6U3MEn Mw8Yb00j7oQ== X-Received: from dldnz11.prod.google.com ([2002:a05:701a:ca0b:b0:143:8cae:a783]) (user=irogers job=prod-delivery.src-stubby-dispatcher) by 2002:a17:90a:10c1:b0:39e:1ec9:dc8d with SMTP id 98e67ed59e1d1-39e1ec9e479mr8502759a91.24.1789627377782; Wed, 16 Sep 2026 23:42:57 -0700 (PDT) Date: Wed, 16 Sep 2026 23:42:27 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <301baf74713ea1b7ea6bce83d83edd8820b6f1d6.1789626978.git.irogers@google.com> Subject: [PATCH v1 05/13] perf trace: Handle fork and exit directly in BPF filter maps From: Ian Rogers To: Arnaldo Carvalho de Melo , Namhyung Kim Cc: Peter Zijlstra , Ingo Molnar , Jiri Olsa , Adrian Hunter , James Clark , linux-perf-users@vger.kernel.org, linux-kernel@vger.kernel.org, Ian Rogers Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Updating target or filtered PIDs in userspace upon processing PERF_RECORD_FORK and PERF_RECORD_EXIT events introduces latency between event occurrence and userspace BPF map updates. If a newly forked child executes system calls before userspace processes PERF_RECORD_FORK, those syscalls may be dropped by BPF PID filtering. Conversely, if userspace evicts PIDs asynchronously on PERF_RECORD_EXIT, the kernel may recycle a PID before userspace processes the exit event, causing the late eviction to silently drop a newly created task that received the recycled PID. Address this by attaching BTF-typed raw tracepoint BPF programs directly to the scheduler task lifetime tracepoints: 1. Attach SEC("tp_btf/sched_process_fork") (sched_process_fork), which runs in copy_process() in the parent's context before wake_up_new_task() wakes the child. Using tp_btf rather than SEC("tp/sched/sched_process_fork") receives the stable TP_PROTO arguments (struct task_struct *parent, struct task_struct *child) rather than the tracepoint ring-buffer record (TP_STRUCT__entry), whose layout changed in Linux 6.16 when parent_comm and child_comm were converted from 16-byte arrays to 4-byte __data_loc strings. When inherit is enabled and the parent's PID or TGID is in pids_to_trace or pids_filtered, insert child->pid into the corresponding map immediately. Because child->pid is task_struct.pid (the global initial-namespace PID), this works accurately across PID namespaces without aliasing host PIDs, and covers both new processes and CLONE_THREAD threads without needing real_parent CO-RE walks or syscall-return heuristics. 2. Attach SEC("tp_btf/sched_process_exit") (sched_process_exit), which runs in do_exit() for every task in its own context, including tasks killed by signals (SIGKILL, SIGSEGV, etc.) and secondary threads torn down implicitly by exit_group. Delete the dying task's PID from pids_to_trace and pids_filtered immediately in kernel space, eliminating both map leaks and any asynchronous userspace eviction window where PID recycling could occur. 3. Attach SEC("tp_btf/sched_process_exec") (sched_process_exec) to follow the one case where a live task's pid changes underneath the maps. When a thread that is not the group leader execs, de_thread() kills the leader and hands the leader's pid, which is the tgid, to the exec'ing thread. The leader dies first, so sched_process_exit() has already dropped exactly the pid the survivor now holds, and the survivor's old entry would be stranded in the map for good. Move the entry from old_pid to p->pid. old_pid is sampled in bprm_execve() before de_thread() runs, so the ordinary group leader exec is a no-op here. 4. With every live task registered before its first syscall and evicted in do_exit(), simplify pid_to_trace__has() and pid_filter__has() to single BPF hash map lookups, and move bpf_probe_read_kernel() in sys_exit back after the PID filter checks. 5. Pass the inherit flag from userspace to BPF .rodata via augmented_syscalls__prepare(!trace.opts.no_inherit), and split attaching out of it into augmented_syscalls__attach(), called from trace__run() once the pid, syscall and program array maps have all been programmed. These are system wide programs, so from the instant they attach they alone decide what is traced: attaching at load time, as before, left a window in which a target could fork without sched_process_fork() knowing the parent was a target, and with the userspace fork handling gone there was nothing left to recover it. The scheduler programs are attached ahead of sys_enter and sys_exit for the same reason. Set has_pids_filtered only after populating pids_filtered. 6. Remove the userspace BPF map updates from PERF_RECORD_FORK and PERF_RECORD_EXIT in trace__process_event(), and delete the now-unused augmented_syscalls__{add,del,has}_target_pid() helpers. No coverage is lost with them: those records only come into being once the ring buffers are mapped by evlist__do_mmap() and the events are switched on by evlist__enable(), both of which run after augmented_syscalls__attach() in trace__run(), and they are then acted on later still, whenever the poll loop gets round to them. The scheduler programs therefore go live strictly earlier than the userspace path could ever have reacted. 7. Gate pid_filter__has() on a has_pids_filtered flag in .bss so the common case without --filter-pids performs no map lookups, and size pids_to_trace and pids_filtered at 16384 entries. pids_filtered is grown from 64 because it is no longer just the handful of pids userspace names: sched_process_fork() adds every descendant of those, so a --filter-pids target that forks or is heavily threaded needs the same headroom as a traced one. A fork or exit is still not seen if it happens before the programs are attached, that is between evlist__create_maps() scanning /proc for a -p target and augmented_syscalls__attach(). Such a window is inherent in programming a system wide filter before switching it on, and as above the userspace handling did not cover it either. What it costs is small: a child forked in the window is still traced through the tracepoints its parent's events were inherited by, only unaugmented, because sys_enter returns 1 for a pid that is not in the map rather than vetoing the tracepoint. A workload started by 'perf trace -- cmd' cannot hit it at all, as evlist__prepare_workload() leaves the child blocked on a pipe until evlist__start_workload(), well after the attach. Assisted-by: Antigravity:gemini-3.1-pro Signed-off-by: Ian Rogers --- tools/perf/builtin-trace.c | 39 ++-- .../bpf_skel/augmented_raw_syscalls.bpf.c | 178 +++++++++++++++++- tools/perf/util/bpf_trace_augment.c | 116 +++++++----- tools/perf/util/trace_augment.h | 28 +-- 4 files changed, 262 insertions(+), 99 deletions(-) diff --git a/tools/perf/builtin-trace.c b/tools/perf/builtin-trace.c index e21b2b4a8794..a30fe273f452 100644 --- a/tools/perf/builtin-trace.c +++ b/tools/perf/builtin-trace.c @@ -2055,23 +2055,6 @@ static int trace__process_event(struct trace *trace,= struct machine *machine, "LOST %" PRIu64 " events!\n", (u64)event->lost.lost); ret =3D machine__process_lost_event(machine, event, sample); break; - case PERF_RECORD_FORK: - if (trace->raw_augmented_syscalls && - (augmented_syscalls__has_target_pid(event->fork.ppid) || - augmented_syscalls__has_target_pid(event->fork.ptid))) { - augmented_syscalls__add_target_pid(event->fork.pid); - } - ret =3D machine__process_fork_event(machine, event, sample); - break; - case PERF_RECORD_EXIT: - if (trace->raw_augmented_syscalls) { - if (event->fork.pid =3D=3D event->fork.tid) - augmented_syscalls__del_target_pid(event->fork.pid); - else - augmented_syscalls__del_target_pid(event->fork.tid); - } - ret =3D machine__process_exit_event(machine, event, sample); - break; default: ret =3D machine__process_event(machine, event, sample); break; @@ -4982,6 +4965,16 @@ static int trace__run(struct trace *trace, int argc,= const char **argv) } } =20 + /* + * Everything the BPF programs filter on is now in their maps, so it is + * safe to let them run. They are attached system wide, so anything + * before this point would have been filtered against a map that was + * still being built up. + */ + err =3D augmented_syscalls__attach(); + if (err < 0) + goto out_errno; + /* * If the "close" syscall is not traced, then we will not have the * opportunity to, in syscall_arg__scnprintf_close_fd() invalidate the @@ -6074,7 +6067,7 @@ int cmd_trace(int argc, const char **argv) goto skip_augmentation; } =20 - err =3D augmented_syscalls__prepare(); + err =3D augmented_syscalls__prepare(!trace.opts.no_inherit); if (err < 0) goto skip_augmentation; =20 @@ -6085,11 +6078,11 @@ int cmd_trace(int argc, const char **argv) trace.syscalls.events.bpf_output =3D evlist__last(trace.evlist); } else { /* - * augmented_syscalls__prepare() already attached sys_enter and - * sys_exit, which are system wide. Falling through to - * skip_augmentation without undoing that would run a BPF - * program for every syscall on the machine, for the whole - * session, with nothing consuming the output. + * Drop the loaded skeleton before falling back to unaugmented + * tracing. Otherwise the setters called from trace__run() would + * still program its maps, and augmented_syscalls__attach() would + * then put system wide BPF programs on raw_syscalls for a + * session with nothing consuming their output. */ pr_debug("Failed to create the augmented syscalls bpf-output event, disa= bling augmentation\n"); augmented_syscalls__cleanup(); diff --git a/tools/perf/util/bpf_skel/augmented_raw_syscalls.bpf.c b/tools/= perf/util/bpf_skel/augmented_raw_syscalls.bpf.c index 6ca9507ecc02..7124ed3c39c8 100644 --- a/tools/perf/util/bpf_skel/augmented_raw_syscalls.bpf.c +++ b/tools/perf/util/bpf_skel/augmented_raw_syscalls.bpf.c @@ -9,6 +9,7 @@ #include "vmlinux.h" =20 #include +#include #include =20 #define PERF_ALIGN(x, a) __PERF_ALIGN_MASK(x, (typeof(x))(a)-1) @@ -107,25 +108,48 @@ struct augmented_arg { }; }; =20 +/* + * Hash map of PIDs/TGIDs whose events must be discarded, e.g. perf trace'= s own + * pid, so that tracing doesn't feed back on itself. + * + * has_pids_filtered: set to true only when the map is populated. Checking= a + * boolean is much cheaper than a map lookup, and sys_e= nter + * runs for every syscall on the system, so the common + * "no pids filtered" case must stay on a fast path. + * + * max_entries matches pids_to_trace: userspace only ever names a handful = of + * pids here, but sched_process_fork() below adds every descendant of thos= e, + * so a --filter-pids target that forks or is heavily threaded needs the s= ame + * headroom as a traced one. + */ struct pids_filtered { __uint(type, BPF_MAP_TYPE_HASH); __type(key, pid_t); __type(value, bool); - __uint(max_entries, 64); + __uint(max_entries, 16384); } pids_filtered SEC(".maps"); =20 +bool has_pids_filtered; + /* * Optional hash map containing specific PIDs/TGIDs to trace (e.g., when * attached to a process with -p or tracing a specific command workload). * * has_pids_to_trace: Set to true if target PID filtering is active. * When false, all processes are eligible for tracing. + * + * max_entries bounds how many tasks can be tracked at once. sched_process= _exit + * below evicts a task as it dies, whatever it died of, so the map holds l= ive + * tasks rather than growing without bound. It is sized well + * above the thread count of realistic traced workloads; should a workload + * still exceed it, bpf_map_update_elem() fails with -E2BIG and the extra + * tasks are simply not traced rather than anything being corrupted. */ struct pids_to_trace { __uint(type, BPF_MAP_TYPE_HASH); __type(key, pid_t); __type(value, bool); - __uint(max_entries, 1024); + __uint(max_entries, 16384); } pids_to_trace SEC(".maps"); =20 bool has_pids_to_trace; @@ -149,6 +173,9 @@ struct syscalls_to_trace { bool has_syscalls_to_trace; bool not_syscalls_to_trace; =20 +/* Inherit tracing for child tasks (set to false if --no-inherit is specif= ied) */ +const volatile bool inherit =3D true; + struct augmented_args_payload { struct syscall_enter_args args; struct augmented_arg arg, arg2; // We have to reserve space for two argum= ents (rename, etc) @@ -471,24 +498,35 @@ static pid_t getpid(void) } =20 /* - * Returns true if a PID is explicitly excluded/filtered out (e.g., via --= filter-pids). + * Checks if a PID is explicitly excluded/filtered out (e.g., via --filter= -pids). + * + * Children of a filtered task are added to the map by sched_process_fork() + * below, so a plain lookup is all that is needed here. */ static bool pid_filter__has(struct pids_filtered *pids, pid_t pid) { + /* + * Fast path: this runs for every syscall on the system, so when no pid + * is filtered do no work at all rather than failing a lookup. + */ + if (!has_pids_filtered) + return false; + return bpf_map_lookup_elem(pids, &pid) !=3D NULL; } =20 /* - * Checks if the current task (thread PID or process TGID) is targeted for= tracing. - * Checks both PID (thread ID) and TGID (process ID) so that all threads o= f a - * target process match. + * Checks if the current task is targeted for tracing. + * + * Every thread that existed when tracing started was named by the target = and + * inserted from userspace, and every task created since was inserted by + * sched_process_fork() below, before it was able to run. So there is noth= ing + * to derive here, and in particular no need to consult the tgid or walk t= o the + * parent: a task is traced if and only if it is in the map. */ static inline bool pid_to_trace__has(pid_t pid) { - pid_t tgid =3D bpf_get_current_pid_tgid() >> 32; - - return bpf_map_lookup_elem(&pids_to_trace, &pid) !=3D NULL || - bpf_map_lookup_elem(&pids_to_trace, &tgid) !=3D NULL; + return bpf_map_lookup_elem(&pids_to_trace, &pid) !=3D NULL; } =20 /* @@ -706,6 +744,7 @@ int sys_exit(struct syscall_exit_args *args) return 1; =20 bpf_probe_read_kernel(&exit_args, sizeof(exit_args), args); + /* * Jump to syscall specific return augmenter, even if the default one, * "!raw_syscalls:unaugmented" that will just return 1 to return the @@ -721,4 +760,123 @@ int sys_exit(struct syscall_exit_args *args) return 1; } =20 +/* + * Propagate tracing to a newly created task. + * + * tp_btf/sched_process_fork is raised by copy_process(), in the parent's + * context and before the child is woken, so the child is in the maps befo= re it + * can issue its first syscall. That removes the need to inspect real_pare= nt + * when a syscall is seen from an unknown task, which could neither tell a + * genuine descendant from a task merely reparented to a traced init, nor = keep + * following a descendant whose parent had already exited. + * + * Using tp_btf rather than tp/sched/sched_process_fork avoids depending o= n the + * tracepoint ring-buffer record layout (TP_STRUCT__entry), which changed = in + * Linux 6.16 when parent_comm and child_comm were converted from fixed 16= -byte + * arrays to 4-byte __data_loc strings (shrinking the tracepoint context f= rom + * 48 to 24 bytes and causing BPF_PROG_TYPE_TRACEPOINT attachment to fail = with + * -EACCES when accessing higher offsets). Instead, tp_btf receives the st= able + * TP_PROTO arguments (struct task_struct *parent, struct task_struct *chi= ld) + * directly. + * + * child->pid is task_struct.pid, i.e. the pid in the initial namespace, w= hich + * is what the maps are keyed by. A clone() return value, in contrast, is = the + * pid in the caller's namespace and would alias an unrelated host task wh= en a + * containerised workload is traced. + * + * CLONE_THREAD needs no special handling: a new thread arrives here like = any + * other task and is inserted under its own pid. + */ +SEC("tp_btf/sched_process_fork") +int BPF_PROG(sched_process_fork, struct task_struct *parent, struct task_s= truct *child) +{ + pid_t parent_tgid, parent_pid, child_pid; + bool val =3D true; + + if (!inherit) + return 0; + + /* + * The parent's own pid and tgid: the thread that called clone() may + * itself only be tracked by the pid of its thread group leader. + */ + parent_pid =3D parent->pid; + parent_tgid =3D parent->tgid; + child_pid =3D child->pid; + + if (has_pids_to_trace && + (bpf_map_lookup_elem(&pids_to_trace, &parent_pid) !=3D NULL || + bpf_map_lookup_elem(&pids_to_trace, &parent_tgid) !=3D NULL)) + bpf_map_update_elem(&pids_to_trace, &child_pid, &val, BPF_ANY); + + if (has_pids_filtered && + (bpf_map_lookup_elem(&pids_filtered, &parent_pid) !=3D NULL || + bpf_map_lookup_elem(&pids_filtered, &parent_tgid) !=3D NULL)) + bpf_map_update_elem(&pids_filtered, &child_pid, &val, BPF_ANY); + + return 0; +} + +/* + * Drop a dying task from the maps. + * + * tp_btf/sched_process_exit is raised by do_exit() for every task, in its= own + * context, so unlike hooking the exit and exit_group syscalls this also c= overs + * tasks killed by a signal and threads torn down implicitly by exit_group. + * + * Doing it here rather than from the userspace PERF_RECORD_EXIT handler a= lso + * means there is no window between the task dying and the map being updat= ed, + * during which the kernel could recycle the pid and the late eviction sil= ently + * stop tracing whichever new task received it. + * + * Each thread is reported separately, including the group leader, whose p= id is + * the thread group's tgid, so one delete per map covers both uses of the = key. + */ +SEC("tp_btf/sched_process_exit") +int BPF_PROG(sched_process_exit, struct task_struct *p) +{ + pid_t pid =3D p->pid; + + bpf_map_delete_elem(&pids_to_trace, &pid); + bpf_map_delete_elem(&pids_filtered, &pid); + + return 0; +} + +/* + * Follow a task whose pid changed under it. + * + * When a thread that is not the thread group leader execs, de_thread() ki= lls + * the rest of the group and then hands the leader's pid, which is the tgi= d, to + * the exec'ing thread. The leader dies first, so sched_process_exit() abo= ve + * has already dropped that pid from the maps, and the survivor is now key= ed by + * a pid nothing knows about while its original entry is left behind for g= ood. + * + * Move the entry across so the task stays tracked and nothing is leaked. + * old_pid is sampled in bprm_execve() before de_thread() runs, so for the + * common case of the group leader exec'ing it simply equals p->pid and th= ere + * is nothing to do. + */ +SEC("tp_btf/sched_process_exec") +int BPF_PROG(sched_process_exec, struct task_struct *p, pid_t old_pid) +{ + pid_t pid =3D p->pid; + bool val =3D true; + + if (pid =3D=3D old_pid) + return 0; + + if (bpf_map_lookup_elem(&pids_to_trace, &old_pid) !=3D NULL) { + bpf_map_update_elem(&pids_to_trace, &pid, &val, BPF_ANY); + bpf_map_delete_elem(&pids_to_trace, &old_pid); + } + + if (bpf_map_lookup_elem(&pids_filtered, &old_pid) !=3D NULL) { + bpf_map_update_elem(&pids_filtered, &pid, &val, BPF_ANY); + bpf_map_delete_elem(&pids_filtered, &old_pid); + } + + return 0; +} + char _license[] SEC("license") =3D "GPL"; diff --git a/tools/perf/util/bpf_trace_augment.c b/tools/perf/util/bpf_trac= e_augment.c index b9d1208e50f2..41a867d60b37 100644 --- a/tools/perf/util/bpf_trace_augment.c +++ b/tools/perf/util/bpf_trace_augment.c @@ -28,7 +28,7 @@ static int attach_prog(struct bpf_link **link, struct bpf= _program *prog, const c return attach_err; } =20 -int augmented_syscalls__prepare(void) +int augmented_syscalls__prepare(bool inherit) { struct bpf_program *prog; char buf[128]; @@ -40,12 +40,18 @@ int augmented_syscalls__prepare(void) return -errno; } =20 + skel->rodata->inherit =3D inherit; + /* - * Disable attaching the BPF programs except for sys_enter and - * sys_exit that tail call into this as necessary. + * Disable attaching the BPF programs other than those attached + * explicitly by augmented_syscalls__attach(), the rest are reached by + * tail calls. */ bpf_object__for_each_program(prog, skel->obj) { - if (prog !=3D skel->progs.sys_enter && prog !=3D skel->progs.sys_exit) + if (prog !=3D skel->progs.sys_enter && prog !=3D skel->progs.sys_exit && + prog !=3D skel->progs.sched_process_fork && + prog !=3D skel->progs.sched_process_exit && + prog !=3D skel->progs.sched_process_exec) bpf_program__set_autoattach(prog, /*autoattach=3D*/false); } =20 @@ -63,13 +69,42 @@ int augmented_syscalls__prepare(void) return err; } =20 + return 0; +} + +int augmented_syscalls__attach(void) +{ + int err; + + if (skel =3D=3D NULL) + return 0; + /* - * Only sys_enter and sys_exit are attached, the remaining programs are - * reached by tail calls. Attach them explicitly and, on failure, undo - * any partial attachment: leaving sys_enter live on - * raw_syscalls:sys_enter would keep running a BPF program for every - * syscall on the system for a perf trace session that never starts. + * Attaching is deliberately separate from, and a lot later than, + * loading: these are system wide tracepoint programs, so from the + * moment they are attached they are the only thing deciding which + * tasks and syscalls are traced. Going live before the pid and + * syscall maps are populated would mean a target that forked in the + * meantime was never picked up by sched_process_fork() below. + * + * Attach explicitly, so that a failure part way through can undo what + * came before it: leaving sys_enter live on raw_syscalls:sys_enter + * would keep running a BPF program for every syscall on the system for + * a perf trace session that never starts. + * + * The scheduler programs maintain the pid maps, and are attached first + * so that no fork, exit or exec can be missed between sys_enter going + * live and the maps being maintained. */ + if (attach_prog(&skel->links.sched_process_fork, skel->progs.sched_proces= s_fork, + "sched_process_fork")) + goto out_cleanup; + if (attach_prog(&skel->links.sched_process_exit, skel->progs.sched_proces= s_exit, + "sched_process_exit")) + goto out_cleanup; + if (attach_prog(&skel->links.sched_process_exec, skel->progs.sched_proces= s_exec, + "sched_process_exec")) + goto out_cleanup; if (attach_prog(&skel->links.sys_enter, skel->progs.sys_enter, "sys_enter= ")) goto out_cleanup; if (attach_prog(&skel->links.sys_exit, skel->progs.sys_exit, "sys_exit")) @@ -81,6 +116,13 @@ int augmented_syscalls__prepare(void) err =3D attach_err; /* Destroys every link attached above along with the skeleton. */ augmented_syscalls__cleanup(); + /* + * Tearing the skeleton down closes file descriptors and frees memory, + * either of which may overwrite errno. Restore it so that a caller + * reporting this with "%m" describes the attach failure rather than + * whatever the teardown happened to do last. + */ + errno =3D -err; return err; } =20 @@ -127,17 +169,29 @@ int augmented_syscalls__set_filter_pids(unsigned int = nr, pid_t *pids) bool value =3D true; int err =3D 0; =20 - if (skel =3D=3D NULL) + if (skel =3D=3D NULL || nr =3D=3D 0) return 0; =20 + /* + * Tell the BPF program that the pids_filtered map is in use. Without + * this it would have to look up every task in an empty map, on every + * syscall on the system, to find out that nothing is filtered. + */ for (size_t i =3D 0; i < nr; ++i) { err =3D bpf_map__update_elem(skel->maps.pids_filtered, &pids[i], sizeof(*pids), &value, sizeof(value), BPF_ANY); if (err) - break; + return err; } - return err; + /* + * Publish the filter only now that the map is fully populated. + * augmented_syscalls__attach() has not run yet, so nothing is reading + * either of them, but keeping the flag and the map consistent means + * the ordering stays correct however the callers are rearranged. + */ + skel->bss->has_pids_filtered =3D true; + return 0; } =20 /* @@ -160,45 +214,15 @@ int augmented_syscalls__set_target_pids(unsigned int = nr, pid_t *pids) return err; } /* - * Set the flag only once every target is in the map. The BPF programs - * are attached by this point, so flipping it first would have them - * filter against a partially populated map and drop syscalls made by - * the targets that had not been added yet. + * Set the flag only once every target is in the map, so that the two + * are never inconsistent. Publishing it first would, once the + * programs are attached, have them filter against a partially + * populated map and drop syscalls made by targets not yet added. */ skel->bss->has_pids_to_trace =3D true; return 0; } =20 -int augmented_syscalls__add_target_pid(pid_t pid) -{ - bool value =3D true; - - if (skel =3D=3D NULL || !skel->bss->has_pids_to_trace || skel->maps.pids_= to_trace =3D=3D NULL) - return 0; - - return bpf_map__update_elem(skel->maps.pids_to_trace, &pid, sizeof(pid), - &value, sizeof(value), BPF_ANY); -} - -int augmented_syscalls__del_target_pid(pid_t pid) -{ - if (skel =3D=3D NULL || !skel->bss->has_pids_to_trace || skel->maps.pids_= to_trace =3D=3D NULL) - return 0; - - return bpf_map__delete_elem(skel->maps.pids_to_trace, &pid, sizeof(pid), = 0); -} - -bool augmented_syscalls__has_target_pid(pid_t pid) -{ - bool value; - - if (skel =3D=3D NULL || !skel->bss->has_pids_to_trace || skel->maps.pids_= to_trace =3D=3D NULL) - return false; - - return bpf_map__lookup_elem(skel->maps.pids_to_trace, &pid, sizeof(pid), - &value, sizeof(value), 0) =3D=3D 0; -} - /* * Populate syscalls in the BPF syscalls_to_trace map: * - not_syscalls: true if '!' prefix was specified (blacklist mode: trace diff --git a/tools/perf/util/trace_augment.h b/tools/perf/util/trace_augmen= t.h index 56f4dbab7d4b..f9eecff7efad 100644 --- a/tools/perf/util/trace_augment.h +++ b/tools/perf/util/trace_augment.h @@ -8,14 +8,12 @@ struct evlist; =20 #ifdef HAVE_BPF_SKEL =20 -int augmented_syscalls__prepare(void); +int augmented_syscalls__prepare(bool inherit); +int augmented_syscalls__attach(void); int augmented_syscalls__create_bpf_output(struct evlist *evlist); void augmented_syscalls__setup_bpf_output(void); int augmented_syscalls__set_filter_pids(unsigned int nr, pid_t *pids); int augmented_syscalls__set_target_pids(unsigned int nr, pid_t *pids); -int augmented_syscalls__add_target_pid(pid_t pid); -int augmented_syscalls__del_target_pid(pid_t pid); -bool augmented_syscalls__has_target_pid(pid_t pid); int augmented_syscalls__set_target_syscalls(unsigned int nr, int *syscall_= ids, bool not_syscalls); int augmented_syscalls__get_map_fds(int *enter_fd, int *exit_fd, int *beau= ty_fd); struct bpf_program *augmented_syscalls__find_by_title(const char *name); @@ -24,11 +22,16 @@ void augmented_syscalls__cleanup(void); =20 #else /* !HAVE_BPF_SKEL */ =20 -static inline int augmented_syscalls__prepare(void) +static inline int augmented_syscalls__prepare(bool inherit __maybe_unused) { return -1; } =20 +static inline int augmented_syscalls__attach(void) +{ + return 0; +} + static inline int augmented_syscalls__create_bpf_output(struct evlist *evl= ist __maybe_unused) { return -1; @@ -50,21 +53,6 @@ static inline int augmented_syscalls__set_target_pids(un= signed int nr __maybe_un return 0; } =20 -static inline int augmented_syscalls__add_target_pid(pid_t pid __maybe_unu= sed) -{ - return 0; -} - -static inline int augmented_syscalls__del_target_pid(pid_t pid __maybe_unu= sed) -{ - return 0; -} - -static inline bool augmented_syscalls__has_target_pid(pid_t pid __maybe_un= used) -{ - return false; -} - static inline int augmented_syscalls__set_target_syscalls(unsigned int nr = __maybe_unused, int *syscall_ids __maybe_unused, bool not_syscalls __maybe_unused) --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Fri Sep 25 03:16:03 2026 Received: from mail-pg1-f198.google.com (mail-pg1-f198.google.com [209.85.215.198]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6576843C05B for ; Thu, 17 Sep 2026 06:43:00 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.198 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789627382; cv=none; b=o2DezmpfHnsR1c2G0x0CkGBzE9DCrWsqoxtY+SSSh6It9PS+7as2xUoaOVEmMM9EKRL2TP/F76kK9hUrgh+MdWZCvC4ElGm2MA5LJNrGW7/vN9sXaMLx5mA1lZHWPN79jsO+TOSsOmJWhjUA1GTJ6YksS28RTBq2FpUuQ5QMon0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789627382; c=relaxed/simple; bh=ijr3LP2P40ZUq1DuBNqRk+AgK81THBcVB9Y2LVWBVcs=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=jy4pBCJXiKkSefM/CRKbG7MauDTdvXFQe6ROP1ciCQjaWId0nfVdaCFL34vh+c6pOShqNAeT9xBqEvsJewvfnS0F3KfzlOjKcbUXAzg8IRCvsolQgcvIdVyqh1aPE6h21xA3KSzZDuk/AK1U7Fb/Qysir6nwDPdVaO5zUJrlVKM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--irogers.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=g1QZU5l9; arc=none smtp.client-ip=209.85.215.198 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--irogers.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="g1QZU5l9" Received: by mail-pg1-f198.google.com with SMTP id 41be03b00d2f7-cc42a07d04aso434922a12.1 for ; Wed, 16 Sep 2026 23:43:00 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789627380; x=1790232180; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=1MnZowd6oTOfaxXZbzT35kcRhO2jRSIVNLLWdFtgUAI=; b=g1QZU5l92FrF56BSE5S8NYmAqQe6Vu9k9YZy5n0saBB26zvahaStdb+/WmCykdutvs VkcYbFWZrq4L3YfS+CniXX2sgKAtzXogN+cid04uPF4JXxMmXty0HJGVDBXSdjvCPK4s ln3wHkFxPQpduckgQNyON3Les6mkZ9ulYb8Rgs8+GIaQiN6QnSC/8dG6znEz155cM+Ic ++H/KZN8CYBqUIZNju/QLLxG8J+3OLMeNJaefZm8eGnJ0VAxaW4PkHcUMsI8Pft2Gktz T07Io4BxshRdCYFLYn1rFquCOddwL1TEuPu/Tr6Rsxi9J8ivpQaubtsnO6wdtfJKWzGl r1QA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789627380; x=1790232180; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=1MnZowd6oTOfaxXZbzT35kcRhO2jRSIVNLLWdFtgUAI=; b=jsZN/qHIzzBZ2Y98opPGVziVnur78qDnwLUtuZEArjOdOTQv+ZsflV+CcLicjEv0OU durJhJgcNfRI7+kpywkoXOok3uXeWeJaPix6BtvoA9ekohgDnIQhupOetJxuXb0hBKTK H+ppdql+R/a/O4uhC/rY2icW16N5VPFJwdoS4e1XZKIEkW3uhrXFyWeKybBxM8qgbM5U azvxJjuvJOUCybXQxKE+SRFxd1xTAlA8pkJDo9bUNK4K4ByRKA4m94lQH7WUHB3Q7nDm gmye+jebgG1/iLF5qJr0zZXr0ufQJqU6LHUHJzVGcjxIvtfyd955URtWha1DL6TQU8GG Yw5w== X-Forwarded-Encrypted: i=1; AKwUvBzh0TaudB4BexnBfdLERTb8p25fHCvisSlLgNxkFu5srYSXXACMLKCxZmV1KdKQY3+fLjFkRw6mWloOPxE=@vger.kernel.org X-Gm-Message-State: AFuF++lJinwwlhtXYaKA5pUfC/0uFuLzOeQnXAIosWEFyDlszXnbB+Cp DRPHdOIpc18DnW7OHXxGry2YpFeUbObu27FVXlf4pXov3rcQilkk3OsBUBPzWs3nbn2plnIbT+p U3XW/WHOfwA== X-Received: from dleb1-n2.prod.google.com ([2002:a05:701b:4241:20b0:144:4896:d41]) (user=irogers job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a20:6a0b:b0:3d3:ae1f:d7f3 with SMTP id adf61e73a8af0-3dd5f73ef30mr14251445637.19.1789627379362; Wed, 16 Sep 2026 23:42:59 -0700 (PDT) Date: Wed, 16 Sep 2026 23:42:28 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: Subject: [PATCH v1 06/13] perf test test_task_analyzer: Isolate in temporary directory and make non-exclusive From: Ian Rogers To: Arnaldo Carvalho de Melo , Namhyung Kim Cc: Peter Zijlstra , Ingo Molnar , Jiri Olsa , Adrian Hunter , James Clark , linux-perf-users@vger.kernel.org, linux-kernel@vger.kernel.org, Ian Rogers Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" test_task_analyzer.sh writes perf.data and temporary files directly into the current working directory, causing collisions when running tests in parallel. As a temporary measure until `perf script report` supports an input file option, resolve perfdir to an absolute path, change directory into $tmpdir for the test duration, and clean up in the exit trap. Remove the (exclusive) tag so the test runs in parallel. perfdir is derived from $0, which may be relative, so it has to be resolved before the cd into $tmpdir, otherwise both PERF_EXEC_PATH and the cleanup trap point at paths that no longer exist. Assisted-by: Antigravity:gemini-3.1-pro Signed-off-by: Ian Rogers --- tools/perf/tests/shell/test_task_analyzer.sh | 11 +++++++---- 1 file changed, 7 insertions(+), 4 deletions(-) diff --git a/tools/perf/tests/shell/test_task_analyzer.sh b/tools/perf/test= s/shell/test_task_analyzer.sh index 0314412e63b4..6f729d9508f8 100755 --- a/tools/perf/tests/shell/test_task_analyzer.sh +++ b/tools/perf/tests/shell/test_task_analyzer.sh @@ -1,8 +1,13 @@ #!/bin/bash -# perf script task-analyzer tests (exclusive) +# perf script task-analyzer tests # SPDX-License-Identifier: GPL-2.0 =20 +# Resolve the source directory before changing the working directory below, +# $0 may be a relative path and would no longer resolve from $tmpdir. +perfdir=3D$(cd "$(dirname "$0")/../.." && pwd) + tmpdir=3D$(mktemp -d /tmp/perf-script-task-analyzer-XXXXX) +cd "$tmpdir" || exit 1 # TODO: perf script report only supports input from the CWD perf.data file= , make # it support input from any file. perfdata=3D"perf.data" @@ -11,7 +16,6 @@ csvsummary=3D"$tmpdir/csvsummary" err=3D0 =20 # set PERF_EXEC_PATH to find scripts in the source directory -perfdir=3D$(dirname "$0")/../.. if [ -e "$perfdir/scripts/python/Perf-Trace-Util" ]; then export PERF_EXEC_PATH=3D$perfdir fi @@ -20,8 +24,7 @@ fi export ASAN_OPTIONS=3Ddetect_leaks=3D0 =20 cleanup() { - rm -f "${perfdata}" - rm -f "${perfdata}".old + cd "$perfdir" || cd /tmp || exit rm -rf "$tmpdir" =20 trap - exit term int --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Fri Sep 25 03:16:03 2026 Received: from mail-pj1-f71.google.com (mail-pj1-f71.google.com [209.85.216.71]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E9E7543D4E9 for ; Thu, 17 Sep 2026 06:43:01 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.71 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789627384; cv=none; b=VoUweisQx0Z4i4Wkcx+UDvB4KCQo2iCxHX4wZ1vsznp2fjxqtLGjlHTvEuIs3U2DQTUJdI8EfKDL01JfLNEcNqvS7nQJM2AohUeDANKWeF12j/D3UyatKx+hDiVyaRbwpCr2hGsrz+C0DHG5Pqewl/7yk0kevtVhkcjivf0wTZk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789627384; c=relaxed/simple; bh=Efwg6Lp4URje4K6ap1NY7WHw9tf5g4FaF0Uk7IwjA5I=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=CGyBq4BhjAViyX5IXbexrKXGzl/MQM/zGZFdSxD4H7BUWmGyChNECV3Gy1xc4FiXaY0Fq6eL/R9RkD+hc4Ggq2XNlqqxtXgOKCSM27S5qq3FaaTNU5zlOnj0faVWHGcVvoYAF1MRWqnM5NqL1HG5S957nwLs/mflZwJ721rhf/c= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--irogers.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=KPcz8wpy; arc=none smtp.client-ip=209.85.216.71 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--irogers.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="KPcz8wpy" Received: by mail-pj1-f71.google.com with SMTP id 98e67ed59e1d1-3968bb86fb7so647976a91.2 for ; Wed, 16 Sep 2026 23:43:01 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789627381; x=1790232181; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=Aji62SWFx1lY4aZnTGM78az2i6brsVbPSUcmX+4ftqQ=; b=KPcz8wpy8e+zWZnq1guMIyqtclqam5iDNFqK67z0LTFx9Zlu+jjWkXe3Gu1W30hGO6 sqBV4UUMrLH5xmbyh3h4vxMHXtDZBfsENCpBkljnlOnfVMHWSXoWlMEU41Q0TXk9nMqz s0POZy52AXVTYgElOs90tGCxneWPPblJkHmgKiqRDWSuZ55biWfuvoUX4dbtYU/s0wfV Y0j20UcIiBe4P5fs+g7ZSiGfSPr5gLBWswrl51u+LTGimAwcXJ1N1AdC1VmDcMFoubhs q65rUQXdd08K49ItKD6N7N3j//gwibf0W+ChFi3IcmcnrDoK9Xrp9h0eL/z/lmOLv3Ts LWjQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789627381; x=1790232181; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=Aji62SWFx1lY4aZnTGM78az2i6brsVbPSUcmX+4ftqQ=; b=EJzGcS+SRewmxvWNpHy4VQC0i0sNzjlfzUw5TW8npgAiqnvBuFGZ8gnodPi6ia9FUA 3yzIMWwuVTyGbsaJ+hLD64Zy3qkKajdF9Q3pXjzDKN85zTr7Gro8ev58JCEaCH1Ya2ss BIHgohYJd4N+rgVCKXRoJj/Xb49heROrQyekfezoWse271weVLM4ry+rpEZ1ppT+8ui9 HAer1VuklSRbaac+OFbRotHqF6VPuQfz2hUe/c/J2eUIFQtKJ924KhxJtnpw3k8dniRI OxxpMd7BQzVr5mp3dfwT7N5UepfPEfOPmUdosh6wfoMtKJIC34jGJH86/aM+lDAjzZqM JBgQ== X-Forwarded-Encrypted: i=1; AKwUvByIAHMkbXkSvYQEKJYJ8ZHJgr06BVY+JmOg1mfd7tZxa8Bk6Etr2w0vL0E2LkkZzMglRB9S/x6RMtU4GiY=@vger.kernel.org X-Gm-Message-State: AFuF++n7VceoxGRY+tZ8bKXLdGL8OfQyXylvk5NqbvSCgfOlTCAlZAcs XCI8iRP3WbYtJ1dSWhPbdcO8cYb2197yRxivua2rgy0Lae0Yq5b2FTYnfygeqsDna/REkX8IK0b VXopAI/vI5Q== X-Received: from dldyq26-n2.prod.google.com ([2002:a05:701b:455a:20b0:144:58aa:2e86]) (user=irogers job=prod-delivery.src-stubby-dispatcher) by 2002:a17:90b:3885:b0:39d:f253:ed66 with SMTP id 98e67ed59e1d1-39e1e51d828mr12013395a91.22.1789627381072; Wed, 16 Sep 2026 23:43:01 -0700 (PDT) Date: Wed, 16 Sep 2026 23:42:29 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <390cfc25923c5be2636ee0bd59a0b33d29ba0e71.1789626978.git.irogers@google.com> Subject: [PATCH v1 07/13] perf test common: Do not globally disable tracing events in clear_all_probes From: Ian Rogers To: Arnaldo Carvalho de Melo , Namhyung Kim Cc: Peter Zijlstra , Ingo Molnar , Jiri Olsa , Adrian Hunter , James Clark , linux-perf-users@vger.kernel.org, linux-kernel@vger.kernel.org, Ian Rogers Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" `echo 0 > /sys/kernel/debug/tracing/events/enable` disables tracepoint events system-wide. When running tests in parallel, this kills active tracing and recording sessions in concurrent tests (such as perf trace and perf record). Remove the global event disable from clear_all_probes(). Assisted-by: Antigravity:gemini-3.1-pro Signed-off-by: Ian Rogers --- tools/perf/tests/shell/common/init.sh | 1 - 1 file changed, 1 deletion(-) diff --git a/tools/perf/tests/shell/common/init.sh b/tools/perf/tests/shell= /common/init.sh index cbfc78bec974..d2c7a31e2c6f 100644 --- a/tools/perf/tests/shell/common/init.sh +++ b/tools/perf/tests/shell/common/init.sh @@ -132,7 +132,6 @@ check_uprobes_available() =20 clear_all_probes() { - echo 0 > /sys/kernel/debug/tracing/events/enable check_kprobes_available && echo > /sys/kernel/debug/tracing/kprobe_events check_uprobes_available && echo > /sys/kernel/debug/tracing/uprobe_events } --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Fri Sep 25 03:16:03 2026 Received: from mail-pg1-f198.google.com (mail-pg1-f198.google.com [209.85.215.198]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CE78B43DA2C for ; Thu, 17 Sep 2026 06:43:03 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.198 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789627389; cv=none; b=Ye0YmTGX/4cjCyd4BTli8qiq+NRtmvkePVWjdh18oJkryS/yAL5MnjzaWzRBzwRlCQYfdnrBkH131GtmSgyInq2Yyt/hwcRBhuC3VXfiriSitarX729kP5cIyQEiUvWmLvY2vJXVdgb0eTiYb7DfzAcW5sAXl3+xbnI2s08Nyfo= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789627389; c=relaxed/simple; bh=/npCTcaijC5/X5gCIcLm38hBGQVHZa7g3OeFrirv27o=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=Z1ZufzH2jP1F0YL1QaF3jGw6C1xoBGP/01bK8izlzLhu4CKbA6HluHoN+vg377KaOw4xTzXCa1icBNd6DEXoCCG/iJG3F7zfRbI0fc+Cj7GByAWE2Tvz/VqbQc11uzYgXKS0aoudGZfnFxYq9MKSHo+Zz3MBgH/kwWY6sVD1fCk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--irogers.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=Vic4k6Z6; arc=none smtp.client-ip=209.85.215.198 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--irogers.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="Vic4k6Z6" Received: by mail-pg1-f198.google.com with SMTP id 41be03b00d2f7-c89704da8c7so756054a12.0 for ; Wed, 16 Sep 2026 23:43:03 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789627383; x=1790232183; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=+++jh0v+Oa7psRCZBUK/IJfrh8ZKmGANC2vEsQ4ZRpc=; b=Vic4k6Z6Yr9lvKrt2VDAQ17SeD4Gw3qG4hIOe6fekGNzrWNyyHJkkgdT/XJtMf7qg1 JJiYyDTTgNdW6TtjW81dzU3NAIxMXZ57IC5YuEGMnWmeOagzeymqKNGFhWhmMzIh0W/b gsdw1mwo375xlRF40Je0xc6kD2YPtFK1FKzOtGvlOQsYerYzC39jnQJq56N4ivwRV3UE 6gRhfQN3pLcB2sUssUxBdoL8idxlaWVdVMwk3BNcBDB5gDNyZvsd/TjvRpvtxU9+GuQy /5GMLjJOFYfmNnUA9JB9O6w+245lq2Yek/s2+vR/zG4DjlbyI6yiu/7sSgsRDkEx4EH1 b8Cw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789627383; x=1790232183; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=+++jh0v+Oa7psRCZBUK/IJfrh8ZKmGANC2vEsQ4ZRpc=; b=sisPlRWYK/T8SfJA0iKHRcb+Yb4t+QAPrQ1J65Yeyu33MjkBVWtCxzI1xyEIsbZxoU hCUWOARnG6xBy+YiWtBxFlIec/4DfYiTCsxiZKcmGqVkEcgaPS8Og4AXAaECIjQp1Pet LxiI5irpB3GakTSlBk3OtlggtJKHFqDhuLX3IamdXiPw/iMNs7T8keP/OClbYE4TDCPQ V3tQ8QTKkVuOBccix4lNqHf0ZrvkQD5356QsgW0tJ41zknu7JzshedsgJj+FbBCbkqnP BuDzri8On7Sqykqk8UgwfIjXzb7p+4AxEUQK6MhoEpPK7/FLDv/2bOAuX43aSHy5Im5Y sYXg== X-Forwarded-Encrypted: i=1; AKwUvBymAhUkMh7qhU+xwDOfXTV8lG6hqu+9M+2IRO7T/p3uAwyCXIiIlddRqLGFNKh4WtiUfU90VHMbJnnqg8A=@vger.kernel.org X-Gm-Message-State: AFuF++nnHyO9wf5twAcTHp7UTLvIKz5o+1ajc1Y4bPHgw9h9bwSQbTnH V/EKud35F0v3iG5lwmxaZ0AIlGqNGZv1lZ7smuZgvKUnnb7mGQO5WlZ08YnkkvAVy51PQZ6/MNm jCfmgWAr3xA== X-Received: from dlg13.prod.google.com ([2002:a05:7022:78d:b0:143:9146:e482]) (user=irogers job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a21:3990:b0:398:8870:b58f with SMTP id adf61e73a8af0-3dd5f57c3a9mr15302756637.14.1789627382884; Wed, 16 Sep 2026 23:43:02 -0700 (PDT) Date: Wed, 16 Sep 2026 23:42:30 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: Subject: [PATCH v1 08/13] perf test probe_vfs_getname: Scope probe name to PID and make non-exclusive From: Ian Rogers To: Arnaldo Carvalho de Melo , Namhyung Kim Cc: Peter Zijlstra , Ingo Molnar , Jiri Olsa , Adrian Hunter , James Clark , linux-perf-users@vger.kernel.org, linux-kernel@vger.kernel.org, Ian Rogers Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" The probe name `vfs_getname` was hardcoded, causing collisions when tests ran concurrently. Furthermore, `cleanup_probe_vfs_getname()` used `perf probe -d probe:vfs_getname*`, deleting probes registered by other parallel tests. Scope the probe name to the pid, and rename it to `getname_flags_$$` so that it no longer begins with "vfs_getname". perf trace calls evlist__add_vfs_getname(), which opens every event matching a hardcoded "probe:vfs_getname*" wildcard, so a perf trace run by any other test would otherwise pin this probe and make `perf probe -d` fail with -EBUSY. That also unblocks making the perf trace tests non-exclusive later in this series. Enumerate the probes to record and to delete from `perf probe -l`, matching `^probe:${vfs_getname}(_[[:digit:]]+)?$` exactly, rather than globbing on `${vfs_getname}*`. perf probe appends _1, _2, ... when getname_flags is inlined at more than one call site, so the variants do have to be matched, but since the name now ends in a pid a trailing wildcard would also match the probes of a test whose pid merely starts with this one's, e.g. 123 and 1234. Remove the `(exclusive)` tag from probe_vfs_getname.sh and record+script_probe_vfs_getname.sh so they run concurrently in pass 1. trace+probe_vfs_getname.sh has to stay exclusive: it is the one test that wants to be discovered by that wildcard, so it sets vfs_getname to a "vfs_getname_$$" name before sourcing the library, and would then pin its siblings' probes if it ran alongside them. A comment in the test records this. Assisted-by: Antigravity:gemini-3.1-pro Signed-off-by: Ian Rogers --- .../perf/tests/shell/lib/probe_vfs_getname.sh | 34 ++++++++++++++++--- tools/perf/tests/shell/probe_vfs_getname.sh | 3 +- .../shell/record+script_probe_vfs_getname.sh | 18 +++++++--- .../tests/shell/trace+probe_vfs_getname.sh | 9 +++++ 4 files changed, 54 insertions(+), 10 deletions(-) diff --git a/tools/perf/tests/shell/lib/probe_vfs_getname.sh b/tools/perf/t= ests/shell/lib/probe_vfs_getname.sh index 88cd0e26d5f6..89a4b6fa5ea1 100644 --- a/tools/perf/tests/shell/lib/probe_vfs_getname.sh +++ b/tools/perf/tests/shell/lib/probe_vfs_getname.sh @@ -1,12 +1,38 @@ #!/bin/bash # Arnaldo Carvalho de Melo , 2017 =20 -perf probe -l 2>&1 | grep -q probe:vfs_getname +# The name of the getname_flags probe added and removed below. +# +# It is scoped to the pid so that tests running in parallel do not collide, +# and it deliberately does not start with "vfs_getname": perf trace calls +# evlist__add_vfs_getname(), which opens everything matching the hardcoded +# "probe:vfs_getname*" wildcard, so a perf trace running in another test w= ould +# otherwise pin this probe and make the 'perf probe -d' below fail with -E= BUSY. +# +# trace+probe_vfs_getname.sh is the one test that does want to be found th= at +# way, so it sets vfs_getname itself before sourcing this file, and is +# (exclusive) as a result. +: "${vfs_getname:=3Dgetname_flags_$$}" + +# Print the probes add_probe_vfs_getname() created. perf probe appends _1,= _2, +# ... when getname_flags is inlined at more than one call site, so there c= an be +# several. Match them exactly rather than with a "${vfs_getname}*" glob: t= he +# name ends in a pid, so such a glob would also match the probes of a test +# whose pid merely starts with this one's, e.g. 123 and 1234. +probes_vfs_getname() { + perf probe -l 2>/dev/null | awk '{print $1}' | + grep -E "^probe:${vfs_getname}(_[[:digit:]]+)?$" +} + +[ -n "$(probes_vfs_getname)" ] had_vfs_getname=3D$? =20 cleanup_probe_vfs_getname() { if [ $had_vfs_getname -eq 1 ] ; then - perf probe -q -d probe:vfs_getname* + local probe + for probe in $(probes_vfs_getname); do + perf probe -q -d "$probe" + done fi } =20 @@ -33,8 +59,8 @@ add_probe_vfs_getname() { return 2 fi =20 - perf probe -q "vfs_getname=3Dgetname_flags:${line} pathname=3Dresu= lt->name:string" || \ - perf probe $add_probe_verbose "vfs_getname=3Dgetname_flags:${line} pathn= ame=3Dfilename:ustring" || return 1 + perf probe -q "${vfs_getname}=3Dgetname_flags:${line} pathname=3Dr= esult->name:string" || \ + perf probe $add_probe_verbose "${vfs_getname}=3Dgetname_flags:${line} pa= thname=3Dfilename:ustring" || return 1 fi } =20 diff --git a/tools/perf/tests/shell/probe_vfs_getname.sh b/tools/perf/tests= /shell/probe_vfs_getname.sh index 5fe5682c28ce..05f1d50732b6 100755 --- a/tools/perf/tests/shell/probe_vfs_getname.sh +++ b/tools/perf/tests/shell/probe_vfs_getname.sh @@ -1,6 +1,5 @@ #!/bin/bash -# Add vfs_getname probe to get syscall args filenames (exclusive) - +# Add vfs_getname probe to get syscall args filenames # SPDX-License-Identifier: GPL-2.0 # Arnaldo Carvalho de Melo , 2017 =20 diff --git a/tools/perf/tests/shell/record+script_probe_vfs_getname.sh b/to= ols/perf/tests/shell/record+script_probe_vfs_getname.sh index 002f7037f182..1d4fb4a4fbfe 100755 --- a/tools/perf/tests/shell/record+script_probe_vfs_getname.sh +++ b/tools/perf/tests/shell/record+script_probe_vfs_getname.sh @@ -1,5 +1,5 @@ #!/bin/bash -# Use vfs_getname probe to get syscall args filenames (exclusive) +# Use vfs_getname probe to get syscall args filenames =20 # Uses the 'perf test shell' library to add probe:vfs_getname to the system # then use it with 'perf record' using 'touch' to write to a temp file, th= en @@ -17,22 +17,32 @@ skip_if_no_perf_probe || exit 2 =20 # shellcheck source=3Dlib/probe_vfs_getname.sh . "$(dirname "$0")/lib/probe_vfs_getname.sh" +# shellcheck disable=3DSC2154 # vfs_getname is assigned in lib/probe_vfs_g= etname.sh =20 record_open_file() { echo "Recording open file:" # Check presence of libtraceevent support to run perf record - skip_no_probe_record_support "probe:vfs_getname*" + skip_no_probe_record_support if [ $? -eq 2 ]; then echo "WARN: Skipping test record_open_file. No libtraceevent support" return 2 fi - perf record -o ${perfdata} -e probe:vfs_getname\* touch $file + # Record every probe the inlining of getname_flags produced, naming + # them exactly rather than with a "${vfs_getname}*" glob, which would + # also match the probes of a test whose pid starts with this one's. + local events + events=3D$(probes_vfs_getname | paste -sd, -) + if [ -z "${events}" ] ; then + echo "FAIL: no ${vfs_getname} probe to record" + return 1 + fi + perf record -o ${perfdata} -e "${events}" touch $file } =20 perf_script_filenames() { echo "Looking at perf.data file for vfs_getname records for the file we t= ouched:" perf script -i ${perfdata} | \ - grep -E " +touch +[0-9]+ +\[[0-9]+\] +[0-9]+\.[0-9]+: +probe:vfs_getname[= _0-9]*: +\([[:xdigit:]]+\) +pathname=3D\"${file}\"" + grep -E " +touch +[0-9]+ +\[[0-9]+\] +[0-9]+\.[0-9]+: +probe:${vfs_getnam= e}(_[0-9]+)?: +\([[:xdigit:]]+\) +pathname=3D\"${file}\"" } =20 add_probe_vfs_getname diff --git a/tools/perf/tests/shell/trace+probe_vfs_getname.sh b/tools/perf= /tests/shell/trace+probe_vfs_getname.sh index 7a0b1145d0cd..146305f4d549 100755 --- a/tools/perf/tests/shell/trace+probe_vfs_getname.sh +++ b/tools/perf/tests/shell/trace+probe_vfs_getname.sh @@ -10,6 +10,13 @@ # SPDX-License-Identifier: GPL-2.0 # Arnaldo Carvalho de Melo , 2017 =20 +# This test must stay exclusive, and is the only one of the probe tests th= at +# does: it does not name the event it uses. perf trace discovers it with t= he +# hardcoded "probe:vfs_getname*" wildcard in evlist__add_vfs_getname(), so= the +# probe has to carry that prefix, and a parallel run of this test would th= en +# also match, and pin, the probes of the other tests. The sibling tests av= oid +# all of this by using a name that the wildcard cannot reach. + # shellcheck source=3Dlib/probe.sh . "$(dirname $0)"/lib/probe.sh =20 @@ -17,6 +24,8 @@ skip_if_no_perf_probe || exit 2 skip_if_no_perf_trace || exit 2 [ "$(id -u)" =3D 0 ] || exit 2 =20 +# shellcheck disable=3DSC2034 # consumed by lib/probe_vfs_getname.sh +vfs_getname=3D"vfs_getname_$$" . "$(dirname $0)"/lib/probe_vfs_getname.sh =20 trace_open_vfs_getname() { --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Fri Sep 25 03:16:03 2026 Received: from mail-pg1-f200.google.com (mail-pg1-f200.google.com [209.85.215.200]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3C18F43F081 for ; Thu, 17 Sep 2026 06:43:05 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.200 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789627390; cv=none; b=Znu9CrxDg4jpsexUQ2My73DUYziTpStZs+zJG1VY3PgiZg6MsNrYblfPEp4G2OUW2t1AXWd6j/Fpq+GaUy0Kdof+MLsc8mSCgT015z9izZE8YAe3Dict75ELY3Yc0LPhZUeXLc00EgoBMxt8IBVLK5y6tFxXpOL9Or6I71bvdS0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789627390; c=relaxed/simple; bh=5Yuc8XAiW6kfWmbvTQBKbdZimkwARfCgP8aby2uGwA0=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=Cn+zPsI22hxBwQYmp9fQSr69o0d4INm8GxvdDEaQrO2PRprrqs9dN/x10QBaKlcQ/Y1Rt+wyX0P4tDe+mJQGY57r3XBweFqQbex2h5Ck8hCCzuxrB8amR7rb6Swn4C5AnIpmZocJWCs0OFF3hkDNWetTw/92TFy6t4JsZbdBwU4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--irogers.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=nbKmfIZU; arc=none smtp.client-ip=209.85.215.200 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--irogers.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="nbKmfIZU" Received: by mail-pg1-f200.google.com with SMTP id 41be03b00d2f7-cbe77d6864dso747075a12.2 for ; Wed, 16 Sep 2026 23:43:05 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789627385; x=1790232185; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=Z3dRiQJZJNOXkNTzggpxwH2qCkPBVHKNDZg538ILUlA=; b=nbKmfIZUMA1B9+a/RO+lQMb8mhmAALhblKzPwSFhpZqObZiz1NvUBSMYCjKRZX2fjZ iFF3CIJdaUuwz+D1SEFmfWG6B3n7SyxxgLkzrknzbHvI1BAZXPfyffEYr16kR1c36KKG N2666d4AoZkkPbgtwbqWMCmmSJuVM2zbUD4K3O5aO61xYXkFaYi1SEjgt5j9+FMPvusS R7ab4GHsm6QcJCs4eae3+jqrhSE8U8SPLe5cbi6FTCgIAQE/s3HpXVtwc37sGmdyVIZs ssTHijX4jZ/LTVZ9HSgfU0YWr0SiG2PPaUVcRiAbLSqxqE+TAMyzPj6v63VP6Htm+zbI 7k7w== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789627385; x=1790232185; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=Z3dRiQJZJNOXkNTzggpxwH2qCkPBVHKNDZg538ILUlA=; b=hU6hZkF2ULsnjnJ1H1Dh1AxBIbl0eCKfFSP4jK+bXAgS9P8a9Sz4NAuGqazDCK8GG1 eNpPmf4h6junXPeDNoTDLa/j8LUG6m7F+mLy8OlIFR+3OiEIEWq968F3t+VMUyqWiLR8 8KKlzFV8NyshujTAPnYlL/bDueU4Fann06rmmy4+E55dWn84qJw9tNJd16+zV0nlydbQ wRR0sAzW8kAyWswnpsDrhx3lZ1CKjXOzIznPPZTAeaWi8txkdyUwdtsvcaa7CMIWLcIl EXzhw80FwRPo2ALjuAdQMBBUdfZpaguC2drt5Rk5N5/7hyTeREJxJHr6vUgqjoKCXmW1 WolA== X-Forwarded-Encrypted: i=1; AKwUvByQWHDlUNIO7RSd8KUF/Mw6DQhYC+GrhfdK5JPKhvxqOEp6fa8SYL2oOu2OO2cIEMgAF71RpyVfyu5YTRo=@vger.kernel.org X-Gm-Message-State: AFuF++m1EpxQMIaaLsIqZCOwvWZjjiOQWiGv47z/sn4gDOmbow5Q1FCx DdgqADI/Iivwam2H1LlPFqDXAPjumwHtTfKZv1WimG7iKer+K7M8mlcUrKi+CLvSZfU0x+svCdK UnVBnCtN5MQ== X-Received: from dlbrn14.prod.google.com ([2002:a05:7022:150e:b0:143:8ccf:e704]) (user=irogers job=prod-delivery.src-stubby-dispatcher) by 2002:a17:90b:3c51:b0:39d:f254:4173 with SMTP id 98e67ed59e1d1-39e1e5047b1mr13970770a91.21.1789627384447; Wed, 16 Sep 2026 23:43:04 -0700 (PDT) Date: Wed, 16 Sep 2026 23:42:31 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <09c5cdf8edb4c418e417d98a996e14d6b5e68f8a.1789626978.git.irogers@google.com> Subject: [PATCH v1 09/13] perf test record+probe_libc_inet_pton: Scope event to PID, add retries, and make non-exclusive From: Ian Rogers To: Arnaldo Carvalho de Melo , Namhyung Kim Cc: Peter Zijlstra , Ingo Molnar , Jiri Olsa , Adrian Hunter , James Clark , linux-perf-users@vger.kernel.org, linux-kernel@vger.kernel.org, Ian Rogers Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" The uprobe name was not scoped to PID, and concurrent writes to `/sys/kernel/debug/tracing/uprobe_events` can occasionally return `-EBUSY` when another process holds the tracefs inode lock. Scope the probe event name with `$$` (`inet_pton_$$=3Dinet_pton`) and add a retry loop with backoff for uprobe addition. Drop the `(exclusive)` tag so the test can run in parallel during pass 1. A PID scoped probe is no longer cleaned up by any other test, so add an EXIT/TERM/INT trap to delete it, otherwise an interrupted run leaks the uprobe into the system. The trap is installed only after the root and IPv6 checks that `exit 2` to skip the test, as trap_cleanup() exits 1 and would otherwise turn those skips into failures. Deletion enumerates the probes from `perf probe -l`, matching `^probe_libc:inet_pton_$$(_[[:digit:]]+)?$` exactly, rather than reading $event_name: a signal arriving after perf probe injected the uprobe but before the assignment completed would leave that variable empty and leak the probe, and an `inet_pton_$$*` glob would reach the probe of a test whose pid merely starts with this one's. While here use mktemp rather than mktemp -u for the temporary files: this test runs as root in a world writable /tmp, and predicting a name without creating it allows another user to win the race and plant a symlink. The perf.data check becomes -s rather than -e as mktemp now pre-creates an empty file. Assisted-by: Antigravity:gemini-3.1-pro Signed-off-by: Ian Rogers --- .../shell/record+probe_libc_inet_pton.sh | 83 ++++++++++++++----- 1 file changed, 63 insertions(+), 20 deletions(-) diff --git a/tools/perf/tests/shell/record+probe_libc_inet_pton.sh b/tools/= perf/tests/shell/record+probe_libc_inet_pton.sh index eca629ee83f0..3eb51426373b 100755 --- a/tools/perf/tests/shell/record+probe_libc_inet_pton.sh +++ b/tools/perf/tests/shell/record+probe_libc_inet_pton.sh @@ -1,5 +1,5 @@ #!/bin/bash -# probe libc's inet_pton & backtrace it with ping (exclusive) +# probe libc's inet_pton & backtrace it with ping =20 # Installs a probe on libc's inet_pton function, that will use uprobes, # then use 'perf trace' on a ping to localhost asking for just one packet @@ -21,20 +21,30 @@ nm -Dg $libc 2>/dev/null | grep -F -q inet_pton || exit= 254 event_pattern=3D'probe_libc:inet_pton(_[[:digit:]]+)?' =20 add_libc_inet_pton_event() { + local attempts=3D0 + while [ $attempts -lt 3 ]; do + event_name=3D$(perf probe -f -x $libc -a "inet_pton_$$=3Dinet_pton" 2>&1= | \ + awk -v ep=3D"$event_pattern" -v l=3D"$libc" '$0 ~ ep && $0 ~ \ + ("\\(on inet_pton in " l "\\)") {print $1}' | head -n 1) + + if [ -n "$event_name" ]; then + return 0 + fi + attempts=3D$((attempts + 1)) + sleep 0.1 + done =20 - event_name=3D$(perf probe -f -x $libc -a inet_pton 2>&1 | \ - awk -v ep=3D"$event_pattern" -v l=3D"$libc" '$0 ~ ep && $0 ~ \ - ("\\(on inet_pton in " l "\\)") {print $1}' | head -n 1) - - if [ $? -ne 0 ] || [ -z "$event_name" ] ; then - printf "FAIL: could not add event\n" - return 1 - fi + printf "FAIL: could not add event\n" + return 1 } =20 trace_libc_inet_pton_backtrace() { =20 - expected=3D`mktemp -u /tmp/expected.XXX` + # Create the files rather than just reserving names with mktemp -u: + # this runs as root and /tmp is world writable, so a predictable name + # that is written to later can be pre-created as a symlink by an + # unprivileged user and used to clobber an arbitrary file. + expected=3D$(mktemp /tmp/expected.XXX) =20 echo "ping[][0-9 \.:]+$event_name: \([[:xdigit:]]+\)" > $expected echo ".*inet_pton\+0x[[:xdigit:]]+[[:space:]]\($libc|inlined\)$" >> $expe= cted @@ -50,8 +60,8 @@ trace_libc_inet_pton_backtrace() { ;; esac =20 - perf_data=3D`mktemp -u /tmp/perf.data.XXX` - perf_script=3D`mktemp -u /tmp/perf.script.XXX` + perf_data=3D$(mktemp /tmp/perf.data.XXX) + perf_script=3D$(mktemp /tmp/perf.script.XXX) =20 # Check presence of libtraceevent support to run perf record skip_no_probe_record_support "$event_name/$eventattr/" @@ -61,9 +71,10 @@ trace_libc_inet_pton_backtrace() { fi =20 perf record -e $event_name/$eventattr/ -o $perf_data ping -6 -c 1 ::1 > /= dev/null 2>&1 - # check if perf data file got created in above step. - if [ ! -e $perf_data ]; then - printf "FAIL: perf record failed to create \"%s\" \n" "$perf_data" + # Check perf record actually wrote data. mktemp already created the + # file, so test that it is non-empty rather than that it exists. + if [ ! -s $perf_data ]; then + printf "FAIL: perf record failed to write \"%s\" \n" "$perf_data" return 1 fi perf script -i $perf_data | tac | grep -m1 ^ping -B9 | tac > $perf_script @@ -97,21 +108,53 @@ trace_libc_inet_pton_backtrace() { # even if the perf script output does not match. } =20 +# Print the pid scoped uprobes this test may have created. perf probe appe= nds +# _1, _2, ... when the name is already taken, so match those too, but anch= or +# the match: an "inet_pton_$$*" glob would also match the probe of a test = whose +# pid merely starts with this one's, e.g. 123 and 1234. +libc_inet_pton_events() { + perf probe -l 2>/dev/null | awk '{print $1}' | + grep -E "^probe_libc:inet_pton_$$(_[[:digit:]]+)?$" +} + delete_libc_inet_pton_event() { + # Ask the kernel what is actually there rather than trusting + # $event_name: a signal arriving after perf probe injected the uprobe + # but before the assignment to event_name completed would otherwise + # leave the variable empty and leak the probe. + local probe + for probe in $(libc_inet_pton_events); do + perf probe -q -d "$probe" + done +} =20 - if [ -n "$event_name" ] ; then - perf probe -q -d $event_name - fi +cleanup() { + rm -f ${perf_data} ${perf_script} ${expected} + delete_libc_inet_pton_event + + trap - EXIT TERM INT +} + +trap_cleanup() { + cleanup + exit 1 } =20 # Check for IPv6 interface existence ip a sh lo | grep -F -q inet6 || exit 2 [ "$(id -u)" =3D 0 ] || exit 2 =20 +# Install the trap only now that the skips above are out of the way: it ex= its +# 1, so arming it any earlier would turn an 'exit 2' skip into a failure. +# +# The event name is pid scoped, so unlike the old fixed name an orphan left +# behind by an interrupted run is never overwritten by a later run: it wou= ld +# stay in the kernel forever. Always clean up, including on a signal. +trap trap_cleanup EXIT TERM INT + skip_if_no_perf_probe && \ add_libc_inet_pton_event && \ trace_libc_inet_pton_backtrace err=3D$? -rm -f ${perf_data} ${perf_script} ${expected} -delete_libc_inet_pton_event +cleanup exit $err --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Fri Sep 25 03:16:03 2026 Received: from mail-pj1-f69.google.com (mail-pj1-f69.google.com [209.85.216.69]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3852543CE71 for ; Thu, 17 Sep 2026 06:43:06 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.69 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789627391; cv=none; b=mdOLZLqpMo9E8xK+BaMr9zBUKqHUUk2v6fvVC1B6C6PDe35kNsX9k8MwEw7tSk8Cf0W5PKeig8XDSsg2AwA9TbXk/HH1KXyoZWa8gOu7aH0V0rvTUNmLULCTBaLL0f2LgN+ze9wgeWqZy8ZPEUuhTeNXr7SpRAQ6Yyo3Jc5Wt7U= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789627391; c=relaxed/simple; bh=P+avWqg0YJ5z30ktLMdV7Tr+I207Djb9tMMIgwKH6dY=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=U95UyWuPreU7Hp7EYoOhGE8QvWtvFeQY4uFBT2a8pKphUdQaj9sjza1FpidkD5o+vpj9QyYHq3bRjokX0zJq9KwW9/+gvX0LD17W2BwdJ65a65JtRNrJ32I7jftTVPOp3oIui8GMROM9q+FuZKqVzriAr7WdTmeJFvLlFlymycM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--irogers.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=nUTVsA9R; arc=none smtp.client-ip=209.85.216.69 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--irogers.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="nUTVsA9R" Received: by mail-pj1-f69.google.com with SMTP id 98e67ed59e1d1-395543dc382so742720a91.0 for ; Wed, 16 Sep 2026 23:43:06 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789627386; x=1790232186; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=SxRrlx2WvAT37qZTLQX7yxCGnnR8nCXcvNsM0kQ33yc=; b=nUTVsA9R+E81pBlr2gl971MNdsIh1jAELoaCeleqB9N+HlsZnOxnOg6BI9ml8j6cpF IUK5SaQaSNp8nGTV+glAsGukXt7PLOAtOQA+eRy4NTVv+wnkOdw6YsNHtqcypwncudAg 2xQpq/zmGHzk6VSZZIZkwqvVJnQXW4Ty0nV0laCU5v/U/x9ir3E0c8vbEkla/qbR1ZXK Vt8Nzw0MOEpgumJIvi15C7r/ir4LWE/VdQCuMEv/IQc7EYRuxOq2RG38A0/qW51vqwTf 0su4dzs415ioRPG+4CZqLqNQGJvJSPwbkESQYvXdA3jr15xet5meGdKcQ6LJOrwE7xpL pIiw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789627386; x=1790232186; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=SxRrlx2WvAT37qZTLQX7yxCGnnR8nCXcvNsM0kQ33yc=; b=Wv5YREoDMXILc+IRNoaiMaXkibQQjWv/f5ztL6egHorUsZXEeJSzS871RPF+P4a0VV PpY4Bir25/jNHx1Mq1x87aWDj274aqglhSGB5+WfKZbb652TAgjgL3L2ImoOm8XnawLG SWvwDhsIkahcaXDK8YemwcQiHRbH1HMHmEZ1FFx0S6vUuXkfbCt3wNBKaJ2tpRc2k4cy t8kIv5zAeGrZKf6QvGcves5NcjJXA3QBCa8uIK97DghKRvuNStc3SGmfFmqPBLu8/Mwn 7RvYM4chuLyaBfhHAVusfxhXEtnhbArLllzK62QMhdcU3uJ+1x9ZejTv/P0jtbi46o6Q GPaw== X-Forwarded-Encrypted: i=1; AKwUvBxM8233XgkEXCI6IHIW2wvnQx79ouzuJEQdu7s4lGg//agWzF/arxn9vXEECLP7BMoF+YEivGgYrhILpIQ=@vger.kernel.org X-Gm-Message-State: AFuF++mj4Odvx8+ChIDq4nJn6UZkqsdB4teP/flAKHQH9VWS+OXpialD TCv0Cgsl09338wMtU32IozBMMpH/iK57Qab4v9xRpNqwhET9EYWtjF2yRb6HFcWc15yd2sSN96G TIHrP/DmU2g== X-Received: from dloo5.prod.google.com ([2002:a05:7023:a45:b0:144:bba4:7daa]) (user=irogers job=prod-delivery.src-stubby-dispatcher) by 2002:a17:90a:e7d2:b0:398:a145:5d3d with SMTP id 98e67ed59e1d1-39e1e2daa81mr14993947a91.6.1789627386022; Wed, 16 Sep 2026 23:43:06 -0700 (PDT) Date: Wed, 16 Sep 2026 23:42:32 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <048255184fea635f244bc8af71c34339085150b7.1789626978.git.irogers@google.com> Subject: [PATCH v1 10/13] perf test trace_summary: Improve error diagnostics From: Ian Rogers To: Arnaldo Carvalho de Melo , Namhyung Kim Cc: Peter Zijlstra , Ingo Molnar , Jiri Olsa , Adrian Hunter , James Clark , linux-perf-users@vger.kernel.org, linux-kernel@vger.kernel.org, Ian Rogers Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" When pattern matching fails in test_perf_trace(), print the command that failed along with the actual match count, the matching lines found, and the first 15 lines of output to aid debugging. Assisted-by: Antigravity:gemini-3.1-pro Signed-off-by: Ian Rogers --- tools/perf/tests/shell/trace_summary.sh | 12 ++++++++---- 1 file changed, 8 insertions(+), 4 deletions(-) diff --git a/tools/perf/tests/shell/trace_summary.sh b/tools/perf/tests/she= ll/trace_summary.sh index b80dea77cec6..f975176247b5 100755 --- a/tools/perf/tests/shell/trace_summary.sh +++ b/tools/perf/tests/shell/trace_summary.sh @@ -28,10 +28,14 @@ test_perf_trace() { =20 count=3D$(grep -E -c -m 3 "${search}" ${OUTPUT}) if [ "${count}" !=3D "3" ]; then - echo "Error: cannot find enough pattern ${search} in the output" - cat ${OUTPUT} - rm -f ${OUTPUT} - exit 1 + echo "Error: cannot find enough pattern ${search} (count=3D${count= }) in output of:" + echo "Error: perf trace ${args} -- ${workload}" + echo "Error: matched lines:" + grep -E "${search}" ${OUTPUT} || echo "none" + echo "Error: first 15 lines of output:" + head -n 15 ${OUTPUT} + rm -f ${OUTPUT} + exit 1 fi } =20 --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Fri Sep 25 03:16:03 2026 Received: from mail-pj1-f70.google.com (mail-pj1-f70.google.com [209.85.216.70]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E34C543F080 for ; Thu, 17 Sep 2026 06:43:08 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.70 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789627390; cv=none; b=UGoXYPpYSwSO1IBt2THsde0QDhBLfZhmTjBKnbMQQ1MuFQgIm4ImQoCwHlU3kn+CCGTiLPiiPXgryj9GZRVjwpE96Jx5SXypEwNlMEjd2+fbZ/3np0vuF4J+UzIqlwh9/uv7SYUXlEtyR8UVZFZXjekKkz5bhRTZ5DazG5Vcjeg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789627390; c=relaxed/simple; bh=uXUyVsWXdGRR9r0Yf+u2xCJJUCEXrufHCD6FF1dYRUw=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=oqdAFciwJxYq9PHphokGyccQ0Beuk/M6kyyYYVnYUpigmAdAqtIYYiaCXTzV6jJ9fALs9/rJNOQLj9IrOmL92GG119qXhnrvtoVH3/TFtjNjVj4/YIXyly1aNZa0vXfYOennZMtK6XfKK6lblsoDTsoPG0vE8e2McidP/djvN3g= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--irogers.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=Xx19uWQW; arc=none smtp.client-ip=209.85.216.70 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--irogers.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="Xx19uWQW" Received: by mail-pj1-f70.google.com with SMTP id 98e67ed59e1d1-38dbf293831so1102899a91.3 for ; Wed, 16 Sep 2026 23:43:08 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789627388; x=1790232188; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=0q/ZlHbTsYZyN2/8M0P/yh3YerzxVzD8tkGrmWQ37IA=; b=Xx19uWQWCX35d3V+qK6dt2sJnFO8CA/WGKOi3J08gmoTp/+eMEF6xkLQSCSaAZYTlx BA8pnXxeiITtleV6mm0TXlSFQpGBEMUQ23yZ3S6tf08VfN9/0LF8DP4A6O9p/bFLwMrD fxlz+PHPHnPRQL7R0m0GOLrtLJ8EZ9k3fz4u+ircu3mbVaCvQQV485Cyh+dis/iCMHDw Z9ZS+Cx2deulaN0r0jEv/mWbsjN+bz5TCIuxgpooNChiJkJzvGxa5jaQO2N3APrj/fWp v9O8/yoD+6THKZvMEv/TvsGzAsLJfKLMFlIpjB6JEi6MzPkw28uVh2istkvqiDo/2f0C fbrA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789627388; x=1790232188; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=0q/ZlHbTsYZyN2/8M0P/yh3YerzxVzD8tkGrmWQ37IA=; b=cTYWBbPrfIFLd7eTCvu3LHDmAKcsNGt/0CvDJTl7SvIUIMj3OVIEMUf+Uol+K2BMc0 zFbShQPSgXyLu0rV9dqkCVDBfUDRzQ3QHvXQkyo9AxuB5OC5S8Uutm9bp/zTrRJ6Kn3Y y90dy08QyHLcyluePrUPacmrmBje+YkUaiSTSw8s0uTXnOJ9pUdzE83ZCNUGXDP6AaL9 +nvil3uG78DVXevTRbbI77KDsnffUXbwey1PSdGKqOlyac8cyF7iT6QvicemwlRNshkf HNFM6aUu7QJWFK+fJOx4zIM/7rglXaVjCZweEuePOIesp1ac+UhmsxW5/5vClQigPgzs qztw== X-Forwarded-Encrypted: i=1; AKwUvBzWAdQrCrGUSZ1Fplf6puApa+6aZAHbITQdwPK2rDe/k4b+MFQARqdRV4RzVLuYJwjo2K0NOEu9HltJMlw=@vger.kernel.org X-Gm-Message-State: AFuF++ml0FdBCiXGCCK8H/NfzmuSXC8XVxSPPxxQS+rvLYbjNnqId9YH pVWDwlX/uO7FCwW/TfeN1iAstyaX7tA8hfZEiumU2DsKWHglYEowJQZsQyWpG3o8TbEUHe+EQrY JUPylepDxEQ== X-Received: from dlbps9.prod.google.com ([2002:a05:7023:889:b0:143:8c65:6162]) (user=irogers job=prod-delivery.src-stubby-dispatcher) by 2002:a17:90b:3a43:b0:39d:f5b1:e365 with SMTP id 98e67ed59e1d1-39e1e559b78mr13296706a91.25.1789627388112; Wed, 16 Sep 2026 23:43:08 -0700 (PDT) Date: Wed, 16 Sep 2026 23:42:33 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <2116386f185b31e0a0e197d37151d585ddb02019.1789626978.git.irogers@google.com> Subject: [PATCH v1 11/13] perf test trace_btf_general: Drop --max-events=1 and make non-exclusive From: Ian Rogers To: Arnaldo Carvalho de Melo , Namhyung Kim Cc: Peter Zijlstra , Ingo Molnar , Jiri Olsa , Adrian Hunter , James Clark , linux-perf-users@vger.kernel.org, linux-kernel@vger.kernel.org, Ian Rogers Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" trace_btf_general.sh used `--max-events=3D1` with `perf trace` on commands such as `mv`, `echo`, and `sleep`. When background activity occurs or tests run in parallel, `perf trace` can capture an event from an unrelated process and exit prematurely before recording the target command's syscalls. Drop `--max-events=3D1` and let tracing run until the command completes, checking for the expected output with grep (matching trace_btf_enum.sh). Remove the (exclusive) tag so the test runs in parallel. Running perf trace in parallel is only safe now that the probe tests no longer name their probes "vfs_getname...": perf trace opens everything matching a hardcoded "probe:vfs_getname*" wildcard, which used to pin those probes and make their cleanup fail with -EBUSY. Assisted-by: Antigravity:gemini-3.1-pro Signed-off-by: Ian Rogers --- tools/perf/tests/shell/trace_btf_general.sh | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/tools/perf/tests/shell/trace_btf_general.sh b/tools/perf/tests= /shell/trace_btf_general.sh index 7a94a5743924..4d654b687a4e 100755 --- a/tools/perf/tests/shell/trace_btf_general.sh +++ b/tools/perf/tests/shell/trace_btf_general.sh @@ -1,5 +1,5 @@ #!/bin/bash -# perf trace BTF general tests (exclusive) +# perf trace BTF general tests # SPDX-License-Identifier: GPL-2.0 =20 err=3D0 @@ -27,7 +27,7 @@ check_vmlinux() { =20 trace_test_string() { echo "Testing perf trace's string augmentation" - output=3D"$(perf trace --sort-events -e renameat* --max-events=3D1 -- mv= ${file1} ${file2} 2>&1)" + output=3D"$(perf trace --sort-events -e renameat* -- mv ${file1} ${file2= } 2>&1)" if ! echo "$output" | grep -q -E "^mv/[0-9]+ renameat(2)?\(.*, \"${file1= }\", .*, \"${file2}\", .*\) +=3D +[0-9]+$" then printf "String augmentation test failed, output:\n$output\n" @@ -38,7 +38,7 @@ trace_test_string() { trace_test_buffer() { echo "Testing perf trace's buffer augmentation" # echo will insert a newline (\10) at the end of the buffer - output=3D"$(perf trace --sort-events -e write --max-events=3D1 -- echo "= ${buffer}" 2>&1)" + output=3D"$(perf trace --sort-events -e write -- echo "${buffer}" 2>&1)" if ! echo "$output" | grep -q -E "^echo/[0-9]+ write\([0-9]+, ${buffer}.= *, [0-9]+\) +=3D +[0-9]+$" then printf "Buffer augmentation test failed, output:\n$output\n" @@ -48,7 +48,7 @@ trace_test_buffer() { =20 trace_test_struct_btf() { echo "Testing perf trace's struct augmentation" - output=3D"$(perf trace --sort-events -e clock_nanosleep --force-btf --ma= x-events=3D1 -- sleep 1 2>&1)" + output=3D"$(perf trace --sort-events -e clock_nanosleep --force-btf -- s= leep 1 2>&1)" if ! echo "$output" | grep -q -E "^sleep/[0-9]+ clock_nanosleep\(0, 0, \= {1,.*\}, 0x[0-9a-f]+\) +=3D +[0-9]+$" then printf "BTF struct augmentation test failed, output:\n$output\n" --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Fri Sep 25 03:16:03 2026 Received: from mail-pj1-f69.google.com (mail-pj1-f69.google.com [209.85.216.69]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 464B543CE4B for ; Thu, 17 Sep 2026 06:43:10 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.69 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789627394; cv=none; b=gqcUPEj9YW9AMh0AaBuQqenuBebJCy7JgA6sNwNZgMCQcVLLuC50GzTKnDkOg0limsTPst7HmW/T8q5FEBr4w973zIerJ7cP3XPKzryyimMfuT9kxXUprOub1yY8w39vpdZRNtq/CeLlfw0us6+Ps5JkuPMsNYq/kaCS/eQPC8c= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789627394; c=relaxed/simple; bh=bYT8WVcurMLOEfyNp8NqaepgaxFS5df3jeYD3CQhJBQ=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=a3iXTipf8W1mO9fIZk85kbqQpMvMvRH3dyTFHXE+J7QgrkYF8X5s8RfTirLIigW/WnFH+kS5penI6mOcJeiJfd5pgE5EJOF77xs6gflGF1gtiE8UMDeJlr5akis6jDTZpMw4fM1SYTJtFcIsNpFup1vbEzTfs5vPvlB5pUR7cv0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--irogers.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=U/51aySQ; arc=none smtp.client-ip=209.85.216.69 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--irogers.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="U/51aySQ" Received: by mail-pj1-f69.google.com with SMTP id 98e67ed59e1d1-39deda201bcso945313a91.2 for ; Wed, 16 Sep 2026 23:43:10 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789627390; x=1790232190; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=nhdENFaJa2NETWrV9GF10gghTb/oZJa/1Sd4bZnoKoc=; b=U/51aySQiXcMWreHnUjMs0rU92pKmi5owYoef7i0YzbffgT5ocxkj0CFHgvih1XCIZ T1CJkpcQJVGJNGhRo8EmaYn1TM1W/WvhjUNI1aanP/NrOv5VNt7JoEVbvGdNsK/BJb6t fA9L5uurhI76atheYoTRipBP2GAipyuoXWsyYsBl+CQiGM/wCVP+KccyKbn3FB6czV/l bLBvs0d4XMOx5KK5PIsSqONVgn36xHP80GEYKYbw+lfzzztDGPYwZw2phQYfKEfS38DI 7hMcPJrL+saW3NptVBMtzQpaT99cyPhJIMrOf03BATCeI4sInKJZcTRht39dBH6B4Y0G FYrA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789627390; x=1790232190; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=nhdENFaJa2NETWrV9GF10gghTb/oZJa/1Sd4bZnoKoc=; b=tB2n+3vu3k4dTLGueBb4ST4y7aPXvFXWtFpzcqEB2/h6wXsEaJzycG8haXo/HuSDXx RHesryVIRj4r3vSKeHIrwMO+EhVVOBZtQgqsn8TeWpVROHw7M/m/QvVR4o6ux/0HhQOI TKUOwIC+WRz1WzNAPNZiImxqQVtJ5l/WhbaEUM0kbGzbBrnbjh5MKhqv7qKiGe+bZzLG jBOtnm7YVsOGbAOsmYPo2vSiCE4Mvj/GlD8z5ggT3d2GuBTj9lFQNFnwGkqsaVl2M/7w VjwiugMaUqcmFmz3CEw5nne2Rb3ELatrqCbrppz/bxKA6inah7xlbW0SnGLbNZpUfsuL ggNA== X-Forwarded-Encrypted: i=1; AKwUvBxF+X7rMHK5lySb8swMXrOEVwjUt/LDgQmDt+1ukf/eJuE6kwUqUfmdG9ZrEgE7QPJonQpTMwr/PStjE94=@vger.kernel.org X-Gm-Message-State: AFuF++mjl6pfViX2Uf7QvCMqbfVElys7uT0m6125peTb3dKSotUqEgpY LftNJ0vTyiq8Mg4HLmZZVHyd3dHguHYLsKGawLHmYUGrGLDeGVFMXxgs1ewbGb/6ZbqovBWARY5 zxSJxrqXVUg== X-Received: from dlbut14.prod.google.com ([2002:a05:7022:7e0e:b0:144:c1be:424b]) (user=irogers job=prod-delivery.src-stubby-dispatcher) by 2002:a17:90b:4a51:b0:393:19a3:4f1 with SMTP id 98e67ed59e1d1-39e1e28db8amr11559485a91.6.1789627390174; Wed, 16 Sep 2026 23:43:10 -0700 (PDT) Date: Wed, 16 Sep 2026 23:42:34 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <3c8cb15f2d1ce25de658a7ea65bab2a5493c22d1.1789626978.git.irogers@google.com> Subject: [PATCH v1 12/13] perf test trace_summary: Make non-exclusive From: Ian Rogers To: Arnaldo Carvalho de Melo , Namhyung Kim Cc: Peter Zijlstra , Ingo Molnar , Jiri Olsa , Adrian Hunter , James Clark , linux-perf-users@vger.kernel.org, linux-kernel@vger.kernel.org, Ian Rogers Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" trace_summary.sh tests various summary modes of `perf trace`. It already directs output to a unique temporary file without polluting the current working directory. Remove the (exclusive) tag so it can run concurrently in parallel test runs. Running perf trace in parallel is only safe now that the probe tests no longer name their probes "vfs_getname...": perf trace opens everything matching a hardcoded "probe:vfs_getname*" wildcard, which used to pin those probes and make their cleanup fail with -EBUSY. Assisted-by: Antigravity:gemini-3.1-pro Signed-off-by: Ian Rogers --- tools/perf/tests/shell/trace_summary.sh | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tools/perf/tests/shell/trace_summary.sh b/tools/perf/tests/she= ll/trace_summary.sh index f975176247b5..e2834bc4eee6 100755 --- a/tools/perf/tests/shell/trace_summary.sh +++ b/tools/perf/tests/shell/trace_summary.sh @@ -1,5 +1,5 @@ #!/bin/bash -# perf trace summary (exclusive) +# perf trace summary # SPDX-License-Identifier: GPL-2.0 =20 # Check that perf trace works with various summary mode --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Fri Sep 25 03:16:03 2026 Received: from mail-pj1-f72.google.com (mail-pj1-f72.google.com [209.85.216.72]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 436224446F5 for ; Thu, 17 Sep 2026 06:43:13 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.72 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789627395; cv=none; b=SOueXu2nt1GtSbF0CAVLQ/wb9+my82OG4Btz3NYhkbwCfHVyg9djbXT/iitVGPOAoY5wsVuyVYLJyPO7W/BunGwvd9JCagk4jRJp6xuIUBq4vFryG0iVEF4pJYxu3jFjjJkZxS2GMMrm2x8Js3/fKrnAKUSryBGB0AFcy4MR3g8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789627395; c=relaxed/simple; bh=Ndq3XZzi/EfqSQhUKNfVgunzhKMGpjb6wAnuZVaabKw=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=bsieG9/K0gaVXwjmvPwuEq1bprmvFYMaRGob8KN/sjyV2yrliLiDCwGQ1LxuYDIPyq3wjh0Y0k3MfX7r6LjU/n5RAt3Os1GwfYMGZ2hG6xmhvXGu1/dtMVOdW1L/29x0PN45YYkzq0vXbKMHqGctOkll4XaMlTe8issO2kNRzbM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--irogers.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=VwvRS+AJ; arc=none smtp.client-ip=209.85.216.72 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--irogers.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="VwvRS+AJ" Received: by mail-pj1-f72.google.com with SMTP id 98e67ed59e1d1-38e8e864ef0so954150a91.0 for ; Wed, 16 Sep 2026 23:43:13 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789627392; x=1790232192; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=rtkpF+2hXoqnckm+Z523eWfhpUjaBTRb2bRSM1mpO+k=; b=VwvRS+AJzgJ6yZowNbliN0GTMTsOS2d09TZjbuTKVMyyMZczGVmJt5A4GBc6x7xXkq 26UGxKrchJWwt9k6HnMHaxIoL8J6llYBAgJudZqCpidY9KmHA0mLTFrX/nxxJENOZZtk nMbaAryqpMDe5fYYz1YhMpUFgs/O120KP92OTDqWUnAVJi+G3iOWMVUp4JPPFpZL8zae 0FAlnWNC3g9oSsa50zq22FRnREuZi3GZaYgvs5ntUs+RCagClUVSSbj6f3FCn9OHGsLm jI780lUOvzZMN5rggoy+biO9nESk/XlwJUkicUVYzibP8etuqahD2jXXuekoebjDlF8u cBIw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789627392; x=1790232192; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=rtkpF+2hXoqnckm+Z523eWfhpUjaBTRb2bRSM1mpO+k=; b=kNBf/pqZf723hBv1IDJ+ZEQpJbz6cfG2+fpTwVBRhRM5pWhlP8jdRmYdr56tgHCt0j A2PIJRNz4MSplKpmFKSnsxe2BdjTU9FPouuL7fs/j18pUnk53oF3WCEBLj6+TlOI0e40 Jq8w6K225PnFpLeVbHgeYMtRgPGqo4uSdVVpMrCgWq2i6iQeFctKNfI2lQ9jHsyp8JaO jdLm181TIPdyaNx44jDy0HcI6Fm6yN99QayPA16Kes6Yfzbg36Qo3zIGeE/0O1eNkRpy LPetIL/4qZkKdT39cM0uzCrRBVvneGIHDJ4No9meaYMI2ynZBLsfTozo1Ku3x8wZzgOc rghQ== X-Forwarded-Encrypted: i=1; AKwUvBxmr4v0gYaY59RP0Lwlmx9ro1Q/jydPI28AvUvVAbeVNu4LK4DLQOOUHh46fq8JDrkTtR6FDmxTn3ea1XI=@vger.kernel.org X-Gm-Message-State: AFuF++mjoYjuMwLR5OnxI4h4oQCpQ5CGdnKaUBBuc+TJ8z+M7EOm0VYQ 7+FhCMTFMwY/mtUtMRhPUcS4TkZpQGj8BXEb2Uwbpgpe1AwNUsmnSIDZW16oV9dukAeris28nZQ ezKs/Rwah+g== X-Received: from dlc21.prod.google.com ([2002:a05:7022:395:b0:144:bebe:a64]) (user=irogers job=prod-delivery.src-stubby-dispatcher) by 2002:a17:90b:1811:b0:39d:e54c:8658 with SMTP id 98e67ed59e1d1-39e1e250230mr16608214a91.5.1789627392115; Wed, 16 Sep 2026 23:43:12 -0700 (PDT) Date: Wed, 16 Sep 2026 23:42:35 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <80beb357d1eebd77f99982fbf4741e8b5fefbafe.1789626978.git.irogers@google.com> Subject: [PATCH v1 13/13] perf test uprobe_from_different_cu: Scope probe name to PID From: Ian Rogers To: Arnaldo Carvalho de Melo , Namhyung Kim Cc: Peter Zijlstra , Ingo Molnar , Jiri Olsa , Adrian Hunter , James Clark , linux-perf-users@vger.kernel.org, linux-kernel@vger.kernel.org, Ian Rogers Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" The test builds a binary in a per-run temporary directory and probes its foo function. The directory name is unique, but perf probe derives the event name from the probed function and the group name from the binary's basename, so every run registers the same probe_testfile:foo event. Running the test concurrently with itself, as 'perf test -r3' does, therefore fails in all but one of the runs with: Error: event "foo" already exists. Hint: Remove existing event by 'perf probe -d' and a losing run's cleanup goes on to delete the winning run's probe out from under it. Name the event after the pid, foo_$$, so that parallel runs no longer collide. This lets the test stay in the parallel pass rather than having to be marked (exclusive). Assisted-by: Antigravity:gemini-3.1-pro Signed-off-by: Ian Rogers --- .../perf/tests/shell/test_uprobe_from_different_cu.sh | 11 +++++++++-- 1 file changed, 9 insertions(+), 2 deletions(-) diff --git a/tools/perf/tests/shell/test_uprobe_from_different_cu.sh b/tool= s/perf/tests/shell/test_uprobe_from_different_cu.sh index 7adf9755d6de..47c99d93436b 100755 --- a/tools/perf/tests/shell/test_uprobe_from_different_cu.sh +++ b/tools/perf/tests/shell/test_uprobe_from_different_cu.sh @@ -18,12 +18,19 @@ fi =20 temp_dir=3D$(mktemp -d /tmp/perf-uprobe-different-cu-sh.XXXXXXXXXX) =20 +# The name of the uprobe added and removed below. The probe is placed on +# ${temp_dir}/testfile, but perf probe derives the event name from the pro= bed +# function and the group name from the binary's basename, so every run wou= ld +# otherwise share one probe_testfile:foo event, and a concurrent run would +# fail with 'event "foo" already exists'. Scope the event name to the pid. +probe_name=3D"foo_$$" + cleanup() { trap - EXIT TERM INT if [[ "${temp_dir}" =3D~ ^/tmp/perf-uprobe-different-cu-sh.*$ ]]; then echo "--- Cleaning up ---" - perf probe -x ${temp_dir}/testfile -d foo || true + perf probe -x ${temp_dir}/testfile -d ${probe_name} || true rm -f "${temp_dir}/"* rmdir "${temp_dir}" fi @@ -84,6 +91,6 @@ gcc -g -Og -c ${temp_dir}/testfile-main.c -o ${temp_dir}/= testfile-main.o gcc -g -Og -o ${temp_dir}/testfile ${temp_dir}/testfile-foo.o ${temp_dir}/= testfile-main.o =20 perf probe -x ${temp_dir}/testfile --funcs foo | grep "foo" -perf probe -x ${temp_dir}/testfile foo +perf probe -x ${temp_dir}/testfile ${probe_name}=3Dfoo =20 cleanup --=20 2.55.0.1082.g2b9226bbc0-goog