From nobody Sat Jul 25 05:23:13 2026 Received: from mail-pg1-f180.google.com (mail-pg1-f180.google.com [209.85.215.180]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7EC193FBEBB for ; Fri, 17 Jul 2026 13:02:45 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.180 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293366; cv=none; b=qdfb32MSlpok9ZUV8G3X3WBoS0kTJHrret69/CqDRYXra3baLMOqKryCgPRagrlEKd89xcIJvx8xgGmCfjb3oGjCiZfipVRfgYX2yfW4EL6L0AP52oeh9isqYxu05JBScHHsXy4aI1uWTrkaUabDY5kZVFJJL9ktU5cimwJPK80= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293366; c=relaxed/simple; bh=PlLf3uatOB4V7qqPGolTM2wzfx6QOLYYVgfstlgW/F4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=YsPkKygnphyQeou5YWi1+oo/bN6OrT8ZnAbByd6Mj2F9XIkVV/y9dIe3IJPfh7uTQQmpfyabvM1JskbJpszsUzxeKKH+/LMb1YvtETZX+WP3D7mujmMPTUJQc02O2UzkRq1KzlSBbGP6Q0y5ibbslDTWZUk3KSMThxWEffSwCL8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=pV8RY9KL; arc=none smtp.client-ip=209.85.215.180 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="pV8RY9KL" Received: by mail-pg1-f180.google.com with SMTP id 41be03b00d2f7-c9fe3c9bd5fso999308a12.0 for ; Fri, 17 Jul 2026 06:02:45 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784293365; x=1784898165; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=/HcjDBHsn65ZCZzOohF40VNQnTu2vZZX71z4vh7Fe7A=; b=pV8RY9KL5NRyhBdnkl8IhflMe9EB9vB18xcJdNOZOzs5/28McoDBlIreED8PiraNl3 icWwx+gQYviKkZd21TqwHoP+bSrbsb3vuXqR1qssn/T5DfmSHm5V5LOQl4SFRj4SxuoU 88+KfXf//S88mNSOGQW462CKES+QGc4kBiWNjLBpUhVZk5oRs8NBGAmFyQJvqoUF7puR J298s1Xg/px1lAHTjezol7Blq9gng0HEFsxox/x+m6ptq/OAY0sfEWcFLJ3kLEjDrEtU QJVV8m1mKlUwuOVD7RGoQ7dJC7ruqNXR9Da+eOjHvH4nCvffbaiacf5LCdEc8v4CTNhP GMPA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784293365; x=1784898165; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=/HcjDBHsn65ZCZzOohF40VNQnTu2vZZX71z4vh7Fe7A=; b=MqDywTjTD/YeLk2hlrgD5DcK74DlazZvcYnNfoW0syS4zrx+00wmVuMgbhqHQLqOD1 fxhBHuUbKIAaC3pZjnEydvlrUhBIm83TbplrTwhIvfwlxmjPkSZr0cPMVAnR937Bmend yqUooxLGSS/dNikq/bIQRN5IsqBZYrnk6UI7Jrp8mJPsD4XTCqOZ4l68aXeYlUT+GKeW Yjeusppkng3VCXCJr2bPJGGLDvcCwbT+TgI65CoApi+RZrwPke7W6eXFjH/1bfna26yY Zgjv1FzBQX36zwWWhyELyZdzpL2p5Qoz8JuNVMOnQuzLYLLiybvbj2YhbRYxJy8gR3ub s6ew== X-Forwarded-Encrypted: i=1; AHgh+RqXgr3uzd7G+eFIKGcAFTPNNVeddXM+3b5r4bdIFpGTYnwUpvBZVLoSRZa9vEffVRNA7vlSEAsp1Ov2xmM=@vger.kernel.org X-Gm-Message-State: AOJu0YyqPomlwqfHugbtNUi2IdBNq2qE+ffJk431Eitbxa24vVyEf5t6 8lG4BU77VSiyJ6JNU5SdzGkQrteLmEdvVwAp8syDkJs+e9oVb5MJqvdT X-Gm-Gg: AfdE7ckBNN5angrt5d1JleKG4lsmskuzlMKhqhqLnu0KSlleZXe2KzZBW3V2jh2PMFC F/iqsrqxJmEYC2dpHSO34OYJhql3wdxGifL9DCcyZ4x6P+ED3AwEwMi6hpaX/mwbJwId6ZHqgy7 a/6TLEhTWHU5prFQX9esov4jsbWn5docG4kJeLPMDMCaehT9KS0s51aFsRRiM0PQ8dkypdTyYnV 4q+9Az9UWNNU5+IewtHvSaGuO02C7aLdB1QW9Gfma9au8kpvGWqAAYz7xP3tmI07cXdl/5q/bJv d0D2e39AL27krz2/mkoNJqyJoCucph1Ji+7aupShQt6545xST1oOQabfOFGE0vELm/KW9Jqe7xd OuaTkdaC2QU1Z6Oh0gftRe/E5nwmRnR70aGG1w+f6X20TACnbl2XL/bJN/LIDsu21xWyRWmlKFw +K2w== X-Received: by 2002:a17:90b:4e85:b0:380:86d8:8162 with SMTP id 98e67ed59e1d1-38e3d299be4mr7567812a91.18.1784293364383; Fri, 17 Jul 2026 06:02:44 -0700 (PDT) Received: from localhost ([144.24.58.22]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-3142a1de301sm7390369eec.24.2026.07.17.06.02.42 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 17 Jul 2026 06:02:42 -0700 (PDT) From: Jinchao Wang To: Andrew Morton , Peter Zijlstra , Thomas Gleixner , Steven Rostedt , Masami Hiramatsu Cc: Ingo Molnar , Borislav Petkov , Dave Hansen , "H . Peter Anvin" , x86@kernel.org, Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Mathieu Desnoyers , David Hildenbrand , Jonathan Corbet , Matthew Wilcox , Alan Stern , Randy Dunlap , Alexander Potapenko , Marco Elver , Mike Rapoport , linux-kernel@vger.kernel.org, linux-mm@kvack.org, linux-trace-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org, linux-doc@vger.kernel.org, Jinchao Wang Subject: [RFC PATCH v2 01/13] arch: add HAVE_REINSTALL_HW_BREAKPOINT Date: Fri, 17 Jul 2026 09:02:27 -0400 Message-ID: <20260717130227.1901488-1-wangjinchao600@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260717125023.1895892-1-wangjinchao600@gmail.com> References: <20260717125023.1895892-1-wangjinchao600@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Some architectures can update the address, length or type of an installed hardware breakpoint in place, without releasing and re-reserving its slot. Add an opt-in Kconfig symbol so generic code can offer such an operation on architectures that implement arch_reinstall_hw_breakpoint(). This is a prerequisite for KWatch, which re-points preallocated per-CPU breakpoints from atomic context, where the register/release path (which may sleep and rebalances slot constraints) cannot be used. Signed-off-by: Jinchao Wang --- arch/Kconfig | 10 ++++++++++ 1 file changed, 10 insertions(+) diff --git a/arch/Kconfig b/arch/Kconfig index fa7507ac8e13..41b3784e0ddd 100644 --- a/arch/Kconfig +++ b/arch/Kconfig @@ -457,6 +457,16 @@ config HAVE_MIXED_BREAKPOINTS_REGS Select this option if your arch implements breakpoints under the latter fashion. =20 +config HAVE_REINSTALL_HW_BREAKPOINT + bool + depends on HAVE_HW_BREAKPOINT + help + Depending on the arch implementation of hardware breakpoints, + some of them are able to update the breakpoint configuration + without release and reserve the hardware breakpoint register. + What configuration is able to update depends on hardware and + software implementation. + config HAVE_USER_RETURN_NOTIFIER bool =20 --=20 2.53.0 From nobody Sat Jul 25 05:23:13 2026 Received: from mail-pj1-f44.google.com (mail-pj1-f44.google.com [209.85.216.44]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AD8A6342C98 for ; Fri, 17 Jul 2026 13:03:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.44 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293391; cv=none; b=TsbFlkxPcMDXnbm4jDXBUcJeNNRxEvjpR3frLMu89BptLwm0pHxaqm5/UH1Qf2m5bE5Bv7rkPuGT+NQyazXKrAIoMww7DOYK0AN1d47yHCEaXgJQa9YtCS289j6Ksrduotz72sEMW6qZa+qwQFKP6AB4cA9YtJULd3cd5sCwIz4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293391; c=relaxed/simple; bh=McJvpjFtHEtr6No4XI/yr1DXNKrvDdPTiwn+hIU5wl0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=kJcMh8HOAVshveDkmhrMdh25UYpqEwcyd2zkPAMW4nHst6z1LQvq/K18LoCcldBuCpYFVLbQIFHoI/f5vRjgK/SNSuxyYIUx7cZFPHO7CghLAk/iMagWeM+zNxC6HuUtRgrZUzp3N5JQBRSKsKGlUUeuuT69awHEVkJ5Qk6COT8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=WSC/ksbv; arc=none smtp.client-ip=209.85.216.44 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="WSC/ksbv" Received: by mail-pj1-f44.google.com with SMTP id 98e67ed59e1d1-38e3617ba36so2201145a91.3 for ; Fri, 17 Jul 2026 06:03:09 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784293389; x=1784898189; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=4gbLfd0R20qfhPC42TsD6whBibdrC59ZB/Kgr873YXk=; b=WSC/ksbvHku704vyMHWjcIfIVjUhKDmkKXrSJI9qL/MUYArt116JWMfYJWLKtCBhZ6 oVRChodDrO3+ZPv7nPX6j1rbRey/qFCUQJG8YUqfmOs8Kmwn8o+dwIvlhyazoIY5bAPC cZjc4VwL+A6ZS40bWj4L9nxzVcH0h9RQXUgD13SMW4/+T88+jVKPyHzeXVIzu5agk2xI v9YIaVZLXeghlphyLFw3ncpY8Xi1fvY3uM+UYMuOxIaSLhPLtaIzArz+QNa8WCEfQXNv uKzezAJZFQwBLxmeJcyiwNiX/Z9wMDm+BKl4JRbQ40HFWnCJulSyOcmMbWESmi1xDk2F rWXQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784293389; x=1784898189; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=4gbLfd0R20qfhPC42TsD6whBibdrC59ZB/Kgr873YXk=; b=JgCO1oeJVkFOUouayMaVBPVKfkDCoPxOnT563niLD7gbWMj9iHzdelTLGzZrKMZmZi gquZOnwNx/L4g8wQKVBow6Ckwn/9u40/PAKwp9cSgzjm9d2b+JL0hAed/4896W5VuMNZ BuyVuOJ279+P6j/0jSs0ch6RVcPK2eYj0k+jsw43x1A3SqP0c/imve6NK8ATeJNt1BYB XqJNAy3Z0TwRP18MqoHy11CE2qNC4AmUIAxFdyvVQTJJ1dUYxkdCU+gR/lDttIOtrCcs MPS2MCr3c/Er3l/6ppRV8YU0aswaPa+4YEWhkAIdajPA818wWfjkXfIPOOIQ1yIY+m3L wLMg== X-Forwarded-Encrypted: i=1; AHgh+RpDMfa3M/2uEZmR7+kwq+6JHaayAaKxHFJ8puX+Ebe7i/NfJo4N8g/lRONI4wqDunMP3jrsgqzjLn037jI=@vger.kernel.org X-Gm-Message-State: AOJu0YwKE34eMq3OHy3SDyrYI60lNAPJm1O/VFWC7E+tA/c3iXgxcvET YZZQghIM5525QntMNgaeAlHPBf1rflowuVXiz1KTg9j9etXXdQeRTr3W X-Gm-Gg: AfdE7cnNbj17cMF+BSKW6bUbtstwNeIgZpT0/XlLR4EEz7ZH8s+Xj1b3EIjHfffe9ZS cXboVwrTQbYRssdrqV3oa1tNdiqPfBS5Jz85TkAIFPLmlL3JyVt+ieM7nD2kkfwhTcXMTC0UvbA lI/9V0hckUMWDIL+ltYGSfjZB2upXl7MRIozeQZN2ZIi4fzzBeZIY9vW0MF9vaB05CQSBdq517V S+LXsrb5olaD2wysKfXFQzAswsBStgEwxKYG8OsfrKMXigJYwG9Me+UDeIm2AmAZVyUWCBhHjPR EGY1PUvG2C1UvRtZdDyHSmdjLNLK++h1npr7Wu97HiGlJayL22fvjzIo1F0EzxBbFBByDs9IBCe dfovbWLSEMhqGHSyT6OpzG1Kwk5sJ6kEgoWeVP2SwmeznErXSJA1TobWxJQe8SYcwcQ5jXSvrvm 6i2g== X-Received: by 2002:a17:90b:1c81:b0:387:e0bb:57f1 with SMTP id 98e67ed59e1d1-38e4b56bd45mr2610926a91.34.1784293388471; Fri, 17 Jul 2026 06:03:08 -0700 (PDT) Received: from localhost ([144.24.58.22]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-31429ff04absm6773202eec.8.2026.07.17.06.03.06 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 17 Jul 2026 06:03:07 -0700 (PDT) From: Jinchao Wang To: Andrew Morton , Peter Zijlstra , Thomas Gleixner , Steven Rostedt , Masami Hiramatsu Cc: Ingo Molnar , Borislav Petkov , Dave Hansen , "H . Peter Anvin" , x86@kernel.org, Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Mathieu Desnoyers , David Hildenbrand , Jonathan Corbet , Matthew Wilcox , Alan Stern , Randy Dunlap , Alexander Potapenko , Marco Elver , Mike Rapoport , linux-kernel@vger.kernel.org, linux-mm@kvack.org, linux-trace-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org, linux-doc@vger.kernel.org, Jinchao Wang Subject: [RFC PATCH v2 02/13] x86/hw_breakpoint: Unify breakpoint install/uninstall Date: Fri, 17 Jul 2026 09:02:51 -0400 Message-ID: <20260717130251.1901695-1-wangjinchao600@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260717125023.1895892-1-wangjinchao600@gmail.com> References: <20260717125023.1895892-1-wangjinchao600@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Consolidate breakpoint management to reduce code duplication. The diffstat was misleading, so the stripped code size is compared instead. After refactoring, it is reduced from 11976 bytes to 11448 bytes on my x86_64 system built with clang. This also makes it easier to introduce arch_reinstall_hw_breakpoint(). In addition, including linux/types.h to fix a missing build dependency. Signed-off-by: Jinchao Wang Reviewed-by: Masami Hiramatsu (Google) --- arch/x86/include/asm/hw_breakpoint.h | 6 ++ arch/x86/kernel/hw_breakpoint.c | 153 ++++++++++++++++----------- 2 files changed, 95 insertions(+), 64 deletions(-) diff --git a/arch/x86/include/asm/hw_breakpoint.h b/arch/x86/include/asm/hw= _breakpoint.h index 0bc931cd0698..aa6adac6c3a2 100644 --- a/arch/x86/include/asm/hw_breakpoint.h +++ b/arch/x86/include/asm/hw_breakpoint.h @@ -5,6 +5,7 @@ #include =20 #define __ARCH_HW_BREAKPOINT_H +#include =20 /* * The name should probably be something dealt in @@ -18,6 +19,11 @@ struct arch_hw_breakpoint { u8 type; }; =20 +enum bp_slot_action { + BP_SLOT_ACTION_INSTALL, + BP_SLOT_ACTION_UNINSTALL, +}; + #include #include #include diff --git a/arch/x86/kernel/hw_breakpoint.c b/arch/x86/kernel/hw_breakpoin= t.c index f846c15f21ca..76886467708b 100644 --- a/arch/x86/kernel/hw_breakpoint.c +++ b/arch/x86/kernel/hw_breakpoint.c @@ -49,7 +49,6 @@ static DEFINE_PER_CPU(unsigned long, cpu_debugreg[HBP_NUM= ]); */ static DEFINE_PER_CPU(struct perf_event *, bp_per_reg[HBP_NUM]); =20 - static inline unsigned long __encode_dr7(int drnum, unsigned int len, unsigned int type) { @@ -86,96 +85,122 @@ int decode_dr7(unsigned long dr7, int bpnum, unsigned = *len, unsigned *type) } =20 /* - * Install a perf counter breakpoint. - * - * We seek a free debug address register and use it for this - * breakpoint. Eventually we enable it in the debug control register. - * - * Atomic: we hold the counter->ctx->lock and we only handle variables - * and registers local to this cpu. + * We seek a slot and change it or keep it based on the action. + * Returns slot number on success, negative error on failure. + * Must be called with IRQs disabled. */ -int arch_install_hw_breakpoint(struct perf_event *bp) +static int manage_bp_slot(struct perf_event *bp, enum bp_slot_action actio= n) { - struct arch_hw_breakpoint *info =3D counter_arch_bp(bp); - unsigned long *dr7; - int i; - - lockdep_assert_irqs_disabled(); + struct perf_event *old_bp; + struct perf_event *new_bp; + int slot; + + switch (action) { + case BP_SLOT_ACTION_INSTALL: + old_bp =3D NULL; + new_bp =3D bp; + break; + case BP_SLOT_ACTION_UNINSTALL: + old_bp =3D bp; + new_bp =3D NULL; + break; + default: + return -EINVAL; + } =20 - for (i =3D 0; i < HBP_NUM; i++) { - struct perf_event **slot =3D this_cpu_ptr(&bp_per_reg[i]); + for (slot =3D 0; slot < HBP_NUM; slot++) { + struct perf_event **curr =3D this_cpu_ptr(&bp_per_reg[slot]); =20 - if (!*slot) { - *slot =3D bp; - break; + if (*curr =3D=3D old_bp) { + *curr =3D new_bp; + return slot; } } =20 - if (WARN_ONCE(i =3D=3D HBP_NUM, "Can't find any breakpoint slot")) - return -EBUSY; + if (old_bp) { + WARN_ONCE(1, "Can't find matching breakpoint slot"); + return -EINVAL; + } =20 - set_debugreg(info->address, i); - __this_cpu_write(cpu_debugreg[i], info->address); + WARN_ONCE(1, "No free breakpoint slots"); + return -EBUSY; +} =20 - dr7 =3D this_cpu_ptr(&cpu_dr7); - *dr7 |=3D encode_dr7(i, info->len, info->type); +static void setup_hwbp(struct arch_hw_breakpoint *info, int slot, bool ena= ble) +{ + unsigned long dr7; + + set_debugreg(info->address, slot); + __this_cpu_write(cpu_debugreg[slot], info->address); + + dr7 =3D this_cpu_read(cpu_dr7); + if (enable) + dr7 |=3D encode_dr7(slot, info->len, info->type); + else + dr7 &=3D ~__encode_dr7(slot, info->len, info->type); =20 /* - * Ensure we first write cpu_dr7 before we set the DR7 register. - * This ensures an NMI never see cpu_dr7 0 when DR7 is not. + * Enabling: + * Ensure we first write cpu_dr7 before we set the DR7 register. + * This ensures an NMI never see cpu_dr7 0 when DR7 is not. */ + if (enable) + this_cpu_write(cpu_dr7, dr7); + barrier(); =20 - set_debugreg(*dr7, 7); - if (info->mask) - amd_set_dr_addr_mask(info->mask, i); + set_debugreg(dr7, 7); =20 - return 0; + /* + * Always push the address mask, even when clearing it (info->mask =3D=3D= 0): + * a REINSTALL from a masked range breakpoint to an exact one must drop + * the stale mask, or the CPU keeps matching the wider range. + * amd_set_dr_addr_mask() is a no-op without X86_FEATURE_BPEXT and skips + * redundant MSR writes, so the unconditional call is cheap. + */ + amd_set_dr_addr_mask(enable ? info->mask : 0, slot); + + /* + * Disabling: + * Ensure the write to cpu_dr7 is after we've set the DR7 register. + * This ensures an NMI never see cpu_dr7 0 when DR7 is not. + * The barrier keeps the compiler from reordering the two: native + * set_debugreg() has no memory clobber of its own. + */ + if (!enable) { + barrier(); + this_cpu_write(cpu_dr7, dr7); + } } =20 /* - * Uninstall the breakpoint contained in the given counter. - * - * First we search the debug address register it uses and then we disable - * it. - * - * Atomic: we hold the counter->ctx->lock and we only handle variables - * and registers local to this cpu. + * find suitable breakpoint slot and set it up based on the action */ -void arch_uninstall_hw_breakpoint(struct perf_event *bp) +static int arch_manage_bp(struct perf_event *bp, enum bp_slot_action actio= n) { - struct arch_hw_breakpoint *info =3D counter_arch_bp(bp); - unsigned long dr7; - int i; + struct arch_hw_breakpoint *info; + int slot; =20 lockdep_assert_irqs_disabled(); =20 - for (i =3D 0; i < HBP_NUM; i++) { - struct perf_event **slot =3D this_cpu_ptr(&bp_per_reg[i]); - - if (*slot =3D=3D bp) { - *slot =3D NULL; - break; - } - } - - if (WARN_ONCE(i =3D=3D HBP_NUM, "Can't find any breakpoint slot")) - return; + slot =3D manage_bp_slot(bp, action); + if (slot < 0) + return slot; =20 - dr7 =3D this_cpu_read(cpu_dr7); - dr7 &=3D ~__encode_dr7(i, info->len, info->type); + info =3D counter_arch_bp(bp); + setup_hwbp(info, slot, action !=3D BP_SLOT_ACTION_UNINSTALL); =20 - set_debugreg(dr7, 7); - if (info->mask) - amd_set_dr_addr_mask(0, i); + return 0; +} =20 - /* - * Ensure the write to cpu_dr7 is after we've set the DR7 register. - * This ensures an NMI never see cpu_dr7 0 when DR7 is not. - */ - barrier(); +int arch_install_hw_breakpoint(struct perf_event *bp) +{ + return arch_manage_bp(bp, BP_SLOT_ACTION_INSTALL); +} =20 - this_cpu_write(cpu_dr7, dr7); +void arch_uninstall_hw_breakpoint(struct perf_event *bp) +{ + arch_manage_bp(bp, BP_SLOT_ACTION_UNINSTALL); } =20 static int arch_bp_generic_len(int x86_len) --=20 2.53.0 From nobody Sat Jul 25 05:23:13 2026 Received: from mail-pj1-f53.google.com (mail-pj1-f53.google.com [209.85.216.53]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 25BEA3FBEC4 for ; Fri, 17 Jul 2026 13:03:33 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.53 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293415; cv=none; b=Hr5Z353FDVs454LkriNKNzLRkzKIf6PWKHbvzYHy1W7ppIMyOrbFttLYQC7pg9QS+sNRmqUdT1NQ+9jRYP/q48pfbQTdrqpB3gfu6gt1S9d/vIUZsdT2MyB7UFjcW3kG84KJBGtv6ay8Iga22eZnWN8HpH2wH/9KWvGgURq0w+8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293415; c=relaxed/simple; bh=BE0O/mKap0ViHNplNVUpYAszwuhBBa7mfDYWdw1MMxs=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=vDz1HpMGO/jrV35eOG8dgRrfdmdbTZmuhGMuiIyzboGYhoG+cxlfwKqp3+teKNumQ34nnjnRYWsanS+1+M+2QnHszOJ241v+6dVu6LzzLHeMs/gcH/ZhLHsyHdsFUvJpW3w7jOl2D2l0ZwNvkfKAFZ8VZ8T4F5LZKnQJDYgyTIw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=VcjSTiw4; arc=none smtp.client-ip=209.85.216.53 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="VcjSTiw4" Received: by mail-pj1-f53.google.com with SMTP id 98e67ed59e1d1-38e42560ebcso1266853a91.1 for ; Fri, 17 Jul 2026 06:03:33 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784293413; x=1784898213; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=usfqleQl6zep7nG+d2WvhlRN371v3Tz2fhzwYC0DlME=; b=VcjSTiw45OYWBy2CgheicZB67Fu4KqTB4zgTPyE8gGyA/pnOoAUbF/V7MA9QSLwh68 trA4/OEbiCCi6Wu6XzXbLPHeRkW3otkukuKh3juosybOnzkg42Y2GnNJvAYtiYe6MHtD gnbly7PvGLhLIHKGJLEFGvEDF8fFKmom7l890cs3bbFxvxLZcj0qacoGQjLI7ORJ/HPT L99oPKBDAj25wGnV7piRDTYZceYSiLAoZ1nS8nZjOBjIjkHQH0ZSWQJzobZ9D5ATfnuj 4c0adUd08fKG5BakDyT0Pnn4mlii6zv5AaTtOT18OCmXXFN51lHqbj+bl4sSP8axq517 Eu0A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784293413; x=1784898213; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=usfqleQl6zep7nG+d2WvhlRN371v3Tz2fhzwYC0DlME=; b=bqqppqFAfxP15yLWP/UOSZ8TqJR4nJD01bvplKPo2iKhUHym2jSp08F+qnOJdxyeG4 sTlgHrDwkedx1jSGPVNatYs0ywjSLy3Dd9uGh6BHIZVWzSxhWgHc3uIxn9DPMfSu/36l Kmi6eQwvzUU3HFtaxRT349dr5C9ZCKaDGfD5wI6mCALW4Eg1ny/KVvwVedyF5dqlLzfH AHwUyUPi8ZHx7YY/R+1syNI3Hv/0gxACCpqx08fU/FfwLO5XW/QfNKegcgcPzI/xRk/x 47ljj12CMPpUrwjU/07VTlEv50Cugn+Lz9nkQTDM4TM2ZtsvgEz4DSXDMd33fBoA94AZ G+6A== X-Forwarded-Encrypted: i=1; AHgh+Rps7wbpR5SgtPJL49e3gIJybNZRZF2DIXnkXZnCTJUQhxnYSxivgvnp/syG5318FZ+Vdku5cTzW4BDqoCY=@vger.kernel.org X-Gm-Message-State: AOJu0Yw5ysrQHl4VJwO1hoi9cah004SFs1hARBkXZvphw/MWOlSjTcyl sdSzgb+kGaGuCo+dmKB7wtoadoZBaqRp081HMZ1vSRTLEQ72jtcECyMc X-Gm-Gg: AfdE7clxR3Wb4i9ONbLBk+6U6EMb874uTtGoM9DMPWBgWTp+4B9nMexYID1WSTgZdkU jjbH3X0+kCYBqG/yixBRAvJw/lvtOKMRy5qRbcEGgNitPVI5c/TaVJiY7y0lMrHE50yg03io+09 kdlsVhhFj9tDO4UGJKNILnGoOx08HKu6truWy7Jw8UwcnMUspHbI5AHm/V0rppiPZCbdThVLF4B XaTwNRral44mJtYv710WPDV1B4BFODXj0XSLqf4KyZlJEcAjUGU2Y6O4zKaGg7m4OxSRyBmKn6b W51xhMcOOmpRv/ErbZOADO0LBn2v2izG4+zHx80uBUwvW6JGjzlHE0tJ+uVeQrvZIx49jNpmDg8 S1+7btGur513WXP6/FKd5jZYcVZtHgjv9IY5oEuk7CzHm3sMYoUFqXdbLBU8PVP30WXDFH/jjBi N4Ow== X-Received: by 2002:a17:90a:ec88:b0:38d:adae:4866 with SMTP id 98e67ed59e1d1-38e4b516544mr2632403a91.21.1784293413065; Fri, 17 Jul 2026 06:03:33 -0700 (PDT) Received: from localhost ([144.24.58.22]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-3142a1dde81sm13705289eec.21.2026.07.17.06.03.31 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 17 Jul 2026 06:03:31 -0700 (PDT) From: Jinchao Wang To: Andrew Morton , Peter Zijlstra , Thomas Gleixner , Steven Rostedt , Masami Hiramatsu Cc: Ingo Molnar , Borislav Petkov , Dave Hansen , "H . Peter Anvin" , x86@kernel.org, Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Mathieu Desnoyers , David Hildenbrand , Jonathan Corbet , Matthew Wilcox , Alan Stern , Randy Dunlap , Alexander Potapenko , Marco Elver , Mike Rapoport , linux-kernel@vger.kernel.org, linux-mm@kvack.org, linux-trace-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org, linux-doc@vger.kernel.org, Jinchao Wang Subject: [RFC PATCH v2 03/13] x86/hw_breakpoint: Add arch_reinstall_hw_breakpoint Date: Fri, 17 Jul 2026 09:03:15 -0400 Message-ID: <20260717130315.1901903-1-wangjinchao600@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260717125023.1895892-1-wangjinchao600@gmail.com> References: <20260717125023.1895892-1-wangjinchao600@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" The new arch_reinstall_hw_breakpoint() function can be used in an atomic context, unlike the more expensive free and re-allocation path. This allows callers to efficiently re-establish an existing breakpoint, and x86 advertises the capability via HAVE_REINSTALL_HW_BREAKPOINT. Since a REINSTALL may change bp_len, setup_hwbp() must clear the slot's stale len/type and enable bits in DR7 before re-encoding: OR-merging the new encoding over the old one would keep the CPU watching with the stale width (verified in QEMU by reading DR7 after re-arming watch_len=3D1 over a len8 breakpoint: 0x999906aa merged without the clearing, 0x199906aa with it). Signed-off-by: Jinchao Wang --- arch/x86/Kconfig | 1 + arch/x86/include/asm/hw_breakpoint.h | 2 ++ arch/x86/kernel/hw_breakpoint.c | 16 ++++++++++++++-- 3 files changed, 17 insertions(+), 2 deletions(-) diff --git a/arch/x86/Kconfig b/arch/x86/Kconfig index bdad90f210e4..5be698db0241 100644 --- a/arch/x86/Kconfig +++ b/arch/x86/Kconfig @@ -246,6 +246,7 @@ config X86 select HAVE_FUNCTION_TRACER select HAVE_GCC_PLUGINS select HAVE_HW_BREAKPOINT + select HAVE_REINSTALL_HW_BREAKPOINT select HAVE_IOREMAP_PROT select HAVE_IRQ_EXIT_ON_IRQ_STACK if X86_64 select HAVE_IRQ_TIME_ACCOUNTING diff --git a/arch/x86/include/asm/hw_breakpoint.h b/arch/x86/include/asm/hw= _breakpoint.h index aa6adac6c3a2..c22cc4e87fc5 100644 --- a/arch/x86/include/asm/hw_breakpoint.h +++ b/arch/x86/include/asm/hw_breakpoint.h @@ -21,6 +21,7 @@ struct arch_hw_breakpoint { =20 enum bp_slot_action { BP_SLOT_ACTION_INSTALL, + BP_SLOT_ACTION_REINSTALL, BP_SLOT_ACTION_UNINSTALL, }; =20 @@ -65,6 +66,7 @@ extern int hw_breakpoint_exceptions_notify(struct notifie= r_block *unused, =20 =20 int arch_install_hw_breakpoint(struct perf_event *bp); +int arch_reinstall_hw_breakpoint(struct perf_event *bp); void arch_uninstall_hw_breakpoint(struct perf_event *bp); void hw_breakpoint_pmu_read(struct perf_event *bp); void hw_breakpoint_pmu_unthrottle(struct perf_event *bp); diff --git a/arch/x86/kernel/hw_breakpoint.c b/arch/x86/kernel/hw_breakpoin= t.c index 76886467708b..e2dbd43d8e39 100644 --- a/arch/x86/kernel/hw_breakpoint.c +++ b/arch/x86/kernel/hw_breakpoint.c @@ -100,6 +100,10 @@ static int manage_bp_slot(struct perf_event *bp, enum = bp_slot_action action) old_bp =3D NULL; new_bp =3D bp; break; + case BP_SLOT_ACTION_REINSTALL: + old_bp =3D bp; + new_bp =3D bp; + break; case BP_SLOT_ACTION_UNINSTALL: old_bp =3D bp; new_bp =3D NULL; @@ -134,10 +138,13 @@ static void setup_hwbp(struct arch_hw_breakpoint *inf= o, int slot, bool enable) __this_cpu_write(cpu_debugreg[slot], info->address); =20 dr7 =3D this_cpu_read(cpu_dr7); + /* + * Clear the slot's stale len/type and enable bits first: a REINSTALL + * with a different bp_len would otherwise OR-merge both encodings. + */ + dr7 &=3D ~__encode_dr7(slot, 0xf, 0); if (enable) dr7 |=3D encode_dr7(slot, info->len, info->type); - else - dr7 &=3D ~__encode_dr7(slot, info->len, info->type); =20 /* * Enabling: @@ -198,6 +205,11 @@ int arch_install_hw_breakpoint(struct perf_event *bp) return arch_manage_bp(bp, BP_SLOT_ACTION_INSTALL); } =20 +int arch_reinstall_hw_breakpoint(struct perf_event *bp) +{ + return arch_manage_bp(bp, BP_SLOT_ACTION_REINSTALL); +} + void arch_uninstall_hw_breakpoint(struct perf_event *bp) { arch_manage_bp(bp, BP_SLOT_ACTION_UNINSTALL); --=20 2.53.0 From nobody Sat Jul 25 05:23:13 2026 Received: from mail-pj1-f49.google.com (mail-pj1-f49.google.com [209.85.216.49]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B9A3A342C98 for ; Fri, 17 Jul 2026 13:03:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.49 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293440; cv=none; b=HWVJJI3MBIFMa1lTqu5xm5KIM0FHQp70ZqE8UdHtElkUplAOZKDEU5SZ2/5liUgXW5z87N+dgmyhyuJE6nnWLilAMvJUSrIFPZenu/7adSytCbIEUiwCFh4GDtjdOB1Z8MinMwkfSXfWwW1do/47QymydCvEP7wLL3ZxrV27i7Q= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293440; c=relaxed/simple; bh=lbvzkRlsaSDQiGshM+ijhIvb7bTedz8kwBQ4nTh8K9I=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=gzWWRklrAE1Y/CnlpUF6XYMMMz4iC7d1lPso8fo7kwMJl1eeW+F9zsq8ADIVxdNR0u3FpOndq8vCPpCJ8prvwlVJs4xWyOl1poQZrJX66eZ+S/6VrvfN5eNZ/5PaMqe+LDGQCri+ELSzuVDVEh/wr77QwW8pvGOXBsDqXKLz2Ao= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=D54CRgnn; arc=none smtp.client-ip=209.85.216.49 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="D54CRgnn" Received: by mail-pj1-f49.google.com with SMTP id 98e67ed59e1d1-38e08baf860so5360522a91.2 for ; Fri, 17 Jul 2026 06:03:58 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784293438; x=1784898238; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=7JZPn8H4RH/LnX03bYCbT81/FjSS1gT/hhL27IRJwnA=; b=D54CRgnnfKxCNTe/r4oVwUW7Y4H3h4oRAdG+BNslFp7936UZzXzbgMPKqNnB7DoJ5+ ZkbNrmjOx7dcamJSrBjzKQE77Q4vuIYucWMDF4JsvOhVpq2dsp18ugaRw8CO+naG6/GI PmC+eHeqwHZbis+kWaI5B67jexK4S21YR4bf7IS4vsQkKuDPL/a8Y/UNI/eSyWahWP/H qutXQuqz09/s/oaV2107nD+iaOMJpx4rVyI36UHCXVOrVbwLk/qjgbnudYb4BmNAGSUf JO5y0e+2P/MSQqLaebvah9LwF7NLyFhtScCwSONNfjbrHWcJqNhQJDj4/2L7LIceeFYd 92JQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784293438; x=1784898238; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=7JZPn8H4RH/LnX03bYCbT81/FjSS1gT/hhL27IRJwnA=; b=AEWWZrY8UQ+ujxNyj2Pw1tLD5N6bUu8/HUtvEZtHg/+/9pre9rvZMYXpUAjCJY3tl7 1ztIsUQDG5GMsqlGvbqhAlNYMHYH+on912HfFv7GaQ6EF4PxW1VoInCqkLRdZsPY/wLn ZA5Ucny/3xj5mCC2JI26MdJTOf3DDtr/pkc10fu1tN4pSIIOYglyfuWWLoCxjgNMFlgP gY6R9KAE8IktDcoKU6a8pqsxz1rfQBZau5bnKA1MwjpfOrjX1mfweE+4F/vmywylMRoE GdmUJCVRCwfEVVBwL9fPi3o8TlbWOdZTj4F1f95CRPglNWtYtffkHGH0h9x2OQ784bUv BPYA== X-Forwarded-Encrypted: i=1; AHgh+RrcNesv9DgPpIyM0I+Z+6pPh6/QSklW3kD3gEbfVzB5sHno+FDV9zAf2t1h0IqNd7j9y/0xdCN8Ads6Mag=@vger.kernel.org X-Gm-Message-State: AOJu0YxOcKuT0JIfL1Lc0fZlzOqQySis4IThQfiKwtoXPpSUKyVs2AIM wQn/Anv/Yvr44TgoVUj9LCiy6oF8AAT7UrBVrm/9EtrsmwzRgYnRuTAt X-Gm-Gg: AfdE7cnpcvUpmMBz7SM9ZsX9L+ux+XDKw5hr89BQChEuA9TGsDPn1Uk0PbnuxH6ilaQ 8kwbAB29BJK/aQNuCAzJCFklrxqBQiUfsQn+SvcA+kI0RlOo7lj34hWkd3GrdL7SPUhwOEt15Cn 0zC94zBd9jRGJLsXH8SAOkXxLuAm67NQKWwYt7d6y2p67NSe/0/BbGTsisMy/1IJm0DRNDNVTNg ih3pdljfMrn1dZEH5qpq0eh4CgXuBFkXb0Jqie9ScnYDq4U9KQoUhYXykzu9Co8W/GtK5zkV5qw 65bl+LXoYhWA3FW6k1HB1t63RU/RqsOVCKyTMSFoLROG3qdclnuznK49V4gCG8zf5lyLzRfuH5Q qOT+J6i6pDZPCQB2/cGV9VT6df3XTBIXTKO3Q5DzU6eoTj3rpe0j6XzTW+twQUgpCJ5UI3bUQ3j /BtA== X-Received: by 2002:a17:90a:e7c1:b0:36b:9e24:c692 with SMTP id 98e67ed59e1d1-38e4b53439emr2169596a91.20.1784293438047; Fri, 17 Jul 2026 06:03:58 -0700 (PDT) Received: from localhost ([144.24.58.22]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-3142a210571sm7374968eec.31.2026.07.17.06.03.56 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 17 Jul 2026 06:03:56 -0700 (PDT) From: Jinchao Wang To: Andrew Morton , Peter Zijlstra , Thomas Gleixner , Steven Rostedt , Masami Hiramatsu Cc: Ingo Molnar , Borislav Petkov , Dave Hansen , "H . Peter Anvin" , x86@kernel.org, Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Mathieu Desnoyers , David Hildenbrand , Jonathan Corbet , Matthew Wilcox , Alan Stern , Randy Dunlap , Alexander Potapenko , Marco Elver , Mike Rapoport , linux-kernel@vger.kernel.org, linux-mm@kvack.org, linux-trace-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org, linux-doc@vger.kernel.org, Jinchao Wang Subject: [RFC PATCH v2 04/13] HWBP: Add modify_wide_hw_breakpoint_local() API Date: Fri, 17 Jul 2026 09:03:40 -0400 Message-ID: <20260717130340.1902076-1-wangjinchao600@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260717125023.1895892-1-wangjinchao600@gmail.com> References: <20260717125023.1895892-1-wangjinchao600@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Masami Hiramatsu (Google)" Add modify_wide_hw_breakpoint_local() arch-wide interface which allows hwbp users to update watch address on-line. This is available if the arch supports CONFIG_HAVE_REINSTALL_HW_BREAKPOINT. Note that this allows to change the type only for compatible types, because it does not release and reserve the hwbp slot based on type. For instance, you can not change HW_BREAKPOINT_W to HW_BREAKPOINT_X. Assisted-by: Antigravity:gemini-3.5-flash Signed-off-by: Masami Hiramatsu (Google) Signed-off-by: Jinchao Wang --- include/linux/hw_breakpoint.h | 6 +++++ kernel/events/hw_breakpoint.c | 43 +++++++++++++++++++++++++++++++++++ 2 files changed, 49 insertions(+) diff --git a/include/linux/hw_breakpoint.h b/include/linux/hw_breakpoint.h index db199d653dd1..6754ffbee9ed 100644 --- a/include/linux/hw_breakpoint.h +++ b/include/linux/hw_breakpoint.h @@ -81,6 +81,9 @@ register_wide_hw_breakpoint(struct perf_event_attr *attr, perf_overflow_handler_t triggered, void *context); =20 +extern int modify_wide_hw_breakpoint_local(struct perf_event *bp, + struct perf_event_attr *attr); + extern int register_perf_hw_breakpoint(struct perf_event *bp); extern void unregister_hw_breakpoint(struct perf_event *bp); extern void unregister_wide_hw_breakpoint(struct perf_event * __percpu *cp= u_events); @@ -124,6 +127,9 @@ register_wide_hw_breakpoint(struct perf_event_attr *att= r, perf_overflow_handler_t triggered, void *context) { return NULL; } static inline int +modify_wide_hw_breakpoint_local(struct perf_event *bp, + struct perf_event_attr *attr) { return -EOPNOTSUPP; } +static inline int register_perf_hw_breakpoint(struct perf_event *bp) { return -ENOSYS; } static inline void unregister_hw_breakpoint(struct perf_event *bp) { } static inline void diff --git a/kernel/events/hw_breakpoint.c b/kernel/events/hw_breakpoint.c index 789add0c185a..f4709c892d67 100644 --- a/kernel/events/hw_breakpoint.c +++ b/kernel/events/hw_breakpoint.c @@ -888,6 +888,49 @@ void unregister_wide_hw_breakpoint(struct perf_event *= __percpu *cpu_events) } EXPORT_SYMBOL_GPL(unregister_wide_hw_breakpoint); =20 +/** + * modify_wide_hw_breakpoint_local - update breakpoint config for local CPU + * @bp: the hwbp perf event for this CPU + * @attr: the new attribute for @bp + * + * This does not release and reserve the slot of a HWBP; it just reuses the + * current slot on local CPU. So the users must update the other CPUs by + * themselves. + * Also, since this does not release/reserve the slot, this can not change= the + * type to incompatible type of the HWBP. + * Return err if attr is invalid or the CPU fails to update debug register + * for new @attr. + */ +#ifdef CONFIG_HAVE_REINSTALL_HW_BREAKPOINT +int modify_wide_hw_breakpoint_local(struct perf_event *bp, + struct perf_event_attr *attr) +{ + struct arch_hw_breakpoint info; + int ret; + + if (find_slot_idx(bp->attr.bp_type) !=3D find_slot_idx(attr->bp_type)) + return -EINVAL; + + ret =3D hw_breakpoint_arch_parse(bp, attr, &info); + if (ret) + return ret; + + *counter_arch_bp(bp) =3D info; + bp->attr.bp_addr =3D attr->bp_addr; + bp->attr.bp_type =3D attr->bp_type; + bp->attr.bp_len =3D attr->bp_len; + + return arch_reinstall_hw_breakpoint(bp); +} +#else +int modify_wide_hw_breakpoint_local(struct perf_event *bp, + struct perf_event_attr *attr) +{ + return -EOPNOTSUPP; +} +#endif +EXPORT_SYMBOL_GPL(modify_wide_hw_breakpoint_local); + /** * hw_breakpoint_is_used - check if breakpoints are currently used * --=20 2.53.0 From nobody Sat Jul 25 05:23:13 2026 Received: from mail-pj1-f52.google.com (mail-pj1-f52.google.com [209.85.216.52]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EA2B43FDC14 for ; Fri, 17 Jul 2026 13:04:23 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.52 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293465; cv=none; b=RCiLPj/zQtbpxLF12YzJ1ZBGIVK98en7siTg8cfnidep2lrJ3PAHe5wFSLX3lxP5tfOY0VMb0qLbKcEWxHIobIgLUXvilIwCpRfbgxRr9Uuob6rY4pdffwQHy8JI9FFho+VDnBx4qWv7HQZO4Wu5flxqpAfJHolRTGqdyL50vS8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293465; c=relaxed/simple; bh=TqoEQEUr0dwFy5/OGLtnT5TOw/H5doFE4/2j9BgHBdA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=K9nYwmYfHy3aKTAsdeLpsmFz83+4/txzJUuEAxRqA/kYZLAj63Tg+q8KKX6N2V421IpHx/ErybRHesq4uHb5tQp4ZL+rwwx+P447qOVoTV/acwoBgejTPB4cqs4MTsb3JuSBXgiicVgUOYNWRFuw9W8kLKSnkd3wX/C/tvAh7xQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=pWgrqcA2; arc=none smtp.client-ip=209.85.216.52 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="pWgrqcA2" Received: by mail-pj1-f52.google.com with SMTP id 98e67ed59e1d1-3811f512167so8022927a91.3 for ; Fri, 17 Jul 2026 06:04:23 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784293463; x=1784898263; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=oZ4T7oI/5myCntWJFqGEBzYfa7to/x/3/pee5fO6Ps8=; b=pWgrqcA2wTQzbEPMa1h+DlJpc/N9UfNaT8DhObK9T7GjXkj5lwrwMwMBEGIbSuM1EG nXciGCWfTyOLHxNJ9vks2qTsJf1/kf1swXCzsKzUBSfmGv495OMGlIIu89JEnCRxcN8c 1QsA7jl1C90uva1rPvfK6HE9DC004bIExN/mI/+H6j3WlkwJGQ0UolAoycsQ0B3fejLU fqqHCW1yiwETu07/nxqZrvTd7LKwX19ojdPY2mzhgIpS+4s7wmTvk+XA9pF7mFZIncEE EpPIVNHFcMgUzUYdeUTsImkv/XiJaVlmV7neIeAi+S+DSNnR42+TWLllcijZ956o+kEz EAYw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784293463; x=1784898263; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=oZ4T7oI/5myCntWJFqGEBzYfa7to/x/3/pee5fO6Ps8=; b=HchKlLl1HWNWDibFyQ8T4Hl1aelv8SY5TmUTlPW5HTghrnsx+QOAuxzlr5IBuzEVHv iXxnld3zv4Vs4j1W2mXoG0o9oxLds5bMq5PCMnDJYfCHwm4FXsl3RHM7YldVZ+gBNDZq WTvDH74R7H3RY6eBMDmAq/LQlEc1EqAdhh2gS66z+NgH32KUbtQVqwNU4zMBCJfW0Vvx r8fZR8eYCqIx/ZLKwpppqxei5juPBRM7rAq7MUbjtcLUGomDiwxhmzwX2549JYz1N1Jy ulf2HvDZMoNIBH2ZU7lj09gKLiotfr171BBH3vDyk92C7TEz5hpoRBfiznzy6mn303UP y/TQ== X-Forwarded-Encrypted: i=1; AHgh+RqVnsaL7Mrwds7tUQqMvIQHndIgDaIC6QZm3sPTL21zhCtv8nEfubjywFfISuxg4Y8WfrPZjyJ1GxCsm2I=@vger.kernel.org X-Gm-Message-State: AOJu0YxyrnaBKq489wlANaEMm+WNyFXrURTBAvh+naFVPAHMm/KgxEjb MvbrXm3KiFAeeT7VgwSYX5JDHhEHntVoXdA54weYBgb1dppOXzEg8p6V X-Gm-Gg: AfdE7clLwM8FMT5SjKw5ro9ToEq/1GaC5kB7jYajh3TQLiej5qg+6alyl69sdhBnIKN kAYcxteAjzK5qu8n2+n+Xw6UgWbyo139TOBqZzu/BebJ+8986jBmf7cKKCt1CgXvRIumpJ24mkk hmIXYxmk887K3qD2L9Z86PU+WbRTpzxLeL+jINKKeiYpL2Xpe2W17ZvvtH01t981xRJm1P0iiqF JdiQIKCBg6anTjEgKzeRG2FVUcgtdhlKpjWDyG18LpnBiB3EBGY5lsExOfKEDiC9f148Z9pCH+p 8akzUv4FHV0sTL6zFfIYQ1Ge7zHqLETzjD4qMIcui8i7crPS+PB5hiC4OMdfL5F/4XoD1QViboq q2S1NMaDDmRbUPH0/kURRRR1vh00t7iZpl+NK2jHw4Zc7n7fG0cFX92z06dhHFfSXT5/W8AU4Pi vX9A== X-Received: by 2002:a17:90b:5784:b0:38d:ecfe:41aa with SMTP id 98e67ed59e1d1-38e4b5a08aemr2660937a91.43.1784293463088; Fri, 17 Jul 2026 06:04:23 -0700 (PDT) Received: from localhost ([144.24.58.22]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-31429ff04absm6782635eec.8.2026.07.17.06.04.20 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 17 Jul 2026 06:04:21 -0700 (PDT) From: Jinchao Wang To: Andrew Morton , Peter Zijlstra , Thomas Gleixner , Steven Rostedt , Masami Hiramatsu Cc: Ingo Molnar , Borislav Petkov , Dave Hansen , "H . Peter Anvin" , x86@kernel.org, Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Mathieu Desnoyers , David Hildenbrand , Jonathan Corbet , Matthew Wilcox , Alan Stern , Randy Dunlap , Alexander Potapenko , Marco Elver , Mike Rapoport , linux-kernel@vger.kernel.org, linux-mm@kvack.org, linux-trace-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org, linux-doc@vger.kernel.org, Jinchao Wang Subject: [RFC PATCH v2 05/13] mm/kwatch: add watch expression parser and dereference engine Date: Fri, 17 Jul 2026 09:04:05 -0400 Message-ID: <20260717130405.1902323-1-wangjinchao600@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260717125023.1895892-1-wangjinchao600@gmail.com> References: <20260717125023.1895892-1-wangjinchao600@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" KWatch watches a memory address that is only known once the target function runs, e.g. "argument 1, plus 8, dereferenced once". Add the two halves of that mechanism: - kwatch_deref_parse() turns a textual watch expression {base}[+-off][->[+-]off]... into a kwatch_config: a base anchor (arg1..arg6, stack, an absolute address or - for built-in KWatch - a symbol name) plus a static offset chain. - kwatch_deref_resolve() replays the chain at probe time against pt_regs. Every pointer load goes through get_kernel_nofault() and the final address must be a kernel address. Also add the internal kwatch.h header shared by the rest of the series. Nothing is built yet; the Kconfig entry comes with the control plane. Signed-off-by: Jinchao Wang --- mm/kwatch/Makefile | 3 + mm/kwatch/deref.c | 174 +++++++++++++++++++++++++++++++++++++++++++++ mm/kwatch/kwatch.h | 101 ++++++++++++++++++++++++++ 3 files changed, 278 insertions(+) create mode 100644 mm/kwatch/Makefile create mode 100644 mm/kwatch/deref.c create mode 100644 mm/kwatch/kwatch.h diff --git a/mm/kwatch/Makefile b/mm/kwatch/Makefile new file mode 100644 index 000000000000..69c21ae62123 --- /dev/null +++ b/mm/kwatch/Makefile @@ -0,0 +1,3 @@ +obj-$(CONFIG_KWATCH) +=3D kwatch.o + +kwatch-y :=3D deref.o diff --git a/mm/kwatch/deref.c b/mm/kwatch/deref.c new file mode 100644 index 000000000000..a93c76139e7c --- /dev/null +++ b/mm/kwatch/deref.c @@ -0,0 +1,174 @@ +// SPDX-License-Identifier: GPL-2.0 +#define pr_fmt(fmt) KBUILD_MODNAME ": " fmt + +#include +#include +#include +#include +#include +#include + +#include "kwatch.h" + +int kwatch_deref_resolve(const struct kwatch_config *cfg, struct pt_regs *= regs, + unsigned long *out_addr, u16 *out_len) +{ + unsigned long addr =3D 0; + int i; + + /* 1. Resolve the Base Anchor */ + if (cfg->base =3D=3D KWATCH_BASE_STACK) { + addr =3D kernel_stack_pointer(regs); + if (unlikely(!addr)) + return -EINVAL; + } else if (cfg->base >=3D KWATCH_BASE_ARG1 && + cfg->base <=3D KWATCH_BASE_ARG6) { + int arg_idx =3D cfg->base - KWATCH_BASE_ARG1; + + addr =3D regs_get_kernel_argument(regs, arg_idx); + } else if (cfg->base =3D=3D KWATCH_BASE_ABS_ADDR || + cfg->base =3D=3D KWATCH_BASE_GLOBAL_SYM) { + /* Zero-latency load of the static symbol location */ + addr =3D cfg->sym_addr; + } else { + return -EINVAL; + } + + /* 2. The Pointer-Chasing FSM */ + for (i =3D 0; i < cfg->offset_count; i++) { + addr +=3D cfg->offsets[i]; + + if (i < cfg->offset_count - 1) { + unsigned long next_addr; + + /* Dynamically read the pointer contents at runtime */ + if (get_kernel_nofault(next_addr, (unsigned long *)addr)) + return -EFAULT; + + addr =3D next_addr; + } + } + + /* Enforce strict Kernel-Space boundary */ + if (unlikely(addr < TASK_SIZE_MAX)) + return -EINVAL; + + *out_addr =3D addr; + *out_len =3D cfg->watch_len; + return 0; +} + +int kwatch_deref_parse(struct kwatch_config *cfg, const char *watch_expr) +{ + char *p, *sep, *dup_expr; + char type =3D '\0'; + bool is_deref =3D false; + int ret =3D 0; + + dup_expr =3D kstrdup(watch_expr, GFP_KERNEL); + if (!dup_expr) + return -ENOMEM; + + cfg->offset_count =3D 1; + cfg->offsets[0] =3D 0; + + /* 1. Isolate and Resolve Base Anchor */ + p =3D dup_expr; + sep =3D NULL; + while (*p) { + if (*p =3D=3D '+') { + sep =3D p; + type =3D '+'; + break; + } + if (*p =3D=3D '-') { + sep =3D p; + type =3D '-'; + if (p[1] =3D=3D '>') + is_deref =3D true; + break; + } + p++; + } + + if (type) + *sep =3D '\0'; + + if (!strcmp(dup_expr, "stack")) { + cfg->base =3D KWATCH_BASE_STACK; + } else if (!strncmp(dup_expr, "arg", 3) && strlen(dup_expr) =3D=3D 4) { + int arg_num; + + if (kstrtoint(dup_expr + 3, 10, &arg_num) || arg_num < 1 || + arg_num > 6) { + ret =3D -EINVAL; + goto out; + } + cfg->base =3D KWATCH_BASE_ARG1 + (arg_num - 1); + } else if (kstrtoul(dup_expr, 0, &cfg->sym_addr) =3D=3D 0) { + cfg->base =3D KWATCH_BASE_ABS_ADDR; + } else { +#if IS_BUILTIN(CONFIG_KWATCH) + cfg->sym_addr =3D kallsyms_lookup_name(dup_expr); + if (!cfg->sym_addr) { + pr_err("Failed to resolve symbol name: %s\n", dup_expr); + ret =3D -EINVAL; + goto out; + } + cfg->base =3D KWATCH_BASE_GLOBAL_SYM; +#else + pr_err("cannot resolve symbol %s when built as a module, use a hex addre= ss\n", + dup_expr); + ret =3D -EINVAL; + goto out; +#endif + } + + if (!type) + goto out; + + /* 2. Resolve Base Offset (if + or - exists) */ + if (!is_deref) { + char *next; + + *sep =3D type; /* Restore the '+' or '-' for kstrtol */ + next =3D strstr(sep, "->"); + if (next) + *next =3D '\0'; + + if (kstrtol(sep, 0, &cfg->offsets[0])) { + ret =3D -EINVAL; + goto out; + } + + p =3D next ? next + 2 : NULL; + } else { + /* Jump directly to the first dereference after '->' */ + p =3D sep + 2; + } + + /* 3. Resolve Dereference Chain */ + while (p) { + char *next; + + if (cfg->offset_count >=3D MAX_DEREF_CHAIN) { + ret =3D -E2BIG; + goto out; + } + + next =3D strstr(p, "->"); + if (next) + *next =3D '\0'; + + if (kstrtol(p, 0, &cfg->offsets[cfg->offset_count++])) { + ret =3D -EINVAL; + goto out; + } + + p =3D next ? next + 2 : NULL; + } + +out: + kfree(dup_expr); + return ret; +} diff --git a/mm/kwatch/kwatch.h b/mm/kwatch/kwatch.h new file mode 100644 index 000000000000..dbe0fd0e6a0d --- /dev/null +++ b/mm/kwatch/kwatch.h @@ -0,0 +1,101 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +#ifndef _MM_KWATCH_H +#define _MM_KWATCH_H + +#include +#include +#include +#include +#include +#include +#include + +#define MAX_CONFIG_STR_LEN 512 +#define MAX_DEREF_CHAIN 4 + +struct kwatch_watchpoint; + +struct kwatch_tsk_ctx { + struct task_struct *task; + struct kwatch_watchpoint *wp; + u16 depth; + u32 epoch; +}; + +struct kwatch_watchpoint { + struct perf_event *__percpu *event; + call_single_data_t __percpu *csd_arm; + call_single_data_t __percpu *csd_disarm; + struct perf_event_attr attr; + atomic_t in_use; // multi-consumer safe get/put + struct list_head list; // for cpu online and offline + + struct task_struct *arm_tsk; + atomic_t pending_ipis; + atomic_t refcount; + bool teardown; +}; + +enum kwatch_base_type { + KWATCH_BASE_STACK, + KWATCH_BASE_ABS_ADDR, + KWATCH_BASE_GLOBAL_SYM, + KWATCH_BASE_ARG1, + KWATCH_BASE_ARG2, + KWATCH_BASE_ARG3, + KWATCH_BASE_ARG4, + KWATCH_BASE_ARG5, + KWATCH_BASE_ARG6, +}; + +struct kwatch_config { + u16 max_watch; + char func_name[KSYM_NAME_LEN]; + u16 func_offset; + u16 depth; + u16 duration; + u16 watch_len; + + /* Unified Deref Engine State */ + enum kwatch_base_type base; + char watch_expr[MAX_CONFIG_STR_LEN]; + unsigned long sym_addr; + long offsets[MAX_DEREF_CHAIN]; + u8 offset_count; + u16 max_concurrency; +}; + +int kwatch_hwbp_prealloc(u16 max_watch); +void kwatch_hwbp_free(void); +int kwatch_hwbp_get(struct kwatch_watchpoint **out_wp); +void kwatch_hwbp_arm(struct kwatch_watchpoint *wp, unsigned long addr, u16= len); +int kwatch_hwbp_put(struct kwatch_watchpoint *wp); + +int kwatch_probe_start(struct kwatch_config *cfg); +void kwatch_probe_stop(void); +void kwatch_probe_mute(bool mute); +bool kwatch_probe_validate_hit(struct pt_regs *regs, struct task_struct *a= rm_tsk); +unsigned long kwatch_probe_nmi_rejected(void); +unsigned long kwatch_hwbp_arm_ipi_suppressed(void); + +int kwatch_tsk_ctx_prealloc(u16 max_concurrency); +struct kwatch_tsk_ctx *kwatch_tsk_ctx_get(bool can_alloc); +void kwatch_tsk_ctx_put(void); +void kwatch_tsk_ctx_release(struct kwatch_tsk_ctx *ctx); +void kwatch_tsk_ctx_reset(struct kwatch_tsk_ctx *ctx, u32 new_epoch); +void kwatch_tsk_ctx_release_wps(void); +void kwatch_tsk_ctx_free(void); + +void kwatch_global_anchor(unsigned long duration_sec); +int kwatch_anchor_start(u16 duration); +void kwatch_anchor_stop(void); +void kwatch_anchor_cancel_work(void); +bool kwatch_anchor_has_expired(void); +void kwatch_anchor_clear_expired(void); +void kwatch_auto_stop(void); + +int kwatch_deref_resolve(const struct kwatch_config *cfg, struct pt_regs *= regs, + unsigned long *out_addr, u16 *out_len); +int kwatch_deref_parse(struct kwatch_config *cfg, const char *watch_expr); + +#endif /* _MM_KWATCH_H */ --=20 2.53.0 From nobody Sat Jul 25 05:23:13 2026 Received: from mail-pl1-f172.google.com (mail-pl1-f172.google.com [209.85.214.172]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D116F3FBEDA for ; Fri, 17 Jul 2026 13:04:47 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.172 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293489; cv=none; b=p1UKyFDCfUfwbnBTbHzE5gVZUqHkvRuqDcPnWpusUyHSoRSKREDhXBsOyGcGeIkdM8LL/TjjpkcrwCdzSPfLMQrFCH8S6A+Zo35yYRXrpgfJgMcAi7GHM1fdZZequy8eiQZOH0aJR1as85f0e4ZqX+GqNglXNsTzlxJZ0YVtva0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293489; c=relaxed/simple; bh=EtJoFWpzZWBdawFoDueuFU6iEwMn2zfmcjtMLM64xIU=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=ojvFemzP/ytp3nZ7vtEsz3VN6wzWRkJ7GyOoo7GJqtYPjFeLF1ELZtcbN4c9QzRURyHD8L4KzqbFRK3gicpu/ExlfpAdBxNWRyifPPugFLZPIlJEty8uACbcDPoUhJaBLRoFJ7oAinQaO7xiByo/iXlYcuP6xgw8ULa78NYLPso= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=HQBPPhBU; arc=none smtp.client-ip=209.85.214.172 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="HQBPPhBU" Received: by mail-pl1-f172.google.com with SMTP id d9443c01a7336-2cc891373e0so93701985ad.2 for ; Fri, 17 Jul 2026 06:04:47 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784293487; x=1784898287; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=4DcODrlK3jRROVEKQp+XA+e79BwmfTHaetARzpTbTQI=; b=HQBPPhBUqwGuBpyqLiWxUXiGJvihqyWKxg6MunhyX4Bg96StAMvwhCahapEjSE/DKO zaN6/yuYruwN23I49T5/cyA89FHiba/mfAJ899negpZZZzwmg8ZR+5ldWa3jH1G23Cvh 7N5ayIxfHwlZmcq3YRqh2Wi/awuqXZMUPXa+y0+5/YzqGxe/LprpxCM5MuVFazlMBFPm 5Hz8PWKa1NkaXwsEA7AW/M+g3Jin2+gPx+Om7vswC4kKGOoq1zjqzKQEVtVn5QrOmwXx aVTqXi8GuT20G46uYtrqgV9owAENAnDWL3sEU6VBBq5BAhRxMz2Wg+XLx1STNAHUjOVi 8AIw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784293487; x=1784898287; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=4DcODrlK3jRROVEKQp+XA+e79BwmfTHaetARzpTbTQI=; b=RyBhu7BbRbeMyHzSqsTRE1k2rLcvfChWGKFFPOFJavOa42A82ImQFEPd+CErGM/8kW 1T01A9UHRHTqMYGU1DlsmG265X9lhx3KAHZNwsg928hNmF1sS4xHFuGBQ9h2tA7lzWaD 1VFWNme1QVoAdMdXEE4m2FkBs8BWNogBuKrMJ6i1S5WCXyuRje0nmzJeDYGVgQDzuxYs 2sEA0I2Ognxt7zKPis73/IkyBjyxk6ZYlmg1u+bxvq78jNob6bYAY+HQbr/2OyDlodkm zYZFQAZgQWW9kVT/0CiFWB6M/XR5lDRuoU8fypiBqsf05aWEIwtUhfA2tWozs0rNxPcJ BFew== X-Forwarded-Encrypted: i=1; AHgh+RojfxEXK/ygDwl1oJjNm4ugI+38G1y166dYxgInzJ3f8LOP2Dk3Yl47Jpln3ytXInDToJB8eOOfUwht4Eg=@vger.kernel.org X-Gm-Message-State: AOJu0Yz86beLl4Itq06PQHg19O21SNFJmkdQO6mxixOT+iHMBbU8qNbX NMkZBbpaWqtEcN/F2o7dyC970D9PzbAGqdffB2Rg+YLqBUWBSikL7fwj X-Gm-Gg: AfdE7ckDpS3UcU8NkftP6pRIasRXbQCStG+1U4Qbj+lPUMek2xJRoD8juJv23CNnCOw Dksmmo88E7gk/ENmZzK25mdmMnWbpyfzKJJ6WVUic1ltXL6w7djAJFDRHLSkZNZDsn7eKeAEXu0 TyYQaa4hK3TClnU3zDadh8EBxXTRU2OnbTmMe5iVK+sgIfbmdhzCOq/h3+oXSgWnr/8Q3K0Z0X/ OtRhKaDsER5nULUqZ7FiwrOk5hZsqDDByf7r9VuhqeP5ejnWihrDBOhItuGmCGNztguvINYhbCU VoW+jUbjVh7rCa2OHQ4r7wWRjofAQkWmOqmfzKbacsycQq6hWUHwTPc0rljDzuFySPcvXxDNdTp 8Zyq3OfcFQ7xpke8Av9AllgLlKJOR+TmLiKoDqlpTAN27dd7t0JEXZfYkPFWaUUJqwA0gicYo7h /K0g== X-Received: by 2002:a17:903:32c3:b0:2cc:670d:9b2f with SMTP id d9443c01a7336-2cf349f4041mr28069715ad.33.1784293487195; Fri, 17 Jul 2026 06:04:47 -0700 (PDT) Received: from localhost ([144.24.58.22]) by smtp.gmail.com with ESMTPSA id a92af1059eb24-13ce2cbb7c5sm5234901c88.10.2026.07.17.06.04.45 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 17 Jul 2026 06:04:45 -0700 (PDT) From: Jinchao Wang To: Andrew Morton , Peter Zijlstra , Thomas Gleixner , Steven Rostedt , Masami Hiramatsu Cc: Ingo Molnar , Borislav Petkov , Dave Hansen , "H . Peter Anvin" , x86@kernel.org, Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Mathieu Desnoyers , David Hildenbrand , Jonathan Corbet , Matthew Wilcox , Alan Stern , Randy Dunlap , Alexander Potapenko , Marco Elver , Mike Rapoport , linux-kernel@vger.kernel.org, linux-mm@kvack.org, linux-trace-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org, linux-doc@vger.kernel.org, Jinchao Wang Subject: [RFC PATCH v2 06/13] mm/kwatch: add lockless per-task context pool Date: Fri, 17 Jul 2026 09:04:30 -0400 Message-ID: <20260717130430.1902527-1-wangjinchao600@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260717125023.1895892-1-wangjinchao600@gmail.com> References: <20260717125023.1895892-1-wangjinchao600@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" A task that enters the watched function needs somewhere to keep its window state (nesting depth, owned watchpoint, config epoch). The lookup runs in kprobe and NMI-like contexts, so it must not allocate or take locks. Use a preallocated open-addressing array hashed by task_struct pointer. Slots are claimed with cmpxchg() and released with smp_store_release(); lookup is a read-only probe sequence. The pool size (max_concurrency) bounds how many tasks can be inside watch windows concurrently; excess tasks are simply not tracked. Signed-off-by: Jinchao Wang --- mm/kwatch/Makefile | 2 +- mm/kwatch/task_ctx.c | 125 +++++++++++++++++++++++++++++++++++++++++++ 2 files changed, 126 insertions(+), 1 deletion(-) create mode 100644 mm/kwatch/task_ctx.c diff --git a/mm/kwatch/Makefile b/mm/kwatch/Makefile index 69c21ae62123..cc6574df0d68 100644 --- a/mm/kwatch/Makefile +++ b/mm/kwatch/Makefile @@ -1,3 +1,3 @@ obj-$(CONFIG_KWATCH) +=3D kwatch.o =20 -kwatch-y :=3D deref.o +kwatch-y :=3D deref.o task_ctx.o diff --git a/mm/kwatch/task_ctx.c b/mm/kwatch/task_ctx.c new file mode 100644 index 000000000000..64383a4429e7 --- /dev/null +++ b/mm/kwatch/task_ctx.c @@ -0,0 +1,125 @@ +// SPDX-License-Identifier: GPL-2.0 +#include +#include +#include +#include +#include "kwatch.h" + +static u16 kwatch_ctx_pool_size; +static u16 kwatch_ctx_pool_mask; + +static struct kwatch_tsk_ctx *kwatch_ctx_pool; + +/* Pool size is a u16 and indexes a power-of-two hash table, so bound the + * request away from both the roundup_pow_of_two() u16 overflow (>32768 wr= aps + * to 0) and the degenerate size-1 case (ilog2(1) =3D=3D 0 breaks hash_ptr= ()). + */ +#define KWATCH_CTX_POOL_MIN 256 +#define KWATCH_CTX_POOL_MAX 32768 + +int kwatch_tsk_ctx_prealloc(u16 max_concurrency) +{ + if (max_concurrency < KWATCH_CTX_POOL_MIN) + max_concurrency =3D KWATCH_CTX_POOL_MIN; + else if (max_concurrency > KWATCH_CTX_POOL_MAX) + max_concurrency =3D KWATCH_CTX_POOL_MAX; + + /* + * Set the size/mask only when actually allocating, so they can never + * drift out of sync with the live pool if prealloc is ever called + * again without a matching free. + */ + if (unlikely(!kwatch_ctx_pool)) { + kwatch_ctx_pool_size =3D roundup_pow_of_two(max_concurrency); + kwatch_ctx_pool_mask =3D kwatch_ctx_pool_size - 1; + + kwatch_ctx_pool =3D kcalloc(kwatch_ctx_pool_size, + sizeof(struct kwatch_tsk_ctx), + GFP_KERNEL); + if (!kwatch_ctx_pool) + return -ENOMEM; + } + return 0; +} + +struct kwatch_tsk_ctx *kwatch_tsk_ctx_get(bool can_alloc) +{ + int start_idx, i, idx; + struct task_struct *t; + + if (unlikely(!kwatch_ctx_pool)) + return NULL; + + start_idx =3D hash_ptr(current, ilog2(kwatch_ctx_pool_size)); + + for (i =3D 0; i < kwatch_ctx_pool_size; i++) { + idx =3D (start_idx + i) & kwatch_ctx_pool_mask; + t =3D READ_ONCE(kwatch_ctx_pool[idx].task); + if (t =3D=3D current) + return &kwatch_ctx_pool[idx]; + } + + if (!can_alloc) + return NULL; + + for (i =3D 0; i < kwatch_ctx_pool_size; i++) { + idx =3D (start_idx + i) & kwatch_ctx_pool_mask; + t =3D READ_ONCE(kwatch_ctx_pool[idx].task); + if (!t) { + if (!cmpxchg(&kwatch_ctx_pool[idx].task, NULL, current)) + return &kwatch_ctx_pool[idx]; + } + } + + return NULL; +} + +void kwatch_tsk_ctx_reset(struct kwatch_tsk_ctx *ctx, u32 new_epoch) +{ + struct kwatch_watchpoint *wp =3D xchg(&ctx->wp, NULL); + + if (wp) + kwatch_hwbp_put(wp); + ctx->depth =3D 0; + ctx->epoch =3D new_epoch; +} + +/* Release a slot we hold a pointer to: disarm its wp and free the slot. */ +void kwatch_tsk_ctx_release(struct kwatch_tsk_ctx *ctx) +{ + kwatch_tsk_ctx_reset(ctx, 0); + + /* Pairs with READ_ONCE() in kwatch_tsk_ctx_get() */ + smp_store_release(&ctx->task, NULL); +} + +void kwatch_tsk_ctx_put(void) +{ + struct kwatch_tsk_ctx *ctx =3D kwatch_tsk_ctx_get(false); + + if (unlikely(!ctx)) + return; + + kwatch_tsk_ctx_release(ctx); +} + +void kwatch_tsk_ctx_release_wps(void) +{ + int i; + + if (!kwatch_ctx_pool) + return; + + for (i =3D 0; i < kwatch_ctx_pool_size; i++) { + struct kwatch_watchpoint *wp =3D xchg(&kwatch_ctx_pool[i].wp, + NULL); + if (wp) + kwatch_hwbp_put(wp); + } +} + +void kwatch_tsk_ctx_free(void) +{ + kfree(kwatch_ctx_pool); + kwatch_ctx_pool =3D NULL; +} --=20 2.53.0 From nobody Sat Jul 25 05:23:13 2026 Received: from mail-pl1-f179.google.com (mail-pl1-f179.google.com [209.85.214.179]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 94A653FC5C1 for ; Fri, 17 Jul 2026 13:05:13 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.179 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293515; cv=none; b=QFUkA8GIaIaiF3nnA4Esc/zCNSWxcNI/2C9EDMmZiSw2KNc5dBVTre4zs+jsHqwLqCmr6g1UGCalo+7wizGgH7hCRSlxaGPiVsFMat4YiX2x0+TdHBSz/kkItBqaKfVOxILBXp1Vn76dfr8+Fc7PpP0NIeZHTs+bLJzV5cBnnco= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293515; c=relaxed/simple; bh=VjO/40vwtlEqoqa0MWjtEwISafs1DuSDXBphWFAVeM8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=OFiMv1LE8xGHnO9/jFvvPNdgQkxfA5BZIQchnzZpJfyJrTClVeB7MUX6ujpsB79mBfCpPVkFl5Yu6rDyxfXguYaHxWiaAOdl3LoA3X1okFhD9DMIB2Vfprs+gkDKFS1T7C4y0NALFitIv+QZ3YywRFZw13XgvIX4Pap68PQIzz0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=oWATUe54; arc=none smtp.client-ip=209.85.214.179 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="oWATUe54" Received: by mail-pl1-f179.google.com with SMTP id d9443c01a7336-2cc73e322dbso94624595ad.1 for ; Fri, 17 Jul 2026 06:05:13 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784293513; x=1784898313; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=KwQ36ycy2027O2OnXTT+jaKQ0lX0IY/IjFba+f+HbbQ=; b=oWATUe54S8qnYmiuOKIynXwkaD4LAHPfEbVj3jj/WIW82VnTmhBeB+KgqRpCGeqlxF zPyByHOu/OutpIdkdjnxxaG1iqlKEANYR2gwViXsr0XYfKYU9i8hJnNhkMKCFkob4Yon Ru5eMt+zMlp58uIQmYtt5RHC9kym1R59wQ/1knr9c45Gd1zoZVI5LUY9G6TstxWzqNzb R7NksC181gf20fARAEmBI2LSuMCDN++WrMjhEYy7IR6RyCgT8HWfLZ2qVDarBXDeF/eF 0K6romtYiqad2VVJSoFQmCG6nA6w/8ElI7BKn9rKCHMhjYbHkO895zFboKsjin9YkDgc uvXg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784293513; x=1784898313; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=KwQ36ycy2027O2OnXTT+jaKQ0lX0IY/IjFba+f+HbbQ=; b=JOzybo3iAKODBY5wMGj0nbHejHq9FWNFUWI/UH+m1rt0jdpfxptrB2zQSE/5JUzlSE en8A0e47FcK1BsmMi/9DRDneqNZ9Ch1Qm9RKaseHbzIQxpAchOytu1wBa53tXl0oH9vs e+ZmqvQp8TJz7HEP/QrGpIfj4cHIj4S4zi5evbxOUL0xJXCFy8V0HdYIWTMFuGLTljMY soxPfFp9IQpAuAf2x/+3tC9654+1buwu6fDHXlAll6r9Gm/sCu/I3cIXC+FMTHZTCoG0 05U2m17EtoI6PKPCANnx/VFJRLVTel4SXlE/7Dxc+1mAClMPrO1jLd7QLW02A3nuUHW0 IahQ== X-Forwarded-Encrypted: i=1; AHgh+RrpmfZ2uhXQCQRq14bi8kBQdqR7U4bOBCov+ifHMp3g40TygVq3oJP9sfbiPn1LF9jwYSxRBiIlAFASwlo=@vger.kernel.org X-Gm-Message-State: AOJu0YwUunlZrPV5U41K97qzjKtNYZlnqm7Y5Nw9ga4h++faOa+GK1FB v5bCbm2DEYtkZoxFw+mglBfKAm2f3yBun1yeWQNIX/z8fcABSN+Gd7VA X-Gm-Gg: AfdE7cmdt8OeA8wbn6KXf28wDknBnq4hqgRaJod7kMtB5BtJmfZ25o0udbnr8VHaGI+ uwsEceZCxw7+iaVlv8zs5JSd4Ut/wSCPPOJn/y4JRgrajzVEvjpZ5EP1MWMiwwaEEiOlNtQaxlT TQvDxil6VXdaXpxMB6BD9gz3/DgKWXInNr+s0XSz+RFkRW9yl2dbCDLZDksTfLtl2mvoA0GdlpS NylcjASKm7Ux+0gjwyTh9wPF9rlWNP+MEk56r40m85G0yYcfsQvM5Vw5NP/dh1uL+rGqtggTlaZ nvlT3f7gJtr1qsdTknKYV6D+gtr8mp2bl/rPnqjLdFhZV7RejIvgsLtP8Luv+Ovfz4g8xH6XSLX sBb0/PUaFuQE2kkz6Mz3x8Mx5jkoEBWm5UgcyZx9fbr9UJjOjA+cq0iKGNQWG56L43HKK6eYVYD r8Zw== X-Received: by 2002:a17:902:d509:b0:2c9:cf62:6f61 with SMTP id d9443c01a7336-2cf348b264cmr25931005ad.17.1784293512798; Fri, 17 Jul 2026 06:05:12 -0700 (PDT) Received: from localhost ([144.24.58.22]) by smtp.gmail.com with ESMTPSA id a92af1059eb24-13ce2cbb6e8sm5156012c88.9.2026.07.17.06.05.11 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 17 Jul 2026 06:05:11 -0700 (PDT) From: Jinchao Wang To: Andrew Morton , Peter Zijlstra , Thomas Gleixner , Steven Rostedt , Masami Hiramatsu Cc: Ingo Molnar , Borislav Petkov , Dave Hansen , "H . Peter Anvin" , x86@kernel.org, Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Mathieu Desnoyers , David Hildenbrand , Jonathan Corbet , Matthew Wilcox , Alan Stern , Randy Dunlap , Alexander Potapenko , Marco Elver , Mike Rapoport , linux-kernel@vger.kernel.org, linux-mm@kvack.org, linux-trace-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org, linux-doc@vger.kernel.org, Jinchao Wang Subject: [RFC PATCH v2 07/13] stacktrace: export stack_trace_save_regs() Date: Fri, 17 Jul 2026 09:04:54 -0400 Message-ID: <20260717130454.1902704-1-wangjinchao600@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260717125023.1895892-1-wangjinchao600@gmail.com> References: <20260717125023.1895892-1-wangjinchao600@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" The other stack_trace_save_*() flavours are exported, but the regs variant is not, so no module can capture a stack trace for a given pt_regs. KWatch, which may be built as a module, uses it to record who wrote to a watched address from the hardware breakpoint handler. Export it like its siblings. Signed-off-by: Jinchao Wang --- kernel/stacktrace.c | 2 ++ 1 file changed, 2 insertions(+) diff --git a/kernel/stacktrace.c b/kernel/stacktrace.c index afb3c116da91..d853c40f916b 100644 --- a/kernel/stacktrace.c +++ b/kernel/stacktrace.c @@ -175,6 +175,7 @@ unsigned int stack_trace_save_regs(struct pt_regs *regs= , unsigned long *store, arch_stack_walk(consume_entry, &c, current, regs); return c.len; } +EXPORT_SYMBOL_GPL(stack_trace_save_regs); =20 #ifdef CONFIG_HAVE_RELIABLE_STACKTRACE /** @@ -325,6 +326,7 @@ unsigned int stack_trace_save_regs(struct pt_regs *regs= , unsigned long *store, save_stack_trace_regs(regs, &trace); return trace.nr_entries; } +EXPORT_SYMBOL_GPL(stack_trace_save_regs); =20 #ifdef CONFIG_HAVE_RELIABLE_STACKTRACE /** --=20 2.53.0 From nobody Sat Jul 25 05:23:13 2026 Received: from mail-pj1-f51.google.com (mail-pj1-f51.google.com [209.85.216.51]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2A08C3E7BCE for ; Fri, 17 Jul 2026 13:05:37 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.51 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293539; cv=none; b=h8e9JMQCp2z1Nj1iSKBvVwAcrRU36XXVOEmPOhQD7uEJXXEnB68FfhDPfwJ8UYrbU9xt0OcVmEYuV3oX/UAwblJ+nFkBPKBvV2ZFqOyWihXW/9D54VCuW4YNpizd3DWeanCapwoA2TPDTxYHTCZVDqwrSd4G7CFpJcPptfjEsSI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293539; c=relaxed/simple; bh=wH8kkta8OF3VXBXdp1N+NHYfBBuMcPNLtXDApPbZ660=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Xl7TuHmM14GeaWEOn/HLJaQSkKL28Ykk5e8R5wqocc3NarmO/f6KI5jTjRs4X5gnOefGxAlD4jp3wdfe5x8S2eXF62VRFTRrSj4K4nenZa3Z8T0y37rQu9JBdL0snjkrAia4PnM+q1eAXYEVypV06BXQofSzLTZIqhwWIsKHbQc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=kwxzabev; arc=none smtp.client-ip=209.85.216.51 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="kwxzabev" Received: by mail-pj1-f51.google.com with SMTP id 98e67ed59e1d1-38125cebfdaso4787163a91.1 for ; Fri, 17 Jul 2026 06:05:37 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784293537; x=1784898337; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=/aEQ2n6TvdCBbgIi3sEsUgz+yDJwxM1XgcBftQT7Uys=; b=kwxzabevZWoJ0jhsYcePYnV6+2mKaT2a7tIUl/hLsHb4KQ6bwMlyHxFvL/HPbW8Jdf odJ+3uM1MSWj8IkgqeEL4YUUkSZXqD9B1BegighR0j0iPP7PVxZj2gE3OjYfMPgWvLQ5 el3Wunz7i6ElxLi9sQApwORmIrxZiOXg8HnUXiZw693BqHx7X7WjT2x8mi0WTkFW0g68 yQNV4S5tO/kmbOz/jQ/q6kSaf5w6vunNKmyfnHUyeO+XJTo6Gbd2e7Q0ClFkUjGO0Uld Os+4SKuplwMIqe/8yo3oa2dix3ztFGkHaa9j4ue3Q+Yg0zoJfqALpKu1ZbZ0NbW0VHnX +LFw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784293537; x=1784898337; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=/aEQ2n6TvdCBbgIi3sEsUgz+yDJwxM1XgcBftQT7Uys=; b=SRayqrUeneJOEYP855i872k+2Y232SHNAT46Xunahq12AB4dgj6Zpk0GcQToRq3EUR 9vfotDNeOtszp8qXcEVxzH6f6cP+uiU7r0onYq14PLV86zYomCM9Y1BS9Y150Mbi1KfB n0pIubMcOysvDhIMhstqSufHK4/fLcNdhklCppm5IVS9nC6bHLPSZ/m5b69kc3zyTRK9 ZEsLIMheOSpO6V+6nPsGYeKoJ4CVd5va+Td4KMLk2GpnBN/isf3y6BqZ3NmnSQdDInLw Y6z7WBVJl3B/iZ4NI1RslRQlkhTDeRlgUPQw0zcOol2reE5cMuZvJt/dlqgSmjLkbN5X iPIQ== X-Forwarded-Encrypted: i=1; AHgh+Rp148p1rUJKD3FcQqha080v+KGSBXixBi6IlxB8ZL5n9+0xtR2MTCHMEt4iT1LzcN/n9fylRY/E/hAftpA=@vger.kernel.org X-Gm-Message-State: AOJu0YxxoSBj65isbjifa+whybg6UrO51UlKgq28SvPdHaKBCD96DcGQ SdQ3xIEMlCqqtyoVmr/NmWs9o+4C2YSEXlxwHsTrNCshn0tpyJxUv7gw X-Gm-Gg: AfdE7clqY8TBut2hCyGSKib6sKaK+kRyvB0D0P221kkOqJC2OVqhg0uO3MbWQpFf0F9 oMhc2N9krAmLq7J+y3pUlQ11ybRE5fTokX/t86HWct5WTRuIobPPaIG2EJkm/G4yLwjssqe60fT nr8N5nkz+Jkh5Pwq8nI+M4kijEipb1qcFO4fB02956prGmrvx2Lan4Kfiy/vpUx/C1WE+b0Q0B9 DovzNrDkyd7fbyuILbyCeHsdKpPrBJJNIdFOPuWOylkyW9zlz3X61fHYZE7Dvl41C197xSJEDY0 GNSZm2JF0B54PBWvUvyCx+f+ZHL7xs+HiMOEwlF2wt1iRnl2CIMIcK8HmPlWMPwze0HRrFIURv3 1MxfaMMDBSvYftAh4E43kPRnNylLewRZZ1blmz4ckPQd2s/mxJ/s9HT4Csy6HUL49DvF3SwyzhL Gx9g== X-Received: by 2002:a17:90b:278e:b0:38a:c3f:3b87 with SMTP id 98e67ed59e1d1-38e4b454f0bmr2580230a91.12.1784293537053; Fri, 17 Jul 2026 06:05:37 -0700 (PDT) Received: from localhost ([144.24.58.22]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-3142a1bb77csm7739750eec.18.2026.07.17.06.05.35 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 17 Jul 2026 06:05:35 -0700 (PDT) From: Jinchao Wang To: Andrew Morton , Peter Zijlstra , Thomas Gleixner , Steven Rostedt , Masami Hiramatsu Cc: Ingo Molnar , Borislav Petkov , Dave Hansen , "H . Peter Anvin" , x86@kernel.org, Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Mathieu Desnoyers , David Hildenbrand , Jonathan Corbet , Matthew Wilcox , Alan Stern , Randy Dunlap , Alexander Potapenko , Marco Elver , Mike Rapoport , linux-kernel@vger.kernel.org, linux-mm@kvack.org, linux-trace-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org, linux-doc@vger.kernel.org, Jinchao Wang Subject: [RFC PATCH v2 08/13] mm/kwatch: add hardware breakpoint backend Date: Fri, 17 Jul 2026 09:05:20 -0400 Message-ID: <20260717130520.1902926-1-wangjinchao600@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260717125023.1895892-1-wangjinchao600@gmail.com> References: <20260717125023.1895892-1-wangjinchao600@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Manage a preallocated pool of wide (per-CPU) perf hardware breakpoints. All breakpoints are registered up front against a dummy address; arming a watchpoint only re-points an already-registered event, so the arm path can run from a kprobe handler. - kwatch_hwbp_get()/put() claim and release pool entries with per-slot cmpxchg, safe for concurrent consumers on any CPU. - kwatch_hwbp_arm() updates the local CPU synchronously via modify_wide_hw_breakpoint_local() and broadcasts asynchronous IPIs to the other CPUs. Arm-side IPIs are rate-limited per CPU; disarm IPIs are refcounted so an entry is only recycled once every CPU has dropped it. - Hits are reported through the kwatch:kwatch_hit tracepoint with a short stack trace: the ftrace ring buffer is usable from NMI-like context and survives a subsequent crash, unlike printk. - A CPU hotplug callback creates/destroys the per-CPU events as CPUs come and go. Signed-off-by: Jinchao Wang --- include/trace/events/kwatch.h | 68 ++++++ mm/kwatch/Makefile | 2 +- mm/kwatch/hwbp.c | 388 ++++++++++++++++++++++++++++++++++ 3 files changed, 457 insertions(+), 1 deletion(-) create mode 100644 include/trace/events/kwatch.h create mode 100644 mm/kwatch/hwbp.c diff --git a/include/trace/events/kwatch.h b/include/trace/events/kwatch.h new file mode 100644 index 000000000000..8a2ec6811ad4 --- /dev/null +++ b/include/trace/events/kwatch.h @@ -0,0 +1,68 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +#undef TRACE_SYSTEM +#define TRACE_SYSTEM kwatch + +#if !defined(_TRACE_KWATCH_H) || defined(TRACE_HEADER_MULTI_READ) +#define _TRACE_KWATCH_H + +#include +#include +#include + +#define KWATCH_STACK_DEPTH 8 + +struct trace_seq; +const char *kwatch_trace_print_stack(struct trace_seq *p, + const unsigned long *stack, + unsigned int nr); + +TRACE_EVENT(kwatch_hit, + TP_PROTO(unsigned long ip, unsigned long sp, unsigned long addr, + u64 time_ns, + unsigned long *stack_entries, unsigned int stack_nr), + TP_ARGS(ip, sp, addr, time_ns, stack_entries, stack_nr), + + TP_STRUCT__entry( + /* + * time_ns first: u64 leading the entry avoids a 4-byte hole + * after the unsigned-long fields on 32-bit. stack_nr trails + * the fixed fields for the same reason; the stack is a + * dynamic array sized to what was actually captured, so a + * short trace neither wastes space nor leaks uninitialized + * tail slots. + */ + __field(u64, time_ns) + __field(unsigned long, ip) + __field(unsigned long, sp) + __field(unsigned long, addr) + __dynamic_array(unsigned long, stack, + min_t(unsigned int, stack_nr, KWATCH_STACK_DEPTH)) + __field(unsigned int, stack_nr) + ), + + TP_fast_assign( + unsigned long *stack =3D __get_dynamic_array(stack); + unsigned int i; + + __entry->time_ns =3D time_ns; + __entry->ip =3D ip; + __entry->sp =3D sp; + __entry->addr =3D addr; + __entry->stack_nr =3D min_t(unsigned int, stack_nr, + KWATCH_STACK_DEPTH); + for (i =3D 0; i < __entry->stack_nr; i++) + stack[i] =3D stack_entries[i]; + ), + + TP_printk("KWatch HIT: time=3D%llu.%06u ip=3D%pS addr=3D0x%lx%s", + div_u64(__entry->time_ns, 1000000000ULL), + (unsigned int)(div_u64(__entry->time_ns, 1000ULL) % 1000000ULL), + (void *)__entry->ip, __entry->addr, + kwatch_trace_print_stack(p, __get_dynamic_array(stack), + __entry->stack_nr)) +); + +#endif /* _TRACE_KWATCH_H */ + +/* This part must be outside protection */ +#include diff --git a/mm/kwatch/Makefile b/mm/kwatch/Makefile index cc6574df0d68..b2bc3003c89b 100644 --- a/mm/kwatch/Makefile +++ b/mm/kwatch/Makefile @@ -1,3 +1,3 @@ obj-$(CONFIG_KWATCH) +=3D kwatch.o =20 -kwatch-y :=3D deref.o task_ctx.o +kwatch-y :=3D deref.o task_ctx.o hwbp.o diff --git a/mm/kwatch/hwbp.c b/mm/kwatch/hwbp.c new file mode 100644 index 000000000000..d1e93754cce8 --- /dev/null +++ b/mm/kwatch/hwbp.c @@ -0,0 +1,388 @@ +// SPDX-License-Identifier: GPL-2.0 +#define pr_fmt(fmt) KBUILD_MODNAME ": " fmt + +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "kwatch.h" + +/* Minimum spacing between cross-CPU arm broadcasts, per CPU. */ +#define KWATCH_ARM_IPI_MIN_INTERVAL_NS 1000000ULL + +static LIST_HEAD(kwatch_all_wp_list); +static struct kwatch_watchpoint **kwatch_wp_slots; +static u16 kwatch_wp_nr; +static DEFINE_MUTEX(kwatch_all_wp_mutex); +static unsigned long kwatch_dummy_holder __aligned(8); +static int kwatch_hwbp_cpuhp_state =3D CPUHP_INVALID; +static atomic_long_t kwatch_arm_ipi_suppressed; + +unsigned long kwatch_hwbp_arm_ipi_suppressed(void) +{ + return atomic_long_read(&kwatch_arm_ipi_suppressed); +} + +#define CREATE_TRACE_POINTS +#include + +/* + * Render the saved stack like the ftrace built-in stacktrace / dump_stack= () + * style. Symbol resolution runs at trace read time, not in the hit path. + */ +const char *kwatch_trace_print_stack(struct trace_seq *p, + const unsigned long *stack, + unsigned int nr) +{ + const char *ret =3D trace_seq_buffer_ptr(p); + unsigned int i; + + for (i =3D 0; i < nr; i++) + trace_seq_printf(p, "\n =3D> %pS", (void *)stack[i]); + trace_seq_putc(p, 0); + return ret; +} + +static void kwatch_hwbp_handler(struct perf_event *bp, + struct perf_sample_data *data, + struct pt_regs *regs) +{ + struct kwatch_watchpoint *wp =3D bp->overflow_handler_context; + unsigned long stack_entries[KWATCH_STACK_DEPTH]; + unsigned int stack_nr; + + if (!kwatch_probe_validate_hit(regs, wp->arm_tsk)) + return; + + stack_nr =3D stack_trace_save_regs(regs, stack_entries, KWATCH_STACK_DEPT= H, 2); + trace_kwatch_hit(instruction_pointer(regs), kernel_stack_pointer(regs), + bp->attr.bp_addr, local_clock(), + stack_entries, stack_nr); +} + +static void kwatch_hwbp_arm_local(void *info) +{ + struct kwatch_watchpoint *wp =3D info; + struct perf_event *bp; + unsigned long flags; + int cpu, err; + + local_irq_save(flags); + + cpu =3D smp_processor_id(); + bp =3D per_cpu(*wp->event, cpu); + + if (unlikely(!bp)) + goto out; + + kwatch_probe_mute(true); + barrier(); + + /* + * On success this also updates the per-CPU bp->attr, so the hit + * handler reports what THIS CPU is watching instead of the shared + * wp->attr, which another CPU may be re-pointing. + */ + err =3D modify_wide_hw_breakpoint_local(bp, &wp->attr); + if (unlikely(err)) + WARN_ONCE(1, + "KWatch: HWBP reinstall failed on CPU%d (err=3D%d, addr=3D0x%llx, len= =3D%llu)\n", + cpu, err, wp->attr.bp_addr, wp->attr.bp_len); + + barrier(); + kwatch_probe_mute(false); + +out: + local_irq_restore(flags); +} + +static inline void kwatch_hwbp_try_recycle(struct kwatch_watchpoint *wp) +{ + if (atomic_dec_and_test(&wp->pending_ipis)) { + if (!READ_ONCE(wp->teardown)) + atomic_set_release(&wp->in_use, 0); + + atomic_dec(&wp->refcount); + } +} + +static void kwatch_hwbp_disarm_local(void *info) +{ + struct kwatch_watchpoint *wp =3D info; + + kwatch_hwbp_arm_local(info); + kwatch_hwbp_try_recycle(wp); +} + +static int kwatch_hwbp_cpu_online(unsigned int cpu) +{ + struct perf_event_attr attr; + struct kwatch_watchpoint *wp; + struct perf_event *bp; + + mutex_lock(&kwatch_all_wp_mutex); + list_for_each_entry(wp, &kwatch_all_wp_list, list) { + attr =3D wp->attr; + attr.bp_addr =3D (unsigned long)&kwatch_dummy_holder; + bp =3D perf_event_create_kernel_counter(&attr, cpu, NULL, + kwatch_hwbp_handler, wp); + if (IS_ERR(bp)) { + pr_warn("%s failed to create watch on CPU %d: %ld\n", + __func__, cpu, PTR_ERR(bp)); + continue; + } + per_cpu(*wp->event, cpu) =3D bp; + } + mutex_unlock(&kwatch_all_wp_mutex); + return 0; +} + +static int kwatch_hwbp_cpu_offline(unsigned int cpu) +{ + struct kwatch_watchpoint *wp; + struct perf_event *bp; + + mutex_lock(&kwatch_all_wp_mutex); + list_for_each_entry(wp, &kwatch_all_wp_list, list) { + bp =3D per_cpu(*wp->event, cpu); + if (bp) { + unregister_hw_breakpoint(bp); + per_cpu(*wp->event, cpu) =3D NULL; + } + } + mutex_unlock(&kwatch_all_wp_mutex); + return 0; +} + +int kwatch_hwbp_get(struct kwatch_watchpoint **out_wp) +{ + struct kwatch_watchpoint *wp; + int i; + + /* + * Per-slot cmpxchg claim: safe for concurrent consumers on any CPU, + * unlike llist_del_first() which requires a single consumer. + */ + for (i =3D 0; i < kwatch_wp_nr; i++) { + wp =3D kwatch_wp_slots[i]; + if (atomic_read(&wp->in_use)) + continue; + if (atomic_cmpxchg(&wp->in_use, 0, 1) =3D=3D 0) { + atomic_inc(&wp->refcount); + *out_wp =3D wp; + return 0; + } + } + return -EBUSY; +} + +void kwatch_hwbp_arm(struct kwatch_watchpoint *wp, unsigned long addr, u16= len) +{ + static DEFINE_PER_CPU(u64, last_ipi_time); + int cur_cpu; + call_single_data_t *csd; + int cpu; + bool is_disarm =3D (addr =3D=3D (unsigned long)&kwatch_dummy_holder); + bool skip_remote =3D false; + + wp->attr.bp_addr =3D addr; + wp->attr.bp_len =3D len; + + if (!is_disarm) + wp->arm_tsk =3D current; + + /* ensure attr update visible to other cpu before sending IPI */ + smp_wmb(); + + atomic_set(&wp->pending_ipis, 1); + cur_cpu =3D get_cpu(); + + /* + * Rate-limit only the cross-CPU broadcast, never the local re-point. + * Arming the current CPU is free and must always reflect this window; + * only the remote IPI fan-out is throttled to keep a hot function from + * storming every CPU. A suppressed broadcast means remote CPUs keep + * watching the previous address for that window (a missed remote-CPU + * writer is possible) - hence the visible counter, and why kwatch + * targets low-frequency functions. Disarm is never throttled: the + * slot must always be released. + */ + if (!is_disarm) { + u64 now =3D local_clock(); + u64 last =3D this_cpu_read(last_ipi_time); + + if (now - last < KWATCH_ARM_IPI_MIN_INTERVAL_NS) { + atomic_long_inc(&kwatch_arm_ipi_suppressed); + skip_remote =3D true; + } else { + this_cpu_write(last_ipi_time, now); + } + } + + if (!skip_remote) { + for_each_online_cpu(cpu) { + if (cpu =3D=3D cur_cpu) + continue; + + if (is_disarm) + atomic_inc(&wp->pending_ipis); + + csd =3D per_cpu_ptr(is_disarm ? wp->csd_disarm : wp->csd_arm, + cpu); + /* + * The arm path ignores a -EBUSY return: a wp has a single + * owner (claimed via kwatch_hwbp_get(), held until exit) + * and is armed once per window, and the per-CPU csd queue + * is FIFO, so this window's csd_arm cannot still be pending + * from a prior window (its disarm, queued later, gates the + * wp's reuse). Do not "fix" this into a retry. + */ + if (smp_call_function_single_async(cpu, csd) && is_disarm) + kwatch_hwbp_try_recycle(wp); + } + } + put_cpu(); + + if (is_disarm) + kwatch_hwbp_disarm_local(wp); + else + kwatch_hwbp_arm_local(wp); +} + +int kwatch_hwbp_put(struct kwatch_watchpoint *wp) +{ + kwatch_hwbp_arm(wp, (unsigned long)&kwatch_dummy_holder, + sizeof(unsigned long)); + + return 0; +} + +void kwatch_hwbp_free(void) +{ + struct kwatch_watchpoint *wp, *tmp; + + kwatch_wp_nr =3D 0; + kfree(kwatch_wp_slots); + kwatch_wp_slots =3D NULL; + + if (kwatch_hwbp_cpuhp_state !=3D CPUHP_INVALID) { + cpuhp_remove_state_nocalls(kwatch_hwbp_cpuhp_state); + kwatch_hwbp_cpuhp_state =3D CPUHP_INVALID; + } + + mutex_lock(&kwatch_all_wp_mutex); + list_for_each_entry_safe(wp, tmp, &kwatch_all_wp_list, list) { + list_del(&wp->list); + + WRITE_ONCE(wp->teardown, true); + atomic_dec(&wp->refcount); + + /* Wait for all async IPIs to finish */ + while (atomic_read(&wp->refcount) > 0) + cpu_relax(); + + unregister_wide_hw_breakpoint(wp->event); + free_percpu(wp->csd_arm); + free_percpu(wp->csd_disarm); + kfree(wp); + } + mutex_unlock(&kwatch_all_wp_mutex); +} + +int kwatch_hwbp_prealloc(u16 max_watch) +{ + struct kwatch_watchpoint *wp; + int success =3D 0, cpu; + int ret; + + atomic_long_set(&kwatch_arm_ipi_suppressed, 0); + + while (!max_watch || success < max_watch) { + wp =3D kzalloc_obj(*wp); + if (!wp) + break; + + wp->csd_arm =3D alloc_percpu(call_single_data_t); + wp->csd_disarm =3D alloc_percpu(call_single_data_t); + if (!wp->csd_arm || !wp->csd_disarm) { + free_percpu(wp->csd_arm); + free_percpu(wp->csd_disarm); + kfree(wp); + break; + } + + for_each_possible_cpu(cpu) { + INIT_CSD(per_cpu_ptr(wp->csd_arm, cpu), + kwatch_hwbp_arm_local, wp); + INIT_CSD(per_cpu_ptr(wp->csd_disarm, cpu), + kwatch_hwbp_disarm_local, wp); + } + + wp->teardown =3D false; + + hw_breakpoint_init(&wp->attr); + wp->attr.bp_addr =3D (unsigned long)&kwatch_dummy_holder; + wp->attr.bp_len =3D sizeof(unsigned long); + /* kwatch localizes corruption: it always watches for writes. */ + wp->attr.bp_type =3D HW_BREAKPOINT_W; + + wp->event =3D register_wide_hw_breakpoint(&wp->attr, + kwatch_hwbp_handler, + wp); + if (IS_ERR_PCPU(wp->event)) { + free_percpu(wp->csd_arm); + free_percpu(wp->csd_disarm); + kfree(wp); + break; + } + + atomic_set(&wp->refcount, 1); + + mutex_lock(&kwatch_all_wp_mutex); + list_add(&wp->list, &kwatch_all_wp_list); + mutex_unlock(&kwatch_all_wp_mutex); + success++; + } + + if (!success) + return -EBUSY; + + /* + * A fresh prealloc must start from an empty slot array; warn if a + * previous session was not torn down, since refilling without a reset + * would index past the freshly sized array. + */ + WARN_ON_ONCE(kwatch_wp_slots || kwatch_wp_nr); + kwatch_wp_nr =3D 0; + + kwatch_wp_slots =3D kcalloc(success, sizeof(*kwatch_wp_slots), + GFP_KERNEL); + if (!kwatch_wp_slots) { + kwatch_hwbp_free(); + return -ENOMEM; + } + mutex_lock(&kwatch_all_wp_mutex); + list_for_each_entry(wp, &kwatch_all_wp_list, list) + kwatch_wp_slots[kwatch_wp_nr++] =3D wp; + mutex_unlock(&kwatch_all_wp_mutex); + + ret =3D cpuhp_setup_state_nocalls(CPUHP_AP_ONLINE_DYN, "kwatch:online", + kwatch_hwbp_cpu_online, + kwatch_hwbp_cpu_offline); + if (ret < 0) { + kwatch_hwbp_free(); + return ret; + } + + kwatch_hwbp_cpuhp_state =3D ret; + return 0; +} --=20 2.53.0 From nobody Sat Jul 25 05:23:13 2026 Received: from mail-pg1-f175.google.com (mail-pg1-f175.google.com [209.85.215.175]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E18A03FBB7E for ; Fri, 17 Jul 2026 13:06:02 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.175 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293566; cv=none; b=fzdYCePTL7vcGnGNoV0DHvmA1xvqz6T3HQp6DPd849TBh8XosWgXSw6o787v13xtZJ3EOkQipmJgBZuewRfkdH45EIQzvhP1QhCJLsoSyQVLequHpJ4/tWpoMPtHs/VrVz57qc9beY1Hk+h+1hppzVriK8mOGjemJ+/IsThSAMc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293566; c=relaxed/simple; bh=VzFZN7zkXZafCAxmduqjsB434NPBXO3a9Es9wEA/9Hk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=CY6HnKlWzd4iaafAm311jV60GMqR3dwVPL9+A/FxUoKymKjcWpj0EzrTna3rQQLpHbnxKRMYKu9jIyzhODZHO/sXoaD/HtHH4FpYp4TZtOvdRAgoP2oM1VK0IQ8x+geFAiXk/2MFCB7Yvhe039yuuOLCJe7X5AW/IWnTaN4cUTc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=ada7Icpn; arc=none smtp.client-ip=209.85.215.175 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="ada7Icpn" Received: by mail-pg1-f175.google.com with SMTP id 41be03b00d2f7-ca913a601fbso5367526a12.3 for ; Fri, 17 Jul 2026 06:06:02 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784293562; x=1784898362; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=gLUXizPmm8lFk+lVj4kGKhusb/cUxotlJSZc5i0VZQA=; b=ada7IcpnkGVsqhB/TlkBnKLL+5bJV8WIIvV0hO583H5onfkTo01xleJ4Ii/+401AcJ DY9dhhJlHq0u9XqwUIYWl2FcCGjnrPbGTF8P0dUECZcGQStv8CmYj/QSx18G14sgahSg NKeFuke9QB7rJdHztqi+55IRJ5mhHF7ySSXshQ89YH4wP1dtp1gjT8JzC/ACeUzrWA40 XGQVIh+szyQeDAjXsx3miu/KJHXlHpj43v2cfxsB+8A5shWgKvjDNiFugqhTabPtnkmI 5gFcPrX3nKld6X1+rhz3uUhURc+5uSYYpn4ikRtQ2MBE72enr6KCddvFa7AgWpqnFfFr 6S0w== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784293562; x=1784898362; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=gLUXizPmm8lFk+lVj4kGKhusb/cUxotlJSZc5i0VZQA=; b=eStd1wwaKR6gkubg5O8uSN/rrkks8Jup3iJ1O9TSULUTirKqdpp6tsYGZ4ixqBB7aq 2pLm3vbynIOb8ENCOTAr4FLXHIuz1v13cmZO8X7uHpf0JBQCC2mr/92NnBiu/cgTtqYE DgamXk+/6WWullIWyAISpXhn/y72sA4YiKm9/0+oXt5SNZtMVfSSbWvURxc72ySlB+cz 7/I0rWHc0mCLnV4PH3Uh62BvQCo4pLLuC2990ZV6ysu2asWj3dq0mDY3LYxfy5CQ2oae qmLZbI+HqGA2ZZ4K6BVTOQeBEFCyAuR8pm0CjYRFN3i80SDnYZ577nBx3Al5lnW6KfI4 TZag== X-Forwarded-Encrypted: i=1; AHgh+RoPvuA4yzipy0bd+otZi7D45wJi6s+qm9k8I6pAAmF/ylKyWNwepuNty5ISRpsj9cRjbJDztej5FYpoEKY=@vger.kernel.org X-Gm-Message-State: AOJu0YwMY8ZiU8yzdJO93NFkqfSYx+LuXMlEp47t+l+uVzNmMbu2HGEB ofmAPmbMJp/H6G7BCF8sGy7JTu76fz0m6gfCGVdGirxJtJ2glax58cSZ X-Gm-Gg: AfdE7cmvsxP2bOmX/896KQOb348yM6QnVqHlkmi79fp9OcCbJ1qVJ6C4cBxSE4kMEM5 ZkFoFZbWR4PH41uIPEfLfT4JCyyl2nh/pVphgM7+oG4ULzP4+QfRBCnZEtsHJSl90mVKq8NykNa 4RpdxSGGM5bSf+TtXkUJUUV2xP8NSmmUCbOYmKdJ9akE+IwUv3I10dqGlU0CgmHWi6zE4EpT3Fr kAepObpdeCfADsnfsg6MVDSnrBw3iXdOC/cK0BgrqZvUSvXFJ2GUW3xeDWt717XgkAw7FEXiiCd rF9YDQH61SOBISqEKcH9gm9T037bFBMXPdEzdWCGszxjoftRDEI5mFG+B6PAZJVOodytkC5rakg nUWPQwCEEJBqbROpeBvVm18QcDk4CSaGJrg/2RecAemIeVKCbNldevlgXCqil+S18P9Xdazmh4m m7mw== X-Received: by 2002:a17:90b:3bcd:b0:387:e0cb:c8e1 with SMTP id 98e67ed59e1d1-38e4b594dadmr2450658a91.42.1784293561881; Fri, 17 Jul 2026 06:06:01 -0700 (PDT) Received: from localhost ([144.24.58.22]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-3142a1bb8a4sm7964292eec.14.2026.07.17.06.05.59 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 17 Jul 2026 06:06:00 -0700 (PDT) From: Jinchao Wang To: Andrew Morton , Peter Zijlstra , Thomas Gleixner , Steven Rostedt , Masami Hiramatsu Cc: Ingo Molnar , Borislav Petkov , Dave Hansen , "H . Peter Anvin" , x86@kernel.org, Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Mathieu Desnoyers , David Hildenbrand , Jonathan Corbet , Matthew Wilcox , Alan Stern , Randy Dunlap , Alexander Potapenko , Marco Elver , Mike Rapoport , linux-kernel@vger.kernel.org, linux-mm@kvack.org, linux-trace-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org, linux-doc@vger.kernel.org, Jinchao Wang Subject: [RFC PATCH v2 09/13] mm/kwatch: add probe lifecycle runtime Date: Fri, 17 Jul 2026 09:05:44 -0400 Message-ID: <20260717130544.1903146-1-wangjinchao600@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260717125023.1895892-1-wangjinchao600@gmail.com> References: <20260717125023.1895892-1-wangjinchao600@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Open and close the watch window with a kretprobe on the target function: the entry handler tracks per-task nesting depth and, when the configured depth is reached, resolves the watch expression and arms a watchpoint; the exit handler disarms it. An optional kprobe at func_offset arms mid-function instead of at entry. Functions running in a real NMI(-like) context are rejected once, at function entry, by comparing the NMI nesting count against the one NMI-like layer that int3-based kprobe delivery itself adds; a companion kprobe with a post_handler pins the probe point so jump optimization cannot change the delivery mechanism after it is sampled. Rejections are counted and exposed to the control plane. A global epoch versioning scheme invalidates stale per-task state across reconfigurations, and a per-CPU mute flag keeps window management quiet while a CPU rewrites its own debug registers. Signed-off-by: Jinchao Wang --- mm/kwatch/Makefile | 2 +- mm/kwatch/probe.c | 275 +++++++++++++++++++++++++++++++++++++++++++++ 2 files changed, 276 insertions(+), 1 deletion(-) create mode 100644 mm/kwatch/probe.c diff --git a/mm/kwatch/Makefile b/mm/kwatch/Makefile index b2bc3003c89b..f04673cc5b1c 100644 --- a/mm/kwatch/Makefile +++ b/mm/kwatch/Makefile @@ -1,3 +1,3 @@ obj-$(CONFIG_KWATCH) +=3D kwatch.o =20 -kwatch-y :=3D deref.o task_ctx.o hwbp.o +kwatch-y :=3D deref.o task_ctx.o hwbp.o probe.o diff --git a/mm/kwatch/probe.c b/mm/kwatch/probe.c new file mode 100644 index 000000000000..249aa50c9f78 --- /dev/null +++ b/mm/kwatch/probe.c @@ -0,0 +1,275 @@ +// SPDX-License-Identifier: GPL-2.0 +#include +#include +#include +#include +#include +#include + +#include "kwatch.h" +#define TRAMPOLINE_CHECK_DEPTH 16 +static DEFINE_PER_CPU(bool, kwatch_probe_cpu_muted); + +struct kwatch_probe_ctx { + struct kprobe kp; + struct kretprobe rp; + struct kprobe pin_kp; + const struct kwatch_config *cfg; + bool rp_via_int3; + + u32 epoch; +}; + +static struct kwatch_probe_ctx kwatch_probe_ctx; +static atomic_long_t kwatch_nmi_rejected; + +unsigned long kwatch_probe_nmi_rejected(void) +{ + return atomic_long_read(&kwatch_nmi_rejected); +} + +/* + * True if the probed function itself runs in an NMI-like context. + * int3-based kprobe delivery adds one NMI-like layer of its own; + * delivery is pinned at registration so the subtraction stays exact. + */ +static bool kwatch_probed_ctx_in_nmi(bool via_int3) +{ + return (preempt_count() & NMI_MASK) > (via_int3 ? NMI_OFFSET : 0); +} + +static void kwatch_pin_post_handler(struct kprobe *p, struct pt_regs *regs, + unsigned long flags) +{ + /* a post_handler pins the probepoint: no jump optimization */ +} + +bool kwatch_probe_validate_hit(struct pt_regs *regs, + struct task_struct *arm_tsk) +{ + struct kwatch_tsk_ctx *ctx =3D kwatch_tsk_ctx_get(false); + const struct kwatch_config *cfg =3D kwatch_probe_ctx.cfg; + + if (unlikely(!ctx || !cfg)) + return true; + + if (arm_tsk !=3D current || ctx->depth !=3D cfg->depth + 1) + return true; + + return false; +} + +void kwatch_probe_mute(bool mute) +{ + __this_cpu_write(kwatch_probe_cpu_muted, mute); +} + +static inline bool kwatch_probe_is_muted(void) +{ + return __this_cpu_read(kwatch_probe_cpu_muted); +} + +enum kwatch_probe_position { + KWATCH_PROBE_POSITION_ENTRY, + KWATCH_PROBE_POSITION_ACTIVE, + KWATCH_PROBE_POSITION_EXIT +}; + +static bool kwatch_tsk_ctx_check(enum kwatch_probe_position pos) +{ + struct kwatch_tsk_ctx *ctx =3D kwatch_tsk_ctx_get(true); + u32 epoch; + + if (unlikely(!ctx)) + return false; + + /* Pairs with smp_store_release() in kwatch_probe_start/stop() */ + epoch =3D smp_load_acquire(&kwatch_probe_ctx.epoch); + + if (unlikely(ctx->epoch !=3D epoch)) + kwatch_tsk_ctx_reset(ctx, epoch); + + if (unlikely(!epoch)) { + /* + * No active session (not yet published, or already stopped): + * kwatch_tsk_ctx_get(true) above may have just claimed a slot + * for current. Release it here, otherwise an entry that lands + * in the register->epoch-publish window leaks the slot until + * the pool is freed. + */ + kwatch_tsk_ctx_release(ctx); + return false; + } + + switch (pos) { + case KWATCH_PROBE_POSITION_ENTRY: + ctx->depth++; + return true; + case KWATCH_PROBE_POSITION_ACTIVE: + return true; + case KWATCH_PROBE_POSITION_EXIT: + if (unlikely(ctx->depth =3D=3D 0)) { + kwatch_tsk_ctx_put(); + return false; + } + + ctx->depth--; + if (ctx->depth =3D=3D 0) { + kwatch_tsk_ctx_put(); + return false; + } + return true; + } + return false; +} + +static int kwatch_activate_handler(struct kprobe *p, struct pt_regs *regs) +{ + struct kwatch_tsk_ctx *ctx =3D kwatch_tsk_ctx_get(false); + unsigned long watch_addr; + u16 watch_len; + + if (unlikely(!ctx)) + return 0; + + if (unlikely(kwatch_probe_is_muted())) + return 0; + + if (unlikely(!kwatch_tsk_ctx_check(KWATCH_PROBE_POSITION_ACTIVE))) + return 0; + + if (ctx->depth !=3D kwatch_probe_ctx.cfg->depth + 1 || ctx->wp) + return 0; + + if (kwatch_deref_resolve(kwatch_probe_ctx.cfg, regs, &watch_addr, + &watch_len)) + return 0; + + if (kwatch_hwbp_get(&ctx->wp)) + return 0; + + kwatch_hwbp_arm(ctx->wp, watch_addr, watch_len); + return 0; +} + +static int kwatch_lifecycle_entry(struct kretprobe_instance *ri, + struct pt_regs *regs) +{ + /* + * Single policy point: the target function's context is judged once + * here. A rejected invocation never increments depth, so the offset + * kprobe path inherits the verdict through the depth check. + */ + if (unlikely(kwatch_probed_ctx_in_nmi(kwatch_probe_ctx.rp_via_int3))) { + atomic_long_inc(&kwatch_nmi_rejected); + return 1; /* NMI context is unsupported: no window, no return hook */ + } + + if (!kwatch_tsk_ctx_check(KWATCH_PROBE_POSITION_ENTRY)) + return 0; + + if (kwatch_probe_ctx.cfg->func_offset =3D=3D 0) + kwatch_activate_handler(NULL, regs); + + return 0; +} + +static int kwatch_lifecycle_exit(struct kretprobe_instance *ri, + struct pt_regs *regs) +{ + struct kwatch_tsk_ctx *ctx =3D kwatch_tsk_ctx_get(false); + + if (unlikely(!ctx)) + return 0; + + if (!kwatch_tsk_ctx_check(KWATCH_PROBE_POSITION_EXIT)) + return 0; + + if (ctx->depth =3D=3D kwatch_probe_ctx.cfg->depth) { + struct kwatch_watchpoint *wp =3D xchg(&ctx->wp, NULL); + + if (wp) + kwatch_hwbp_put(wp); + } + + return 0; +} + +int kwatch_probe_start(struct kwatch_config *cfg) +{ + static u32 next_epoch; + u32 current_epoch; + int ret; + + /* + * Lockless check to prevent concurrent starts. Strictly serialized + * by the control plane mutex, but serves as a sanity check. + */ + if (smp_load_acquire(&kwatch_probe_ctx.epoch) !=3D 0) + return -EBUSY; + + memset(&kwatch_probe_ctx, 0, sizeof(kwatch_probe_ctx)); + kwatch_probe_ctx.cfg =3D cfg; + + /* Session-scoped, like arm_ipi_suppressed in kwatch_hwbp_prealloc() */ + atomic_long_set(&kwatch_nmi_rejected, 0); + + /* + * Pin the entry probepoint before the kretprobe registers, so its + * delivery (int3 vs ftrace) can never change under jump optimization. + * register_kretprobe() clears kp.post_handler, hence the companion. + */ + kwatch_probe_ctx.pin_kp.symbol_name =3D cfg->func_name; + kwatch_probe_ctx.pin_kp.post_handler =3D kwatch_pin_post_handler; + ret =3D register_kprobe(&kwatch_probe_ctx.pin_kp); + if (ret < 0) + return ret; + + kwatch_probe_ctx.rp.entry_handler =3D kwatch_lifecycle_entry; + kwatch_probe_ctx.rp.handler =3D kwatch_lifecycle_exit; + kwatch_probe_ctx.rp.kp.symbol_name =3D cfg->func_name; + + ret =3D register_kretprobe(&kwatch_probe_ctx.rp); + if (ret < 0) { + unregister_kprobe(&kwatch_probe_ctx.pin_kp); + return ret; + } + kwatch_probe_ctx.rp_via_int3 =3D !kprobe_ftrace(&kwatch_probe_ctx.rp.kp); + + if (cfg->func_offset) { + kwatch_probe_ctx.kp.symbol_name =3D cfg->func_name; + kwatch_probe_ctx.kp.offset =3D cfg->func_offset; + kwatch_probe_ctx.kp.pre_handler =3D kwatch_activate_handler; + + ret =3D register_kprobe(&kwatch_probe_ctx.kp); + if (ret) { + unregister_kretprobe(&kwatch_probe_ctx.rp); + unregister_kprobe(&kwatch_probe_ctx.pin_kp); + return ret; + } + } + + current_epoch =3D ++next_epoch; + if (unlikely(!current_epoch)) + current_epoch =3D ++next_epoch; + + /* Pairs with smp_load_acquire() in kwatch_tsk_ctx_check() */ + smp_store_release(&kwatch_probe_ctx.epoch, current_epoch); + + return 0; +} + +void kwatch_probe_stop(void) +{ + if (!kwatch_probe_ctx.epoch) + return; + + /* Pairs with smp_load_acquire() in kwatch_tsk_ctx_check() */ + smp_store_release(&kwatch_probe_ctx.epoch, 0); + + if (kwatch_probe_ctx.cfg->func_offset > 0) + unregister_kprobe(&kwatch_probe_ctx.kp); + + unregister_kretprobe(&kwatch_probe_ctx.rp); + unregister_kprobe(&kwatch_probe_ctx.pin_kp); +} --=20 2.53.0 From nobody Sat Jul 25 05:23:13 2026 Received: from mail-pg1-f172.google.com (mail-pg1-f172.google.com [209.85.215.172]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 563DD3F7A94 for ; Fri, 17 Jul 2026 13:06:30 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.172 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293591; cv=none; b=qqbUA6WXcIY0BgD0Sc6Z9vzmID6HOdRwWWJezqr5eshONvpPZGvkZXQmGPdbPavR28pvMIvu8wnFyhmaFQqL5M5PFsYc1V3VqwdKEkZkcBUR/w89XvRQ0pD4C8KUPgDltsxmdq80jRy/zy9gCsUgTAzglTnHi4cUj2i/NOvUnbI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293591; c=relaxed/simple; bh=EZAvJ5i0NW+Kg4ruSRBOsnEDn2AFbFSgVSv6jhiiLDM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=G0RFDLGdnY/Q2ykANKsBaAu+Zj/Be6VPi1EC7e6XdBe0Vl3aYAc2gepgzrRm9J3BbOf/GM629tGDssNFOE5LgFgW0Z0avcOflrLV+cmTmpMaXCGOSmEJOqj6M4YPpjt8OKptacPg7XNolhGVj1AImhNvdtJdlgpuYDQLXW46eO0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=TpsK9HBA; arc=none smtp.client-ip=209.85.215.172 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="TpsK9HBA" Received: by mail-pg1-f172.google.com with SMTP id 41be03b00d2f7-c9c26a5fb98so1161667a12.0 for ; Fri, 17 Jul 2026 06:06:30 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784293590; x=1784898390; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=TEU9n4OqT0ozu0Vj195yfeupUKeUqzKPFRIVzbMgbeQ=; b=TpsK9HBAckMRtpTGNKK8yuyMPojyb6uzVp8p1fIGX9YeSaX5J/N6tOJSyeRwdWFhQ3 B/p0BN2J1p47lNRzE7w7Si3ZQjMM7bC8dHknSHW5i6yC1sG42iVNe96s07glkHuzfR1o DeyoB0TeahpxfnxhCPHsiMF7BFIOcKFCoFuyJiHyu5s8LO8qT6Kp5QxHyfsTufpfj+yy Vv+lrXt0onqkPCUJ+qBbaqZZKh8+P5kCVKTLDs5Ey179EjuNE4aVG8n8dQc/O2c/nqgO P7kXcLDmu16OQisEatsq7OOvt9YSqnryThr2HbLNSHbZXtg0xL0avniR57xd4HLI+A9Q qw+w== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784293590; x=1784898390; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=TEU9n4OqT0ozu0Vj195yfeupUKeUqzKPFRIVzbMgbeQ=; b=c8P3CYfG8cf6x1Qe2A/7qf695PUeY1Z5A632jcQ8borLYX52J1S0ohDEYJlwUohE9Q f7paeXNAugGzb6ssBbM+GbGBEqzDyHDXvOoMvIhe0eunx/S9ZRhYRQCyF/KtTbz3wP6T fpm29rz43wyrJxpVvH+lM/eLGtXeektbf+HRsYXz4dnmCamAHzvyWoZzkBw3z7JDDonp AQlZKWbj7Cd1/gtgTOERGJs7ZQGccrhDftqhNQW/+G6fghCWnEUNx9PbH8YN1EUm8ffG 4VZQraV84NdT5+axQJ6H7qSI7xFcIqIRs1pMJnm0slwL8W3svCCeZuJxCqjQ+LxZrv8M gWBw== X-Forwarded-Encrypted: i=1; AHgh+RrlfzEjX3td5gQGMtKTrlnBjj6FOPsoaqfaQTruAEFCdqdR+7K0DDVewiIS5WsH3q41qCYniBOQM2ZaEII=@vger.kernel.org X-Gm-Message-State: AOJu0YyIvPekbyfqtead/yS1POiYqjtVMpuDrYEMsxp5QKaUmHE1gP8Z E8wSlHScrP41L7R15Szy0MfwhlUuwYDbQyIKVztMIuUJoz3Nf6ug0Jqo X-Gm-Gg: AfdE7cngzNL6/4mwhIXFKQPdKMTS/ZjVEux+M80ZHMCigLV4bezuT3GGb1fyf9rokNk U2IDrnK6AmnG6O7uFAiIA8Is4i5nmjz9ob5fx8zmosfeMMWGRR7jnqmNJbiqOVggSkjuwHqKgiq j+s8bGpV02SzKb6qsWmNefeEDVOvWlNgWED6Z0TufRQOQzFM/OKnnSI9kfQxV8Wzjikw3UUKQHW /AW5pyHZHh9cX9+CTQxUOeJy80CnEyt3XoJVmnIQ2BZfBzntU/dQkbi5bWQpPf1st5MKxeh3WNz T6GsCC1BfCgtSzDsChCvKL1Eh0GPGIN7boyGW2a7cbhTTY4DOGRKhJXFB0rb8GNFEZWfLqVuBj2 RsQ4L8JY2BQKSptnZXqGxrjKKa13ZxQL5x3HstqBexJVrnhhks0RYXJQA3wdofHZa7lb7fkoxM3 x8VA== X-Received: by 2002:a05:6a20:430f:b0:3c3:8315:80b6 with SMTP id adf61e73a8af0-3c38dbff098mr7615722637.32.1784293589583; Fri, 17 Jul 2026 06:06:29 -0700 (PDT) Received: from localhost ([144.24.58.22]) by smtp.gmail.com with ESMTPSA id a92af1059eb24-13ce29ffb13sm5473497c88.6.2026.07.17.06.06.24 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 17 Jul 2026 06:06:25 -0700 (PDT) From: Jinchao Wang To: Andrew Morton , Peter Zijlstra , Thomas Gleixner , Steven Rostedt , Masami Hiramatsu Cc: Ingo Molnar , Borislav Petkov , Dave Hansen , "H . Peter Anvin" , x86@kernel.org, Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Mathieu Desnoyers , David Hildenbrand , Jonathan Corbet , Matthew Wilcox , Alan Stern , Randy Dunlap , Alexander Potapenko , Marco Elver , Mike Rapoport , linux-kernel@vger.kernel.org, linux-mm@kvack.org, linux-trace-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org, linux-doc@vger.kernel.org, Jinchao Wang Subject: [RFC PATCH v2 10/13] mm/kwatch: add anchor thread for global watchpoints Date: Fri, 17 Jul 2026 09:06:09 -0400 Message-ID: <20260717130609.1903443-1-wangjinchao600@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260717125023.1895892-1-wangjinchao600@gmail.com> References: <20260717125023.1895892-1-wangjinchao600@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Global variables have no function whose execution can bound the watch window. Provide one: a kernel thread sleeps for the configured duration inside a dedicated noinline function, kwatch_global_anchor(), and the probe runtime hooks that function like any other target. When the duration expires the thread schedules a work item that tears the session down; the expired flag is cleared under the control-plane mutex so a stale work item from a previous session cannot stop a new one. Signed-off-by: Jinchao Wang --- mm/kwatch/Makefile | 2 +- mm/kwatch/anchor.c | 85 ++++++++++++++++++++++++++++++++++++++++++++++ 2 files changed, 86 insertions(+), 1 deletion(-) create mode 100644 mm/kwatch/anchor.c diff --git a/mm/kwatch/Makefile b/mm/kwatch/Makefile index f04673cc5b1c..b196c794619a 100644 --- a/mm/kwatch/Makefile +++ b/mm/kwatch/Makefile @@ -1,3 +1,3 @@ obj-$(CONFIG_KWATCH) +=3D kwatch.o =20 -kwatch-y :=3D deref.o task_ctx.o hwbp.o probe.o +kwatch-y :=3D deref.o task_ctx.o hwbp.o probe.o anchor.o diff --git a/mm/kwatch/anchor.c b/mm/kwatch/anchor.c new file mode 100644 index 000000000000..e87eb5e813ff --- /dev/null +++ b/mm/kwatch/anchor.c @@ -0,0 +1,85 @@ +// SPDX-License-Identifier: GPL-2.0 +#define pr_fmt(fmt) KBUILD_MODNAME ": " fmt + +#include +#include +#include +#include +#include + +#include "kwatch.h" + +static DECLARE_WAIT_QUEUE_HEAD(kwatch_anchor_wq); +static struct task_struct *kwatch_anchor_tsk; +static bool kwatch_anchor_expired; + +bool kwatch_anchor_has_expired(void) +{ + return READ_ONCE(kwatch_anchor_expired); +} + +void kwatch_anchor_clear_expired(void) +{ + WRITE_ONCE(kwatch_anchor_expired, false); +} + +static void kwatch_auto_stop_handler(struct work_struct *work) +{ + kwatch_auto_stop(); +} + +static DECLARE_WORK(kwatch_auto_stop_work, kwatch_auto_stop_handler); + +noinline void kwatch_global_anchor(unsigned long duration_sec) +{ + /* TASK_IDLE: a long timed sleep must not inflate loadavg or trip the + * hung-task detector the way TASK_UNINTERRUPTIBLE would. + */ + wait_event_idle_timeout(kwatch_anchor_wq, kthread_should_stop(), + duration_sec * HZ); +} + +static int kwatch_anchor_thread_fn(void *data) +{ + unsigned long duration =3D (unsigned long)data; + + kwatch_global_anchor(duration); + + if (!kthread_should_stop()) { + /* mark before scheduling; cleared under the control mutex */ + WRITE_ONCE(kwatch_anchor_expired, true); + schedule_work(&kwatch_auto_stop_work); + } + + while (!kthread_should_stop()) + schedule_timeout_idle(HZ); + + return 0; +} + +int kwatch_anchor_start(u16 duration) +{ + kwatch_anchor_tsk =3D kthread_run(kwatch_anchor_thread_fn, + (void *)(unsigned long)duration, + "kwatch_anchor"); + if (IS_ERR(kwatch_anchor_tsk)) { + int ret =3D PTR_ERR(kwatch_anchor_tsk); + + kwatch_anchor_tsk =3D NULL; + return ret; + } + return 0; +} + +void kwatch_anchor_stop(void) +{ + if (kwatch_anchor_tsk) { + kthread_stop(kwatch_anchor_tsk); + kwatch_anchor_tsk =3D NULL; + } +} + +void kwatch_anchor_cancel_work(void) +{ + cancel_work_sync(&kwatch_auto_stop_work); +} --=20 2.53.0 From nobody Sat Jul 25 05:23:13 2026 Received: from mail-pg1-f170.google.com (mail-pg1-f170.google.com [209.85.215.170]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A8CD13FBEBC for ; Fri, 17 Jul 2026 13:06:55 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.170 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293617; cv=none; b=p1mvSRZ3OZ2uRN7rrzq7+ODCkuJB4Co0P9XQdG48mIr2H91DVvzQwvGSCp/KFpdm3RtpE4oIkuzjl3VU/3HbVb8zsBMyNkZ8JqTHYMVwIWnG7oE8JuD6vsOmHoZeT8jEM+ti+zoIRav+NCwMglBvT16RfnWuio4wCigU/Tpi4jQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293617; c=relaxed/simple; bh=mSZSjX34G0DYKny1FJRnV53asv5EFLj78ugJoIQTb8I=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Kq0+BsUcUKOjwVS66zPnDxSKpe6fq/xB9zUsbnbjmR81KaRs7caHpTDHskcIGskiY3LfnT19M7MYKxKXnXake92zmos4wxsXeQH1P8mBIihW4B3SVf5z9fLJS3drzMLFuSTkXFf++0RD673CbzxHALWsuLzRR1nYEbeJS1Pg1Lk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=SPkfJx8M; arc=none smtp.client-ip=209.85.215.170 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="SPkfJx8M" Received: by mail-pg1-f170.google.com with SMTP id 41be03b00d2f7-ca88130e09aso5102029a12.3 for ; Fri, 17 Jul 2026 06:06:55 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784293615; x=1784898415; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=pe2aLEJuh1t7h3+VwKXqsCr8PrbeTbQwnX+mP9fCQIw=; b=SPkfJx8MFuIykPcL/5FcwNzS8Me8HEZ8Xhrn4PnbclqpX6qfkJlEjjwzU5iuLM8IVn u/szNJKS8rd3ddwvWrUR8CSgGOw31oVW1/dFXwps7bOmocHIxSRDjjUB1awMZT08MJ7u EVACB22NuTCo8oVk8pUBjpgi7/1u5dldj2HdyfOHhRVCMjrJYpTFyy1kXfCWwQ4w6G2z T65gC1DlnF3XfHrU12ugXcFswkshdkqxcVxuoEJWOomFMfFCaf83uch89L9t2MiQLO11 lzCiCSchQzDqvrh60IR6/Jz9CzQGh1Gtt7IyipN9QU9qQjl1JFJYBuPAGurAkL3HH+37 g5kw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784293615; x=1784898415; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=pe2aLEJuh1t7h3+VwKXqsCr8PrbeTbQwnX+mP9fCQIw=; b=IkR+AorxyIVA7kiNQ0/B7bIvV+ErnahmNvL4LaukOfA62mx/1PPEfU83kjyscQJfIZ 6TF/lgExKciPbCZhuSlI7ERFRXYhf7LeTw0MV1sI79QaAMLNfVEb2WSjZ66TX5gl+wqZ ahKPnKBekkqhy+C5Q/lJq4pRbnEjnQzBUanj4EUhva+IMawjoKwKUtSGck0+iVBc8/XU px+TzeHhE1mOGP8n2K2YNIpyz/6xfvXyMsEtz405J07JZ3eV/MV3f+CRCWvp/jn/O/Bc 4tkh9Dg+fUWeDnHEUWTyeWzvGVVW+LuzxyGi8DR8ze7mwA4vNIPPJdcacGeS2N+aw5+j ke+w== X-Forwarded-Encrypted: i=1; AHgh+Rp7VR6gV4dz//6AhejVitr+S4f2jHjKwnOotOo37Y8KPqmHsoXSV1LpjIPSLGWhrxrTlDl2UAjHqgpx78E=@vger.kernel.org X-Gm-Message-State: AOJu0YzTeaQLXjGyQwtoy6KBTY6Tjl6J6GGwsohCC03a8uwTupbJEKhj fgJ+HcWQtw2VWcwXhVWiBsW7tNoQhp6EwRsRdEG3H1zHfoSntW8iB4gm X-Gm-Gg: AfdE7ckdFvCw3PBu0ROseD4G5q3/DQrCix88wDTrvAs5MWb/KKjp7yotiWEImTyZ0l9 Gg2mQFbFHNBP43DxqaQf3PwzW8g9Mv/V9uwrSk68oc9BPpZdSAScKfFgmrnWSDaqU5E21N+IoWN yFt6XjySZVHi+pHezVYNedTCnvt2O11fxNu0/mV88wo9XT8YJ/7QWybYCwUVyjP/y0bJEVOydQ2 beym9wZTMnkHtrxbBUFsmoEF62yKqVkXdj1p7JdaGcompkCxRR+GokZ/xUIMWDvtRd+FzZoqB8k QmgX1qAaXmq5NUGjZE0lZ+fgo5hb+hRIeIXpYxUxLaEed1XyRnzLlUSDDZ/jalYjeP7GzSFeTaf 35Qv9RkF2YoYrybWejMsOgK6pAxK3l80SIJ2YLbCmPi3zU5YCeaqzm44lGTsOy0lXNISMa79EIq A98w== X-Received: by 2002:a17:90b:2dc1:b0:37f:fdc8:71b4 with SMTP id 98e67ed59e1d1-38e4b3e1424mr2506032a91.2.1784293614807; Fri, 17 Jul 2026 06:06:54 -0700 (PDT) Received: from localhost ([144.24.58.22]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-3142a1bbda9sm6885623eec.19.2026.07.17.06.06.52 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 17 Jul 2026 06:06:53 -0700 (PDT) From: Jinchao Wang To: Andrew Morton , Peter Zijlstra , Thomas Gleixner , Steven Rostedt , Masami Hiramatsu Cc: Ingo Molnar , Borislav Petkov , Dave Hansen , "H . Peter Anvin" , x86@kernel.org, Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Mathieu Desnoyers , David Hildenbrand , Jonathan Corbet , Matthew Wilcox , Alan Stern , Randy Dunlap , Alexander Potapenko , Marco Elver , Mike Rapoport , linux-kernel@vger.kernel.org, linux-mm@kvack.org, linux-trace-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org, linux-doc@vger.kernel.org, Jinchao Wang Subject: [RFC PATCH v2 11/13] mm/kwatch: add debugfs control plane Date: Fri, 17 Jul 2026 09:06:37 -0400 Message-ID: <20260717130637.1903667-1-wangjinchao600@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260717125023.1895892-1-wangjinchao600@gmail.com> References: <20260717125023.1895892-1-wangjinchao600@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Wire the pieces together behind a single debugfs file, /sys/kernel/debug/kwatch/config. Writing a key=3Dvalue configuration string stops any active session and starts a new one; reading shows the active configuration and the nmi_rejected counter. An open-count guard keeps the file single-open and a mutex serializes start/stop/auto-stop against each other. Add the Kconfig entry and hook mm/kwatch into the mm build. KWatch can be built in or as a module; symbol-name watch expressions need the built-in flavour (kallsyms_lookup_name is not exported). Signed-off-by: Jinchao Wang --- MAINTAINERS | 8 ++ mm/Kconfig | 1 + mm/Makefile | 1 + mm/kwatch/Kconfig | 16 +++ mm/kwatch/Makefile | 2 +- mm/kwatch/core.c | 324 +++++++++++++++++++++++++++++++++++++++++++++ 6 files changed, 351 insertions(+), 1 deletion(-) create mode 100644 mm/kwatch/Kconfig create mode 100644 mm/kwatch/core.c diff --git a/MAINTAINERS b/MAINTAINERS index 7cc4bca5a2c5..b6371f92fe5c 100644 --- a/MAINTAINERS +++ b/MAINTAINERS @@ -14578,6 +14578,14 @@ S: Supported T: git git://git.kernel.org/pub/scm/virt/kvm/kvm.git F: arch/x86/kvm/xen.* =20 +KWATCH +M: Jinchao Wang +L: linux-mm@kvack.org +S: Maintained +F: Documentation/dev-tools/kwatch.rst +F: include/trace/events/kwatch.h +F: mm/kwatch/ + L3MDEV M: David Ahern L: netdev@vger.kernel.org diff --git a/mm/Kconfig b/mm/Kconfig index 9e0ca4824905..cac75a46e21a 100644 --- a/mm/Kconfig +++ b/mm/Kconfig @@ -1510,5 +1510,6 @@ config LAZY_MMU_MODE_KUNIT_TEST If unsure, say N. =20 source "mm/damon/Kconfig" +source "mm/kwatch/Kconfig" =20 endmenu diff --git a/mm/Makefile b/mm/Makefile index eff9f9e7e061..80c688330358 100644 --- a/mm/Makefile +++ b/mm/Makefile @@ -92,6 +92,7 @@ obj-$(CONFIG_PAGE_POISONING) +=3D page_poison.o obj-$(CONFIG_KASAN) +=3D kasan/ obj-$(CONFIG_KFENCE) +=3D kfence/ obj-$(CONFIG_KMSAN) +=3D kmsan/ +obj-$(CONFIG_KWATCH) +=3D kwatch/ obj-$(CONFIG_FAILSLAB) +=3D failslab.o obj-$(CONFIG_FAIL_PAGE_ALLOC) +=3D fail_page_alloc.o obj-$(CONFIG_MEMTEST) +=3D memtest.o diff --git a/mm/kwatch/Kconfig b/mm/kwatch/Kconfig new file mode 100644 index 000000000000..9daf6d4463ef --- /dev/null +++ b/mm/kwatch/Kconfig @@ -0,0 +1,16 @@ +config KWATCH + tristate "Kernel Watch Framework" + depends on PERF_EVENTS && HAVE_HW_BREAKPOINT && DEBUG_FS + depends on HAVE_REINSTALL_HW_BREAKPOINT + depends on KPROBES && KRETPROBES + depends on STACKTRACE + help + A generalized hardware-assisted memory monitor utility. + It provides a low-overhead, real-time trigger mechanism to monitor + kernel memory safely in atomic contexts using hardware breakpoints. + + KWatch is designed to catch silent memory corruptions, stack + overwrites, and complex Heisenbugs by synchronously trapping the + exact instruction causing the illegal access. + + If unsure, say N. diff --git a/mm/kwatch/Makefile b/mm/kwatch/Makefile index b196c794619a..02d7917602f1 100644 --- a/mm/kwatch/Makefile +++ b/mm/kwatch/Makefile @@ -1,3 +1,3 @@ obj-$(CONFIG_KWATCH) +=3D kwatch.o =20 -kwatch-y :=3D deref.o task_ctx.o hwbp.o probe.o anchor.o +kwatch-y :=3D core.o deref.o task_ctx.o hwbp.o probe.o anchor.o diff --git a/mm/kwatch/core.c b/mm/kwatch/core.c new file mode 100644 index 000000000000..d8526d5aae5c --- /dev/null +++ b/mm/kwatch/core.c @@ -0,0 +1,324 @@ +// SPDX-License-Identifier: GPL-2.0 +#define pr_fmt(fmt) KBUILD_MODNAME ": " fmt + +#include +#include +#include +#include +#include +#include +#include +#include +#include "kwatch.h" + +static struct kwatch_config kwatch_config; +static bool watching_active; + +static struct dentry *dbgfs_dir; +static struct dentry *dbgfs_config; +static DEFINE_MUTEX(kwatch_dbgfs_mutex); +static atomic_t dbgfs_config_busy =3D ATOMIC_INIT(0); + +static int kwatch_start_watching(void) +{ + int ret; + + if (!strlen(kwatch_config.func_name)) { + if (kwatch_config.duration > 0) { + strscpy(kwatch_config.func_name, "kwatch_global_anchor", + sizeof(kwatch_config.func_name)); + } else { + pr_err("func_name or duration is required\n"); + return -EINVAL; + } + } else if (kwatch_config.duration > 0 && + strcmp(kwatch_config.func_name, "kwatch_global_anchor")) { + pr_warn("duration is ignored when watching a specific function\n"); + } + + ret =3D kwatch_hwbp_prealloc(kwatch_config.max_watch); + if (ret) { + pr_err("kwatch_hwbp_prealloc ret: %d\n", ret); + return ret; + } + + ret =3D kwatch_tsk_ctx_prealloc(kwatch_config.max_concurrency); + if (ret) { + kwatch_hwbp_free(); + return ret; + } + + ret =3D kwatch_probe_start(&kwatch_config); + if (ret) { + pr_err("kwatch_probe_start ret: %d\n", ret); + kwatch_tsk_ctx_free(); + kwatch_hwbp_free(); + return ret; + } + + if (!strcmp(kwatch_config.func_name, "kwatch_global_anchor")) { + ret =3D kwatch_anchor_start(kwatch_config.duration); + if (ret) { + kwatch_probe_stop(); + synchronize_rcu(); + kwatch_tsk_ctx_release_wps(); + kwatch_hwbp_free(); + kwatch_tsk_ctx_free(); + return ret; + } + } + + watching_active =3D true; + return 0; +} + +static void kwatch_stop_watching(void) +{ + watching_active =3D false; + + kwatch_anchor_stop(); + /* after kthread_stop: the dead thread cannot re-mark expiry */ + kwatch_anchor_clear_expired(); + + kwatch_probe_stop(); + synchronize_rcu(); + kwatch_tsk_ctx_release_wps(); + /* + * Waits for disarm IPIs and unregisters breakpoints: no #DB can + * reach the ctx pool once this returns. + */ + kwatch_hwbp_free(); + kwatch_tsk_ctx_free(); +} + +void kwatch_auto_stop(void) +{ + mutex_lock(&kwatch_dbgfs_mutex); + /* the expired check neutralizes work items from torn-down sessions */ + if (watching_active && kwatch_anchor_has_expired()) { + kwatch_stop_watching(); + pr_info("watch duration expired, stopped watching\n"); + } + mutex_unlock(&kwatch_dbgfs_mutex); +} + +static int kwatch_config_parse(char *buf, struct kwatch_config *cfg) +{ + char *token, *key, *val; + int ret =3D 0; + + memset(cfg, 0, sizeof(*cfg)); + cfg->max_concurrency =3D 256; + cfg->max_watch =3D 4; + cfg->watch_len =3D 8; + + while ((token =3D strsep(&buf, " \t\n")) !=3D NULL) { + if (!*token) + continue; + key =3D strsep(&token, "=3D"); + val =3D token; + if (!key || !val) + return -EINVAL; + + if (!strcmp(key, "func_name")) { + strscpy(cfg->func_name, val, sizeof(cfg->func_name)); + } else if (!strcmp(key, "func_offset")) { + ret =3D kstrtou16(val, 0, &cfg->func_offset); + } else if (!strcmp(key, "depth")) { + ret =3D kstrtou16(val, 0, &cfg->depth); + } else if (!strcmp(key, "max_concurrency")) { + ret =3D kstrtou16(val, 0, &cfg->max_concurrency); + } else if (!strcmp(key, "max_watch")) { + ret =3D kstrtou16(val, 0, &cfg->max_watch); + } else if (!strcmp(key, "watch_len")) { + ret =3D kstrtou16(val, 0, &cfg->watch_len); + if (!ret && cfg->watch_len !=3D 1 && + cfg->watch_len !=3D 2 && cfg->watch_len !=3D 4 && + cfg->watch_len !=3D 8) + ret =3D -EINVAL; + } else if (!strcmp(key, "duration")) { + ret =3D kstrtou16(val, 0, &cfg->duration); + } else if (!strcmp(key, "watch_expr")) { + strscpy(cfg->watch_expr, val, sizeof(cfg->watch_expr)); + ret =3D kwatch_deref_parse(cfg, val); + } + + if (ret) + return ret; + } + return 0; +} + +static int kwatch_dbgfs_open(struct inode *inode, struct file *file) +{ + if (atomic_cmpxchg(&dbgfs_config_busy, 0, 1)) + return -EBUSY; + return 0; +} + +static int kwatch_dbgfs_release(struct inode *inode, struct file *file) +{ + atomic_set(&dbgfs_config_busy, 0); + return 0; +} + +static ssize_t kwatch_dbgfs_read(struct file *file, char __user *user_buf, + size_t count, loff_t *ppos) +{ + char *out_buf; + size_t len =3D 0; + ssize_t ret; + + out_buf =3D kzalloc(MAX_CONFIG_STR_LEN, GFP_KERNEL); + if (!out_buf) + return -ENOMEM; + + /* + * Serialize against the write path and the auto-stop work item so the + * config snapshot cannot tear or race a session teardown. + */ + mutex_lock(&kwatch_dbgfs_mutex); + + if (watching_active) { + len +=3D scnprintf(out_buf + len, MAX_CONFIG_STR_LEN - len, + "func_name=3D%s\n" + "func_offset=3D%u\n" + "depth=3D%u\n" + "duration=3D%u\n" + "max_concurrency=3D%u\n" + "max_watch=3D%u\n" + "watch_len=3D%u\n", + kwatch_config.func_name, + kwatch_config.func_offset, kwatch_config.depth, + kwatch_config.duration, + kwatch_config.max_concurrency, + kwatch_config.max_watch, + kwatch_config.watch_len); + + if (kwatch_config.base =3D=3D KWATCH_BASE_GLOBAL_SYM) { + len +=3D scnprintf(out_buf + len, MAX_CONFIG_STR_LEN - len, + "sym_addr=3D0x%lx\n", kwatch_config.sym_addr); + } + + len +=3D scnprintf(out_buf + len, MAX_CONFIG_STR_LEN - len, + "watch_expr=3D%s\n" + "nmi_rejected=3D%lu\n" + "arm_ipi_suppressed=3D%lu\n", + kwatch_config.watch_expr, + kwatch_probe_nmi_rejected(), + kwatch_hwbp_arm_ipi_suppressed()); + } else { + len =3D scnprintf(out_buf, MAX_CONFIG_STR_LEN, "not watching\n"); + } + + mutex_unlock(&kwatch_dbgfs_mutex); + + ret =3D simple_read_from_buffer(user_buf, count, ppos, out_buf, len); + kfree(out_buf); + return ret; +} + +static ssize_t kwatch_dbgfs_write(struct file *file, const char __user *bu= ffer, + size_t count, loff_t *ppos) +{ + char *input_alloc; + char *parse_str; + int ret; + + if (count =3D=3D 0 || count >=3D MAX_CONFIG_STR_LEN) + return -EINVAL; + + input_alloc =3D memdup_user_nul(buffer, count); + if (IS_ERR(input_alloc)) + return PTR_ERR(input_alloc); + + mutex_lock(&kwatch_dbgfs_mutex); + + if (watching_active) + kwatch_stop_watching(); + + parse_str =3D strim(input_alloc); + + if (!strlen(parse_str)) { + ret =3D -EINVAL; + goto out; + } + + ret =3D kwatch_config_parse(parse_str, &kwatch_config); + if (ret) { + pr_err("Failed to parse config %d\n", ret); + goto out; + } + + ret =3D kwatch_start_watching(); + if (ret) { + pr_err("Failed to start watching with %d\n", ret); + goto out; + } + + ret =3D count; + +out: + mutex_unlock(&kwatch_dbgfs_mutex); + kfree(input_alloc); + return ret; +} + +static const struct file_operations kwatch_fops =3D { + .owner =3D THIS_MODULE, + .open =3D kwatch_dbgfs_open, + .release =3D kwatch_dbgfs_release, + .read =3D kwatch_dbgfs_read, + .write =3D kwatch_dbgfs_write, +}; + +static int __init kwatch_init(void) +{ + int ret =3D 0; + + memset(&kwatch_config, 0, sizeof(kwatch_config)); + + dbgfs_dir =3D debugfs_create_dir("kwatch", NULL); + if (IS_ERR(dbgfs_dir)) { + ret =3D PTR_ERR(dbgfs_dir); + goto err_dir; + } + + dbgfs_config =3D debugfs_create_file("config", 0600, dbgfs_dir, NULL, + &kwatch_fops); + if (IS_ERR(dbgfs_config)) { + ret =3D PTR_ERR(dbgfs_config); + goto err_file; + } + + pr_info("module loaded\n"); + return 0; + +err_file: + debugfs_remove_recursive(dbgfs_dir); + dbgfs_dir =3D NULL; +err_dir: + return ret; +} +module_init(kwatch_init); + +static void __exit kwatch_exit(void) +{ + mutex_lock(&kwatch_dbgfs_mutex); + if (watching_active) + kwatch_stop_watching(); + mutex_unlock(&kwatch_dbgfs_mutex); + + /* the anchor thread is dead: nothing can schedule new work now */ + kwatch_anchor_cancel_work(); + + debugfs_remove_recursive(dbgfs_dir); + dbgfs_dir =3D NULL; + + pr_info("kwatch unloaded\n"); +} +module_exit(kwatch_exit); + +MODULE_AUTHOR("Jinchao Wang "); +MODULE_DESCRIPTION("Kernel watchpoint"); +MODULE_LICENSE("GPL"); --=20 2.53.0 From nobody Sat Jul 25 05:23:13 2026 Received: from mail-pj1-f42.google.com (mail-pj1-f42.google.com [209.85.216.42]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3E0C73E6DE3 for ; Fri, 17 Jul 2026 13:07:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.42 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293641; cv=none; b=sVmXo9mImnOHAN8s+C0WmgywVL4tuar25abFFMqKLtVNLAFETLug8tVQRTcZEFQO38JNqSANPtUASCJhuJJj3UZq+fEiiaRv8Xmgl1Du9LuuxPHO6jaEFgWkCAWLgGL02QN5sJ54NfuosJxPlnjuvG7o+WDoYRzrT4ck31nOZlw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293641; c=relaxed/simple; bh=1TWFwrqLLPhi4hOF+lQYiAR8gfx3VWRwxzpCtpmYYnQ=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=LQ4321db1n7ZvI95f1YTm1I1vGctHBXuruXoDDWQ94t+YpochtRD81b36kDzgw23QnvOSdZuwUmYC6+fCr3i7o6j5t7R16HNU3pm/HYw7euV3MiBIa9Am2/osr1qyMWqTEHCku6VB7fG0FP3Pu8ubtUeE/TGDA3gw5HWtEMwJLQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=CginYwjn; arc=none smtp.client-ip=209.85.216.42 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="CginYwjn" Received: by mail-pj1-f42.google.com with SMTP id 98e67ed59e1d1-381216921aaso7762991a91.1 for ; Fri, 17 Jul 2026 06:07:20 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784293640; x=1784898440; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=GlWa8mIQcYhfASTlpT5AvSeo0bnNLaBVHf432AwkEfc=; b=CginYwjn34ITNmcQnAuTKlrIxoc+eFRpVQEKNnynWaFvrgPt6LJ2+N2D6oOF4i6tXS 8Q66nvEBWkifMoBgkbBprsAXDF3vh3Ecfrk+CvR8IqZd5rfsMEZok+uwsh5BnqSxpq9M COpQnWnSC/vpMsDjiTw0pCoMwDmcXoNgsOVf08gRfL4KTTVHTBDh2OXEwX6QNy7xhcp5 /RvJs0dICzmQaF6oes1IQN6Jb/fmcYDClFTiVtPXbyB0JX1xEuHCfwiJRpwNja8/bwhJ gUqgEPSKrnTFQXnxo6/T3vAAY99CkZ43SqnNyXHjaP4p9Cy68C52yASkK1Z4ufQ0D/MG DSjQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784293640; x=1784898440; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=GlWa8mIQcYhfASTlpT5AvSeo0bnNLaBVHf432AwkEfc=; b=rEoQGIsY1MEicuY0ImsNYIoY1G34YSBAXODXVMkjAOCFqkDGEGJhCshu3aL4fQDaGH sPCsJrq7c2LyNUv1OEC9PAqZyPFeA8Uibluf7/LEFLGIb/pvKyZfVxjl6wfdBd9pbRVM T6/JrT0grNsRYt9Y2sEoT/6rMZ71eVRYwFFGqtKctRYaz/sf8R8/GDWFau5q0EUFLMO5 1+Qnv0BQ9yZr7K52Cc60ciQmKvjOLvLfkbbCKK5phJdx0jagRU2wMsXuoaRHrh53qzH/ J63vFsNiYUbs11sjiiJ2OwE0GEW3iRyN8wPD+WEpVKV887/2xaBQXpyjoOU66AYyMnB6 13ng== X-Forwarded-Encrypted: i=1; AHgh+RrK+NS8+RoErQP8wU3JFVLvNN+rxiXLOn/rFo1aM1qcIOi+aWjmhQhu4EUL8i0f64Cq1cUYRXJQhwEwlmM=@vger.kernel.org X-Gm-Message-State: AOJu0YxYc3d/Ws7hy5+4Jw8ufMcaGN8Xlal6sRY3b0ks5fganhEWVQXd EbHOn9AmvqrsFCMN46bgzKqnknZx6+9NpuJ7h/zoCMhRDx5VZx8+X9hI X-Gm-Gg: AfdE7cmXNUJGt2cvWXpRMLSX67Dl6WdXCkXoKyiXpg2bnmTCpvfdLOgz3VUnoa+e8pw g0ZdZDExi78npA+QMI2XC4u6NIevWZZ/9TiODIqBerxyyzudrWgAq2sFz3IDhMXKYNOCDhzDMWH 1zZ8/9POFnAgoBHtPr6nmjHcAyjYTs40tQTTX4tPBXZODaa41rdnS8seam1OBHB6sxMMRupoeEg rmqucPeCTZlNJXfeH3ELfqt1Qo8oYJUAfvAYr2S0A6XzXKP6RbIHiVDHUwETnqRCch42N2WXyVu zMuDJs2nCyGer5d9UZXQ5wWxm7E0oC/DzJc4oWVym4hD+ZfFsdICGNprTV26FlNUWaMOKkWwDji fo7K2IGrm8OWsiFP+UBRtcx2xBphq1vAZhHCPuQLfdXWSrcWxzIYlFjlQo49C+1gLU/G5veiS7G YG70r8pl8SkpRA X-Received: by 2002:a17:90b:2dc1:b0:37f:c97a:939f with SMTP id 98e67ed59e1d1-38e4b3e11a5mr2543730a91.7.1784293639476; Fri, 17 Jul 2026 06:07:19 -0700 (PDT) Received: from localhost ([144.24.58.22]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-31429cf22b3sm7709544eec.0.2026.07.17.06.07.17 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 17 Jul 2026 06:07:18 -0700 (PDT) From: Jinchao Wang To: Andrew Morton , Peter Zijlstra , Thomas Gleixner , Steven Rostedt , Masami Hiramatsu Cc: Ingo Molnar , Borislav Petkov , Dave Hansen , "H . Peter Anvin" , x86@kernel.org, Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Mathieu Desnoyers , David Hildenbrand , Jonathan Corbet , Matthew Wilcox , Alan Stern , Randy Dunlap , Alexander Potapenko , Marco Elver , Mike Rapoport , linux-kernel@vger.kernel.org, linux-mm@kvack.org, linux-trace-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org, linux-doc@vger.kernel.org, Jinchao Wang Subject: [RFC PATCH v2 12/13] mm/kwatch: add KUnit tests for the watch expression parser Date: Fri, 17 Jul 2026 09:07:02 -0400 Message-ID: <20260717130702.1903917-1-wangjinchao600@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260717125023.1895892-1-wangjinchao600@gmail.com> References: <20260717125023.1895892-1-wangjinchao600@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Cover base anchors (stack, argN, absolute address), positive and negative offsets, dereference chains, and rejection of malformed expressions (missing offsets, bad argument index, junk offsets). Signed-off-by: Jinchao Wang --- mm/kwatch/.kunitconfig | 9 +++ mm/kwatch/Kconfig | 12 ++++ mm/kwatch/Makefile | 1 + mm/kwatch/deref_test.c | 146 +++++++++++++++++++++++++++++++++++++++++ 4 files changed, 168 insertions(+) create mode 100644 mm/kwatch/.kunitconfig create mode 100644 mm/kwatch/deref_test.c diff --git a/mm/kwatch/.kunitconfig b/mm/kwatch/.kunitconfig new file mode 100644 index 000000000000..7e977ddf0da1 --- /dev/null +++ b/mm/kwatch/.kunitconfig @@ -0,0 +1,9 @@ +CONFIG_KUNIT=3Dy +CONFIG_KWATCH=3Dy +CONFIG_KWATCH_KUNIT_TEST=3Dy +CONFIG_PERF_EVENTS=3Dy +CONFIG_HAVE_HW_BREAKPOINT=3Dy +CONFIG_HAVE_REINSTALL_HW_BREAKPOINT=3Dy +CONFIG_KPROBES=3Dy +CONFIG_KRETPROBES=3Dy +CONFIG_PRINTK=3Dy diff --git a/mm/kwatch/Kconfig b/mm/kwatch/Kconfig index 9daf6d4463ef..6ec9aa448ece 100644 --- a/mm/kwatch/Kconfig +++ b/mm/kwatch/Kconfig @@ -14,3 +14,15 @@ config KWATCH exact instruction causing the illegal access. =20 If unsure, say N. + +config KWATCH_KUNIT_TEST + bool "KUnit tests for KWatch" if !KUNIT_ALL_TESTS + # Built into the kwatch module, so it must be y; a bool cannot be + # enabled when KWATCH is a module (KWATCH=3Dm would force it off). + depends on KWATCH=3Dy && KUNIT + default KUNIT_ALL_TESTS + help + Enable KUnit tests for the KWatch kernel module. + This suite tests the core parsing logic, the pointer-chasing + finite state machine, and edge cases involving complex watchpoint + expressions. If unsure, say N. diff --git a/mm/kwatch/Makefile b/mm/kwatch/Makefile index 02d7917602f1..1d223d73b461 100644 --- a/mm/kwatch/Makefile +++ b/mm/kwatch/Makefile @@ -1,3 +1,4 @@ obj-$(CONFIG_KWATCH) +=3D kwatch.o =20 kwatch-y :=3D core.o deref.o task_ctx.o hwbp.o probe.o anchor.o +kwatch-$(CONFIG_KWATCH_KUNIT_TEST) +=3D deref_test.o diff --git a/mm/kwatch/deref_test.c b/mm/kwatch/deref_test.c new file mode 100644 index 000000000000..35919dd24d92 --- /dev/null +++ b/mm/kwatch/deref_test.c @@ -0,0 +1,146 @@ +// SPDX-License-Identifier: GPL-2.0 +#include +#include "kwatch.h" +#include + +static void kwatch_test_parse_deref_chain(struct kunit *test) +{ + struct kwatch_config cfg; + int ret; + + // Test 1: stack + memset(&cfg, 0, sizeof(cfg)); + ret =3D kwatch_deref_parse(&cfg, "stack"); + KUNIT_EXPECT_EQ(test, ret, 0); + KUNIT_EXPECT_EQ(test, cfg.base, KWATCH_BASE_STACK); + KUNIT_EXPECT_EQ(test, cfg.offset_count, 1); + KUNIT_EXPECT_EQ(test, cfg.offsets[0], 0); + + // Test 2: arg1 + memset(&cfg, 0, sizeof(cfg)); + ret =3D kwatch_deref_parse(&cfg, "arg1"); + KUNIT_EXPECT_EQ(test, ret, 0); + KUNIT_EXPECT_EQ(test, cfg.base, KWATCH_BASE_ARG1); + KUNIT_EXPECT_EQ(test, cfg.offset_count, 1); + KUNIT_EXPECT_EQ(test, cfg.offsets[0], 0); + + // Test 3: arg6+8 + memset(&cfg, 0, sizeof(cfg)); + ret =3D kwatch_deref_parse(&cfg, "arg6+8"); + KUNIT_EXPECT_EQ(test, ret, 0); + KUNIT_EXPECT_EQ(test, cfg.base, KWATCH_BASE_ARG6); + KUNIT_EXPECT_EQ(test, cfg.offset_count, 1); + KUNIT_EXPECT_EQ(test, cfg.offsets[0], 8); + + // Test 4: arg2-16 + memset(&cfg, 0, sizeof(cfg)); + ret =3D kwatch_deref_parse(&cfg, "arg2-16"); + KUNIT_EXPECT_EQ(test, ret, 0); + KUNIT_EXPECT_EQ(test, cfg.base, KWATCH_BASE_ARG2); + KUNIT_EXPECT_EQ(test, cfg.offset_count, 1); + KUNIT_EXPECT_EQ(test, cfg.offsets[0], -16); + + // Test 5: arg3->8 + memset(&cfg, 0, sizeof(cfg)); + ret =3D kwatch_deref_parse(&cfg, "arg3->8"); + KUNIT_EXPECT_EQ(test, ret, 0); + KUNIT_EXPECT_EQ(test, cfg.base, KWATCH_BASE_ARG3); + KUNIT_EXPECT_EQ(test, cfg.offset_count, 2); + KUNIT_EXPECT_EQ(test, cfg.offsets[0], 0); + KUNIT_EXPECT_EQ(test, cfg.offsets[1], 8); + + // Test 6: arg4+8->16 + memset(&cfg, 0, sizeof(cfg)); + ret =3D kwatch_deref_parse(&cfg, "arg4+8->16"); + KUNIT_EXPECT_EQ(test, ret, 0); + KUNIT_EXPECT_EQ(test, cfg.base, KWATCH_BASE_ARG4); + KUNIT_EXPECT_EQ(test, cfg.offset_count, 2); + KUNIT_EXPECT_EQ(test, cfg.offsets[0], 8); + KUNIT_EXPECT_EQ(test, cfg.offsets[1], 16); + + // Test 7: arg5-8->-16 + memset(&cfg, 0, sizeof(cfg)); + ret =3D kwatch_deref_parse(&cfg, "arg5-8->-16"); + KUNIT_EXPECT_EQ(test, ret, 0); + KUNIT_EXPECT_EQ(test, cfg.base, KWATCH_BASE_ARG5); + KUNIT_EXPECT_EQ(test, cfg.offset_count, 2); + KUNIT_EXPECT_EQ(test, cfg.offsets[0], -8); + KUNIT_EXPECT_EQ(test, cfg.offsets[1], -16); + + // Test 8: stack->0->8 + memset(&cfg, 0, sizeof(cfg)); + ret =3D kwatch_deref_parse(&cfg, "stack->0->8"); + KUNIT_EXPECT_EQ(test, ret, 0); + KUNIT_EXPECT_EQ(test, cfg.base, KWATCH_BASE_STACK); + KUNIT_EXPECT_EQ(test, cfg.offset_count, 3); + KUNIT_EXPECT_EQ(test, cfg.offsets[0], 0); + KUNIT_EXPECT_EQ(test, cfg.offsets[1], 0); + KUNIT_EXPECT_EQ(test, cfg.offsets[2], 8); + + // Test 9: arg1->+8 + memset(&cfg, 0, sizeof(cfg)); + ret =3D kwatch_deref_parse(&cfg, "arg1->+8"); + KUNIT_EXPECT_EQ(test, ret, 0); + KUNIT_EXPECT_EQ(test, cfg.base, KWATCH_BASE_ARG1); + KUNIT_EXPECT_EQ(test, cfg.offset_count, 2); + KUNIT_EXPECT_EQ(test, cfg.offsets[0], 0); + KUNIT_EXPECT_EQ(test, cfg.offsets[1], 8); + + // Test 9.1: arg1-> (implicit 0 should fail) + memset(&cfg, 0, sizeof(cfg)); + ret =3D kwatch_deref_parse(&cfg, "arg1->"); + KUNIT_EXPECT_EQ(test, ret, -EINVAL); + + // Test 9.2: stack->->8 (implicit 0 should fail) + memset(&cfg, 0, sizeof(cfg)); + ret =3D kwatch_deref_parse(&cfg, "stack->->8"); + KUNIT_EXPECT_EQ(test, ret, -EINVAL); + + // Test 10: Invalid base + memset(&cfg, 0, sizeof(cfg)); + ret =3D kwatch_deref_parse(&cfg, "invalid_base"); + KUNIT_EXPECT_EQ(test, ret, -EINVAL); + + // Test 11: Invalid offset + memset(&cfg, 0, sizeof(cfg)); + ret =3D kwatch_deref_parse(&cfg, "arg1+abc"); + KUNIT_EXPECT_EQ(test, ret, -EINVAL); + + // Test 12: Invalid arg + memset(&cfg, 0, sizeof(cfg)); + ret =3D kwatch_deref_parse(&cfg, "arg7"); + KUNIT_EXPECT_EQ(test, ret, -EINVAL); + + // Test 13: Absolute address. Use a width-appropriate literal: a 64-bit + // address would overflow unsigned long and fail kstrtoul() on 32-bit. + memset(&cfg, 0, sizeof(cfg)); +#if BITS_PER_LONG =3D=3D 64 + ret =3D kwatch_deref_parse(&cfg, "0xffffffff81000000+8"); +#else + ret =3D kwatch_deref_parse(&cfg, "0xc1000000+8"); +#endif + KUNIT_EXPECT_EQ(test, ret, 0); + KUNIT_EXPECT_EQ(test, cfg.base, KWATCH_BASE_ABS_ADDR); +#if BITS_PER_LONG =3D=3D 64 + KUNIT_EXPECT_EQ(test, cfg.sym_addr, 0xffffffff81000000UL); +#else + KUNIT_EXPECT_EQ(test, cfg.sym_addr, 0xc1000000UL); +#endif + KUNIT_EXPECT_EQ(test, cfg.offset_count, 1); + KUNIT_EXPECT_EQ(test, cfg.offsets[0], 8); +} + +static struct kunit_case kwatch_deref_test_cases[] =3D { + KUNIT_CASE(kwatch_test_parse_deref_chain), + {} +}; + +static struct kunit_suite kwatch_deref_test_suite =3D { + .name =3D "kwatch_deref", + .test_cases =3D kwatch_deref_test_cases, +}; + +kunit_test_suite(kwatch_deref_test_suite); + +MODULE_DESCRIPTION("KUnit tests for the KWatch watch expression parser"); +MODULE_LICENSE("GPL"); --=20 2.53.0 From nobody Sat Jul 25 05:23:13 2026 Received: from mail-pj1-f44.google.com (mail-pj1-f44.google.com [209.85.216.44]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 69CB13FD15F for ; Fri, 17 Jul 2026 13:07:44 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.44 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293669; cv=none; b=Na7NcVxVpIF1w4KRHm3hA5Cb5tbaap9LTvSOf0ld/vkQ67EMZHhWCUNNlsb2M4qXVTJ4zLPlLndOxqUTrzvueqR1WnO+wNTc+E7rhz5SDNEkE8JDC8XtYNFjFe5jyKH7B5ZgV0iTePFq0QWauCPxReIs9QdArYsTAJmCx9Nc1HE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293669; c=relaxed/simple; bh=PBhnT9R86+Lgrugmh2xVUA0skBZMR4Bzkwn46FP4ujY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=NMeRg0yeNOG6W/Fa6tr8AmBeaIspU9iCna4yz043abAaH5w5HMrWyu8KliNB3TzdqkWwwZKsK7f/Pp1cVIXDd77s4L0n078NXRS6qEcBBZMgYAAEejRr0nN3jLxWQXQ7DC2SMVs+uiSZSmL61GUBompuaU/6W2h33VlQLDhMQiY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=ojrEWIK6; arc=none smtp.client-ip=209.85.216.44 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="ojrEWIK6" Received: by mail-pj1-f44.google.com with SMTP id 98e67ed59e1d1-38de840f2f0so4418507a91.0 for ; Fri, 17 Jul 2026 06:07:44 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784293664; x=1784898464; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=YAVyRkTkPaWk8kIYvBYAepg/O4FB7G2KaVbrN0SgkUc=; b=ojrEWIK6hQwJGiq4eK8Dloh1pfU5NKHC6nc1v/8eWlLpuQueOENXZIwDFAeTH6tkbb TlCbFCAZLea17qbN9TyF1JSEcj7unHzZIgBIJZlMX4Yp5B/+Z0+Tuc2WxJDPFb6S07Fv /jv2K2G9OR7lRkdBW7bqTsrBjJgaQ5wdmvqdoI2EEMGACOgqdaLmNItumKF7d/gSbDPq Ch/PNjYpqzPGFMiZWlSkO5RwUQAA6XR5Y+hKKbIGjkbfkoJLFzxtmuKXWJWwKMu+3NRQ J+ndyPXHmuivQTGwdaeZZAnDn3O1hY790VWF3IGUqZ9prcaT/s54dSr4Pbz7DGgQCeYr S/Vw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784293664; x=1784898464; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=YAVyRkTkPaWk8kIYvBYAepg/O4FB7G2KaVbrN0SgkUc=; b=Nw0YnEI1u3ln/RIrzwS96aJdZ3hXU3k19jieUFUnpEFTB9lxdSher8HMllgITm7Z5m N3uiG4BpjBzIlNPsfRMFjjXSlg84ii5BlyUgJHCldPdwtP9fw5eHSr/PSHA91pAjn0qV p0qzMDIeiA1uxftngBog7jTCwGD671gb+pU7UURsNZze31luLG3v8pv3XEXsKkVUzOcb kkbNCcYkN9EjYuQsupAQPJZ3fl6rDc30l4irbHc57rr8dP8C9uz+fSs+cTqpzl8Kb80N wLmkm1UQPiWgfxtnuOmDKv7jpe8QH1hBE+ia4+872fkNmGxXnPtZkOpBdVIWMLcZOb3T U2Uw== X-Forwarded-Encrypted: i=1; AHgh+RpqPDoegMyoI7C7DW72aKGav01SQ+bFMhDnhFyISlZQYOAd5vhW0UDF71k6rqUNSLnp2l2NIdp8/CaJWa8=@vger.kernel.org X-Gm-Message-State: AOJu0YxxvWUp6lp0ndsXbeCem4ro/Ctz+Q2gfQYHK9EixA+C/6YxjQSW Q5q5HUmS8dsxx9pBB2speHMxzBjd3aS+7+g1ki56rWueUV9t7BEAWmjh X-Gm-Gg: AfdE7cnCP7lt5WZLzjz8nesRj1iIUIvPVu/F8PsrIcCu3FsA7na8KSL2SQxg0cDXBiU BQE2N/3IF0C8lffeC6Uz7i/gkG7IYmzzarLaHQeUVRR256ussLUDn3iJMfPcxzTsPR6zVfO3+7W YSrIU7l0qGPd/DLvUsCeh3grkBFVW3SWwBquVAVX3YBl8x0YeUduyH8odcmuu339ZDWhbteQAqd bvOuUwmjPI9LS68WmBJpLlDW6pJNAkqWYZF6piwKx4BaBBejurNeLsNQi8cA7oHFD8fi8x/+7Ur kjEffv7ufBQilYCjfBcTmgZhWBnpUatjnSqxho+3VFagWCHN1O3xAXzgqfNiOcRpiWhsKoLE8Vi ub4wZjqIiBNtM/6IwCHudzsMszRW7Z2jX8YjWD3OM1wHyaocVf/y5vthQ4SAJ4POAkSMxJ4x43N QOO2pMCXRhY/00 X-Received: by 2002:a17:90b:3904:b0:387:e0db:bc34 with SMTP id 98e67ed59e1d1-38e4b5793d2mr2622266a91.42.1784293663521; Fri, 17 Jul 2026 06:07:43 -0700 (PDT) Received: from localhost ([144.24.58.22]) by smtp.gmail.com with ESMTPSA id a92af1059eb24-13ce29e340esm5521158c88.5.2026.07.17.06.07.41 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 17 Jul 2026 06:07:42 -0700 (PDT) From: Jinchao Wang To: Andrew Morton , Peter Zijlstra , Thomas Gleixner , Steven Rostedt , Masami Hiramatsu Cc: Ingo Molnar , Borislav Petkov , Dave Hansen , "H . Peter Anvin" , x86@kernel.org, Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Mathieu Desnoyers , David Hildenbrand , Jonathan Corbet , Matthew Wilcox , Alan Stern , Randy Dunlap , Alexander Potapenko , Marco Elver , Mike Rapoport , linux-kernel@vger.kernel.org, linux-mm@kvack.org, linux-trace-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org, linux-doc@vger.kernel.org, Jinchao Wang Subject: [RFC PATCH v2 13/13] Documentation/dev-tools: document KWatch Date: Fri, 17 Jul 2026 09:07:26 -0400 Message-ID: <20260717130726.1904101-1-wangjinchao600@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260717125023.1895892-1-wangjinchao600@gmail.com> References: <20260717125023.1895892-1-wangjinchao600@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Describe what KWatch is for, how it compares with KASAN and KFENCE, the debugfs configuration interface, the watch expression syntax, how to read hits from the trace buffer (including after a crash), and the current limitations. Signed-off-by: Jinchao Wang --- Documentation/dev-tools/index.rst | 1 + Documentation/dev-tools/kwatch.rst | 207 +++++++++++++++++++++++++++++ 2 files changed, 208 insertions(+) create mode 100644 Documentation/dev-tools/kwatch.rst diff --git a/Documentation/dev-tools/index.rst b/Documentation/dev-tools/in= dex.rst index 59cbb77b33ff..f4c748da63db 100644 --- a/Documentation/dev-tools/index.rst +++ b/Documentation/dev-tools/index.rst @@ -30,6 +30,7 @@ Documentation/process/debugging/index.rst ubsan kmemleak kcsan + kwatch lkmm/index kfence kselftest diff --git a/Documentation/dev-tools/kwatch.rst b/Documentation/dev-tools/k= watch.rst new file mode 100644 index 000000000000..e58f3185ebbd --- /dev/null +++ b/Documentation/dev-tools/kwatch.rst @@ -0,0 +1,207 @@ +.. SPDX-License-Identifier: GPL-2.0 + +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D +KWatch - Kernel Memory Watchpoint Tool +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +Overview +=3D=3D=3D=3D=3D=3D=3D=3D + +KWatch is a runtime-configurable debugging tool for locating kernel memory +corruption. It arms hardware breakpoints (watchpoints) on a target address +while a chosen function is executing, and reports the exact instruction th= at +touches the watched memory, together with a stack trace, through a +tracepoint. + +Unlike shadow-memory sanitizers, KWatch does not detect invalid accesses in +general; it answers a narrower but common question during corruption hunts: +"who writes to this address?". This includes in-bounds logical overwrites +that KASAN cannot see, because the rogue writer modifies valid memory +through a valid pointer, just at the wrong time or with the wrong data. + +Comparison with other tools: + +* KASAN detects out-of-bounds and use-after-free accesses, but reports the + symptom (the invalid access), not the writer that corrupted the data + earlier. It requires a rebuild and has significant CPU and memory + overhead, and its redzones perturb memory layout, which can hide + timing-sensitive bugs. +* KFENCE is a low-overhead sampling detector for slab objects; it cannot be + pointed at one specific address. +* Hardware breakpoints via kgdb or perf can watch an address, but only a + fixed one, system-wide, for the whole run. KWatch resolves the address + dynamically at function entry (for example "argument 2 of this function, + plus offset 8, dereferenced once") and disarms it again at function exit, + so short-lived and per-invocation objects can be watched too. + +KWatch has near-zero overhead while armed: the watched function pays for +one kprobe/kretprobe pair plus programming of the debug registers; the rest +of the system runs at full speed. + +Requirements +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +* ``CONFIG_KWATCH=3Dy`` or ``m``. The Kconfig symbol depends on + ``CONFIG_PERF_EVENTS``, ``CONFIG_DEBUG_FS`` and an architecture that + provides ``HAVE_REINSTALL_HW_BREAKPOINT`` (currently x86 only). +* Resolving symbol names in watch expressions requires ``CONFIG_KWATCH=3Dy= `` + (built-in); a module can only watch absolute hexadecimal addresses. + +Usage +=3D=3D=3D=3D=3D + +KWatch is configured through a single debugfs file:: + + /sys/kernel/debug/kwatch/config + +Writing a configuration string starts a watch session (stopping any previo= us +one); reading the file shows the active configuration and hit-rejection +counters. The configuration is a whitespace-separated list of ``key=3Dvalu= e`` +tokens: + +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D +Key Meaning +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D +``func_name`` Function whose execution opens the watch window. +``func_offset`` Instruction offset inside ``func_name`` at which the + watchpoint is armed (default 0 =3D function entry). +``watch_expr`` Expression describing the address to watch (see below). +``watch_len`` Watched length in bytes: 1, 2, 4 or 8 (default 8). +``depth`` Recursion depth at which the window opens (default 0). +``max_watch`` Number of hardware watchpoints to preallocate + (default 4). +``max_concurrency`` Maximum number of tasks concurrently inside the watch + window (default 256). +``duration`` For global watches: seconds until automatic stop. +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D + +Watch expressions +----------------- + +The address to watch is computed at function entry from:: + + watch_expr=3D{base}[+-offset][->[+-]offset]... + +* ``base`` is one of: + + - ``arg1`` ... ``arg6``: a function argument (register calling + convention). These are only meaningful at function entry, i.e. with + ``func_offset`` unset; combining ``argN`` with ``func_offset`` reads + the argument registers mid-function, where they no longer hold the + original arguments, + - ``stack``: the kernel stack pointer at the probe point, + - an absolute hexadecimal address, e.g. ``0xffffffff81234567``, + - a global symbol name (built-in KWatch only). + +* ``+offset`` / ``-offset`` adjusts the current address. +* ``->offset`` loads the pointer stored at the current address (via + ``get_kernel_nofault()``) and then applies the offset. Up to four chain + elements are supported; offsets must be explicit (``->`` alone is + rejected). + +Given:: + + struct some_struct { + struct some_struct *ptr; /* offset 0 */ + int num; /* offset 8 */ + }; + + void target_function(struct some_struct *arg1); + +typical expressions are: + +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D +Expression Watches +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D +``watch_expr=3Darg1`` ``&arg1->ptr`` (the pointer field itself) +``watch_expr=3Darg1+8`` ``&arg1->num`` +``watch_expr=3Darg1->0`` ``&arg1->ptr->ptr`` (one dereference) +``watch_expr=3Darg1->8`` ``&arg1->ptr->num`` +``watch_expr=3D0xffff...+8`` absolute address plus 8 +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +Example: catch whoever overwrites ``arg1->num`` of a function while that +function runs:: + + echo "func_name=3Dtarget_function watch_expr=3Darg1+8 watch_len=3D4" \ + > /sys/kernel/debug/kwatch/config + +Watching global variables +------------------------- + +A global variable has no natural function window. When ``duration`` is +given without ``func_name``, KWatch starts an internal anchor kernel thread +that sleeps inside a dummy function, and uses that function as the window:: + + echo "watch_expr=3Djiffies_wobble duration=3D60 watch_len=3D8" \ + > /sys/kernel/debug/kwatch/config + +The session tears itself down when the duration expires. + +Reading hits +------------ + +Hits are emitted as the ``kwatch:kwatch_hit`` tracepoint, which is safe in +NMI-like contexts where printk is not. Each event carries the timestamp, +the instruction pointer, the watched address and a short stack trace:: + + echo 1 > /sys/kernel/debug/tracing/events/kwatch/kwatch_hit/enable + cat /sys/kernel/debug/tracing/trace_pipe + +If the corruption crashes the machine, the ring buffer can still be +recovered: + +* ``echo 1 > /proc/sys/kernel/ftrace_dump_on_oops`` (or the + ``ftrace_dump_on_oops`` boot parameter) dumps the buffer to the console + on an oops. +* With kdump, the buffer is present in the vmcore and can be read with + ``crash> trace``. +* ``CONFIG_PSTORE_FTRACE`` persists it across reboots on supported + platforms. + +Limitations +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +* Functions that run in a genuine NMI(-like) context are rejected at + function entry; rejected invocations never open a watch window and are + counted in the ``nmi_rejected`` field of the config file. Watching + functions reachable from NMI handlers is out of scope. +* The number of concurrent watchpoints is bounded by the CPU's debug + registers (typically 4). +* Cross-CPU re-arming of a watchpoint is rate-limited per CPU: the local + CPU is always re-pointed, but the broadcast to other CPUs is throttled + so a very hot watched function cannot storm the system with IPIs. While + a broadcast is suppressed the other CPUs keep watching the previous + address, so a writer that runs on another CPU during that window can be + missed; the number of suppressed broadcasts is reported in the + ``arm_ipi_suppressed`` field of the config file. KWatch therefore + targets functions that are entered at a moderate rate. +* If the target address cannot be resolved at arming time (for example a + ``get_kernel_nofault()`` failure on a swapped or unmapped page), the + watchpoint is not armed for that invocation. +* Offsets in watch expressions are static; dynamic indexing such as + ``arg1->ptr[arg2]`` is not supported. +* If a task is torn down while still inside the watched function without + the function returning (an oops or BUG in the window, which abandons the + stack), its watch window is not closed until the session stops. This is + the same best-effort cleanup that applies to every resource a task holds + when it dies abnormally. Do not target the task-exit path itself. +* arm64 is not yet supported: stepping over a hit that has a custom + overflow handler needs a generic mechanism in the arch code, which is + planned as a follow-up series. + +Implementation notes +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +The implementation lives in ``mm/kwatch/`` and is split into a control +plane (``core.c``, the debugfs interface), an execution plane (``probe.c`` +and ``deref.c``: kprobe/kretprobe window management and address +resolution), and a resource plane (``hwbp.c`` and ``task_ctx.c``). + +Hardware watchpoints are preallocated as perf events on every CPU and +re-pointed at hit time with ``modify_wide_hw_breakpoint_local()``, a new +hw_breakpoint API that updates the breakpoint on the local CPU without +releasing its slot; other CPUs are updated by asynchronous IPIs. Per-task +window state is kept in a fixed-size, lockless open-addressing array +claimed with ``cmpxchg()``, so the hit path performs no allocation and +takes no locks, which keeps it safe in atomic and NMI-like contexts. --=20 2.53.0