From nobody Fri Sep 25 21:02:40 2026 Received: from mail-pj2-f13.google.com (mail-pj2-f13.google.com [74.125.227.141]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D39D7488DAF for ; Mon, 21 Sep 2026 11:16:03 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.227.141 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789989365; cv=none; b=b+/cYdV6MuTI6V8DwXoqF4cgaQKwXRoUCJNUwd+LNCw86XSZ0+1/AboS0qh5Rwkh50Q3TJ0C2GAFVxMtjYX3t72YGQFnVgKeErIJYBNzT5mabIZWjH9RnSf398ww3diKyaBsxziUAOyE3dnhPFONo5BnmLyuwlPZQXjNdggsNF8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789989365; c=relaxed/simple; bh=Uupe71rKqVRL9KIaXB//FIQMoxw3iyok3V2rnxr/HZo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=QhWDBkv72EwnA73fRy+eZGRbxVHuHIAtFxzYcdPU0wD41N+KhuLEXsq+xEbRzOnFDBfyK12ak4JPgoZysPXljumbb0tiHqDZgulV43g35dRe7WB3pBKz7GC8wPV7qxnDDSlO6+dI5Nt5Nj/FWuK3G5Inz8QdIA3mtGWisyLc+FI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com; spf=pass smtp.mailfrom=bytedance.com; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b=R7K8Qgte; arc=none smtp.client-ip=74.125.227.141 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=bytedance.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b="R7K8Qgte" Received: by mail-pj2-f13.google.com with SMTP id d9443c01a7336-2d747ed9866so25981925ad.2 for ; Mon, 21 Sep 2026 04:16:03 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=bytedance.com; s=google; t=1789989363; x=1790594163; darn=vger.kernel.org; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:from:to:cc:subject :date:message-id:reply-to:content-type; bh=LkdiBF2IWGfWcg5OnqET18ZOvzF9PCNRUzrg/Kdn0K8=; b=R7K8QgteLtXX2wftJxqtI1HhieF1hpyuyZb0XEFirQB/3EMTFf5kirQlikthxR0mi0 zlI7KVTbUlX2NcJHn3X/c8LSc6csUzp/3sTKUrKkKKPvg9Dz4WeGthkyqyM8E0WqSL8y d3oIZH2+/4VRxyizVArPAaW3gCq6uiu94qqzEAPyTmqGApUX1Ow5BhC42qc5DN2HDlkH r1hq/A+Vtm6gO8cvpyBx49XvvuGtuXWbMbisodElZdWybchUrCIiL/KCN4Lgo8RwxuAK WxZuLBpHxDilN1HHLuOxKu0FOYqrFa8AX2/oKBaHJhI90v1b6Va1KHtLU5fPeAYLOfQf jU1g== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789989363; x=1790594163; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=LkdiBF2IWGfWcg5OnqET18ZOvzF9PCNRUzrg/Kdn0K8=; b=RJcDsiC7IrBLR0uidAZ3lQX8Uq3iZsXmoSBjFgBpwVpbL65Wpitdd1C17P6UVwJeYb YJGxrRhNy5sdRibUCxQA0UHZaeekNann/jKqyYhl0JJGx48/ao/b2hh1glpySRiJj0mH Lffe4Qwq3rQr6NkWy9ToU1WednQBHII/fVOdG6Z08IQgaQBOF2FVcl/XV72uhqm3uSZ+ +l6b08Cm11JJqcPxzfz0MTWKGHBggeemHeRNbmAgvZ8n25oeSCF9W0Q7MMq2tthi5iKU aFmRx/KEVA1Ycl0Q7/r4C0wti3H/Eyf8Su+6mF+6UqJjcolvPiqJ06c1DBDUQqe+20hk h/3g== X-Forwarded-Encrypted: i=1; AKwUvBzr/hRdSw5LvTNWPIXYbOZF2RU+RFkLaMDFzWM2WR1q6gnDh9gRe8l/evjH0si4o+GaJ5JmmbjoEZSfAss=@vger.kernel.org X-Gm-Message-State: AFuF++lSLCKVYK4hPJCIc9KyIjRb6cAtOcErKyJP1Y2iHvQGGhskbWKD KD8sg4IkaiI1d84SZo+UJCTLTzZJo+PfuO3EhBZVuHZBCymGHjNBOnezz23Lon52vSw= X-Gm-Gg: AYBFou3T/Sh2n5t4Z4VPpma2pEWlzy21MKxBD3xgR6v3539cQTldJkIp1MhWV/FwJwI mOKPnVyLaK5s9nv92ix8OMNA5SOreXhfX2AoybLqw/JnZwKBJaD+4uj+YFKDlGDhY2sVroBMYIw oY1Sx3IBhrP7gyYbaNhR6x9xcUzaevPILaTu4WTLK4QhyKqAt97yPa545eS+XJXzLVVq3tW10cE sjkzlsAGySHDBOhThYV9n5fZSrCGK6JP4Ozn1Du01ipdtmXpnTg9Q3kC4s2YdihkYnRmLJxlDNw zh7ZFBPU711XG/yK/Z0N+Dv5mmzqvRkdi4PgLYzICXg3JY1ate9vM+XrzhXNr5o7c9Y/14xKLv0 hW/RjeJbpbCm+SoWkAH69sHKc1t4jH+jI4UhNY5v4p+013c+4dDx2XdoB4JK9PA1Awli676zVz2 7dQwE5S57M1Bv2C2Bd9NRgPljp/lp+kvuVRdEoZAlXrY+6/J9KwAYBvHtnw33Cq2u2IPQjeY12m wDnAX/4c2NO4qrXpQMt2hS1Ys0tk1L/Z14= X-Received: by 2002:a17:903:3c05:b0:2dd:4070:8eb with SMTP id d9443c01a7336-2ddb1b97f0amr160128205ad.20.1789989362987; Mon, 21 Sep 2026 04:16:02 -0700 (PDT) Received: from FJ7FR2JRQ3.bytedance.net ([240e:694:e20:401::8]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2ddc17ba651sm31831125ad.53.2026.09.21.04.15.41 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Mon, 21 Sep 2026 04:16:02 -0700 (PDT) From: Zhanpeng Zhang To: Paul Walmsley , Palmer Dabbelt , Albert Ou , Alexandre Ghiti , Himanshu Chauhan , Conor Dooley , Anup Patel Cc: =?UTF-8?q?Cl=C3=A9ment=20L=C3=A9ger?= , Yunhui Cui , Atish Patra , Peter Zijlstra , Ingo Molnar , Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Alexander Shishkin , Jiri Olsa , Ian Rogers , Adrian Hunter , James Clark , Will Deacon , Thomas Gleixner , Jonathan Corbet , Randy Dunlap , Shuah Khan , Shuah Khan , Yuanzhu , Yicong Yang , Susheng Yang , linux-riscv@lists.infradead.org, linux-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-arm-kernel@lists.infradead.org, Zhanpeng Zhang Subject: [PATCH v10 RESEND 1/9] riscv: add SBI SSE extension definitions Date: Mon, 21 Sep 2026 19:14:58 +0800 Message-ID: <990ca6db19bc476f7533f97d48e2030b509633aa.1789974241.git.zhangzhanpeng.jasper@bytedance.com> X-Mailer: git-send-email 2.50.1 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable From: Cl=C3=A9ment L=C3=A9ger Add definitions for the SBI Supervisor Software Events extension [1]. This extension enables the SBI to inject events into supervisor software much like ARM SDEI. [1] https://lists.riscv.org/g/tech-prs/message/515 Signed-off-by: Cl=C3=A9ment L=C3=A9ger Co-developed-by: Himanshu Chauhan Signed-off-by: Himanshu Chauhan Co-developed-by: Zhanpeng Zhang Signed-off-by: Zhanpeng Zhang --- arch/riscv/include/asm/sbi.h | 63 ++++++++++++++++++++++++++++++++++++ 1 file changed, 63 insertions(+) diff --git a/arch/riscv/include/asm/sbi.h b/arch/riscv/include/asm/sbi.h index 5725e0ca4dda..4e54d79ba543 100644 --- a/arch/riscv/include/asm/sbi.h +++ b/arch/riscv/include/asm/sbi.h @@ -38,6 +38,7 @@ enum sbi_ext_id { SBI_EXT_FWFT =3D 0x46574654, SBI_EXT_MPXY =3D 0x4D505859, SBI_EXT_DBTR =3D 0x44425452, + SBI_EXT_SSE =3D 0x535345, =20 /* Experimentals extensions must lie within this range */ SBI_EXT_EXPERIMENTAL_START =3D 0x08000000, @@ -506,6 +507,68 @@ enum sbi_mpxy_rpmi_attribute_id { #define SBI_MPXY_CHAN_CAP_SEND_WITHOUT_RESP BIT(4) #define SBI_MPXY_CHAN_CAP_GET_NOTIFICATIONS BIT(5) =20 +enum sbi_ext_sse_fid { + SBI_SSE_EVENT_ATTR_READ =3D 0, + SBI_SSE_EVENT_ATTR_WRITE, + SBI_SSE_EVENT_REGISTER, + SBI_SSE_EVENT_UNREGISTER, + SBI_SSE_EVENT_ENABLE, + SBI_SSE_EVENT_DISABLE, + SBI_SSE_EVENT_COMPLETE, + SBI_SSE_EVENT_INJECT, + SBI_SSE_HART_UNMASK, + SBI_SSE_HART_MASK, +}; + +enum sbi_sse_state { + SBI_SSE_STATE_UNUSED =3D 0, + SBI_SSE_STATE_REGISTERED =3D 1, + SBI_SSE_STATE_ENABLED =3D 2, + SBI_SSE_STATE_RUNNING =3D 3, +}; + +/* SBI SSE Event Attributes. */ +enum sbi_sse_attr_id { + SBI_SSE_ATTR_STATUS =3D 0x00000000, + SBI_SSE_ATTR_PRIO =3D 0x00000001, + SBI_SSE_ATTR_CONFIG =3D 0x00000002, + SBI_SSE_ATTR_PREFERRED_HART =3D 0x00000003, + SBI_SSE_ATTR_ENTRY_PC =3D 0x00000004, + SBI_SSE_ATTR_ENTRY_ARG =3D 0x00000005, + SBI_SSE_ATTR_INTERRUPTED_SEPC =3D 0x00000006, + SBI_SSE_ATTR_INTERRUPTED_FLAGS =3D 0x00000007, + SBI_SSE_ATTR_INTERRUPTED_A6 =3D 0x00000008, + SBI_SSE_ATTR_INTERRUPTED_A7 =3D 0x00000009, + + SBI_SSE_ATTR_MAX =3D 0x0000000A +}; + +#define SBI_SSE_ATTR_STATUS_STATE_OFFSET 0 +#define SBI_SSE_ATTR_STATUS_STATE_MASK 0x3 +#define SBI_SSE_ATTR_STATUS_PENDING_OFFSET 2 +#define SBI_SSE_ATTR_STATUS_INJECT_OFFSET 3 + +#define SBI_SSE_ATTR_CONFIG_ONESHOT BIT(0) + +#define SBI_SSE_ATTR_INTERRUPTED_FLAGS_SSTATUS_SPP BIT(0) +#define SBI_SSE_ATTR_INTERRUPTED_FLAGS_SSTATUS_SPIE BIT(1) +#define SBI_SSE_ATTR_INTERRUPTED_FLAGS_HSTATUS_SPV BIT(2) +#define SBI_SSE_ATTR_INTERRUPTED_FLAGS_HSTATUS_SPVP BIT(3) +#define SBI_SSE_ATTR_INTERRUPTED_FLAGS_SSTATUS_SPELP BIT(4) +#define SBI_SSE_ATTR_INTERRUPTED_FLAGS_SSTATUS_SDT BIT(5) + +#define SBI_SSE_EVENT_LOCAL_HIGH_PRIO_RAS 0x00000000 +#define SBI_SSE_EVENT_LOCAL_DOUBLE_TRAP 0x00000001 +#define SBI_SSE_EVENT_GLOBAL_HIGH_PRIO_RAS 0x00008000 +#define SBI_SSE_EVENT_LOCAL_PMU_OVERFLOW 0x00010000 +#define SBI_SSE_EVENT_LOCAL_LOW_PRIO_RAS 0x00100000 +#define SBI_SSE_EVENT_GLOBAL_LOW_PRIO_RAS 0x00108000 +#define SBI_SSE_EVENT_LOCAL_SOFTWARE_INJECTED 0xffff0000 +#define SBI_SSE_EVENT_GLOBAL_SOFTWARE_INJECTED 0xffff8000 + +#define SBI_SSE_EVENT_PLATFORM BIT(14) +#define SBI_SSE_EVENT_GLOBAL BIT(15) + /* SBI debug triggers function IDs */ enum sbi_ext_dbtr_fid { SBI_EXT_DBTR_NUM_TRIGGERS =3D 0, --=20 2.50.1 (Apple Git-155) From nobody Fri Sep 25 21:02:40 2026 Received: from mail-pj2-f12.google.com (mail-pj2-f12.google.com [74.125.227.140]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1D78E489FD8 for ; Mon, 21 Sep 2026 11:16:26 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.227.140 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789989390; cv=none; b=QNlt7VDhVy/+/sfy1kPpEQ6SPA/KTBYw9b9F1TcX9chlKPDsdsYGrXX3ORjVA2KMBnoIyZTB/6YrYq4J1I7Kwxr3+TtgfqypBRRjWG3vohHm4nyt47QE9Fc15PQSb77BQNLYINs2vdxJXX5cqYMzgC72K9GGsZyZm7pP3L/CFKk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789989390; c=relaxed/simple; bh=G6CSjY6qjOYN2nw4/gPkOoeK/3pp8ZCwCWvOb3IKMKk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=kk+bDKfLrfSpxEagTckhI25ONZqebOYB9Y2HAl6db+EU83j2gBVoUAYyWCWMUgS7eqXNjGejVam/b26vUT0HgKq3nOLo44KHjozGMr6SX3Czdf9BMYJ4Y0SL7cUAT+XFbDDutair+djD1GhnfnXnc3cRfqIflUN73Wign0Lta24= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com; spf=pass smtp.mailfrom=bytedance.com; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b=d8HTh4g+; arc=none smtp.client-ip=74.125.227.140 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=bytedance.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b="d8HTh4g+" Received: by mail-pj2-f12.google.com with SMTP id d9443c01a7336-2dd58e1e2c7so24976055ad.0 for ; Mon, 21 Sep 2026 04:16:26 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=bytedance.com; s=google; t=1789989386; x=1790594186; darn=vger.kernel.org; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:from:to:cc:subject :date:message-id:reply-to:content-type; bh=xeMiGJ2SVpHxKgYYhwuFjvWQcTsb5P9v/SPH3HjgQHc=; b=d8HTh4g+ySkJ0H0l8ags+0GJB1q1cbrFexcXiALcsMupV92nVvUwYL1a5JKpgVe3Q7 wlqmK0sKp3Ugbj4VeOuCUH1VqJgwjZtEaf7lGC04bqVHFwUG86NcyRNalR+g2kmIy7Ht PLNNrjJDVGf4GZYAZZc2GtYgkwOQ1waJpTcpWEUBhsDV8M/aGl8HoiJUKmmTS+kowkwf 8IxMt4IdZ40D/NCMXu+RY6Wt2ZQ6cUjZaoMdlUXDlLIIEWThJ9GIMoEvBxxNXONU8KGI haC85p+tVARvheTXH4FpLMv24+1uau6lrPOUH6FG+8gwPDrIFvHhDfrRw7mWM16sPiVr Y2GA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789989386; x=1790594186; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=xeMiGJ2SVpHxKgYYhwuFjvWQcTsb5P9v/SPH3HjgQHc=; b=RmDJlQEukaKCVCHrLw+yD4FMDCbRrGXndvQInZ3Wm6At8DOgAQRCLz2DjwHl+UTHKV DlEOYtp8qGmJFTIw5taqWHvBV40/zpv5JFg1UWDyo7fe0GkIf6kD7OAsshZWxUqXG+sd nU0q70t7DXL7GzUgz2vLaYtuImMbRijpFcdHnxpRRdSFJj9lgYnrHNzm0jz8HU8cAfUD +8DvHU/O5EnRzau1vEwv9523Sjb4sWaebIu9mGE+4Mat3I3IcF3rO7GabQ0BqjuYH/9s hycSF+XPzj0FRSqMoEe6DxJUEAl8DHnrDbu0+b3qwfYi5OSRawaqZ8aMXM5vd3iIr0Zb JG/Q== X-Forwarded-Encrypted: i=1; AKwUvBxhv/QDaQU1pBOK8Mu6paX2zr2PW0chmRzRTkVHw9r4ToSgNfSdmVo93tdnAdF3gCutCPXA15cyEDy+R00=@vger.kernel.org X-Gm-Message-State: AFuF++kKNdQN+q6tbJdcfuuNJ+83iLUmU2/8eqW41BJwKbl/7zz/+qLh hGKzcsDyM0Pcapx4rGYxmu68RGbh0ozPRLETI0suP4SlFlYZXaMtZe5OlpGHC1HKJWs= X-Gm-Gg: AYBFou0vqi0gC0ERoLxhsPLeNyumHPdwwc4hTn+XDw6XiNaRGQs1eN+wwVsMAEAElex BqvHEqgkkEaIAGrOoX9oSjGKbOSSmdbokWRmCUe37WWXkisCRk2lFTRjPuc/QlpoFFGdSeWxz5+ RYD0MuO1G2ALepo5q8FBOyHWZd2FyP0n9YkbUzPE/P95l4amQ2ZlZu5Ou1/4B7+ibViseQ5plRK zAYNVY543h/JKdFFDgaiv7zo0yyWMaodWH3WNW/OsHR6pLhn2dH13Jrmr8hmsN4ufg7O8S3JOmc dStttHoOYBI3D3IaoEV4O5QMtQqv9JI5ZNDzsjTqH7Hl4iYRnw9/ZGWmRLzlFtMK0FGypEI7mdt c4A74b+NsKRLlcHvjBm2FxDXfKD8PFeyhIc+nuahw4Q7nvq+KilbrCPrzIFKfAxmUk1jwslTyn+ SFc0XJ2KbqgQ2RP/Q6VMg80HgG/x3vo7zedeK0utcVLhzCd1RdXG/+MYAmxOpLzd9Udmk8Y37DK 9E8wpA+38tWdSBLuy9D4aCU0/GMUHGNSB0C X-Received: by 2002:a17:903:1103:b0:2dd:ad74:6d16 with SMTP id d9443c01a7336-2ddb1baf462mr156833115ad.28.1789989386189; Mon, 21 Sep 2026 04:16:26 -0700 (PDT) Received: from FJ7FR2JRQ3.bytedance.net ([240e:694:e20:401::8]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2ddc17ba651sm31831125ad.53.2026.09.21.04.16.03 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Mon, 21 Sep 2026 04:16:25 -0700 (PDT) From: Zhanpeng Zhang To: Paul Walmsley , Palmer Dabbelt , Albert Ou , Alexandre Ghiti , Himanshu Chauhan , Conor Dooley , Anup Patel Cc: =?UTF-8?q?Cl=C3=A9ment=20L=C3=A9ger?= , Yunhui Cui , Atish Patra , Peter Zijlstra , Ingo Molnar , Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Alexander Shishkin , Jiri Olsa , Ian Rogers , Adrian Hunter , James Clark , Will Deacon , Thomas Gleixner , Jonathan Corbet , Randy Dunlap , Shuah Khan , Shuah Khan , Yuanzhu , Yicong Yang , Susheng Yang , linux-riscv@lists.infradead.org, linux-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-arm-kernel@lists.infradead.org, Zhanpeng Zhang Subject: [PATCH v10 RESEND 2/9] riscv: add support for SBI Supervisor Software Events extension Date: Mon, 21 Sep 2026 19:14:59 +0800 Message-ID: X-Mailer: git-send-email 2.50.1 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable From: Cl=C3=A9ment L=C3=A9ger The SBI SSE extension allows firmware to notify supervisor software of events that must be delivered independently of normal S-mode interrupts. Firmware saves the minimal state required to enter the supervisor handler, and Linux builds the synthetic handler context around it. SSE can arrive while Linux is already in an exception entry path. At that point sscratch and tp may be in the middle of the normal trap-entry exchange, so they cannot always identify current. Store current in a per-CPU slot and use the hart ID passed by firmware to recover it. Give each event, including each CPU instance of a local event, a dedicated stack and shadow call stack. Synchronize vmapped stack ranges before unmasking events so that the handler cannot take a vmalloc fault while running in an NMI-like context. The handler is a synthetic supervisor episode, but completion must resume the context interrupted by the SSE. Preserve stvec and, when the hypervisor extension is present, hstatus across the handler. Read the interrupted a6 and a7 values from the SSE attributes, construct pt_regs for the interrupted context, and write back any changes made by the handler. Nested exceptions on the SSE event stack temporarily replace both TASK_TI_KERNEL_SP and TASK_TI_USER_SP. Preserve their original values across the handler and restore them before completing the event, so a nested exception cannot leave the interrupted task referring to the event stack. Keep an explicit EVENT_REGISTER not-supported result distinct from other firmware failures. This lets clients select another delivery mechanism only when firmware has positively rejected the requested event. Signed-off-by: Cl=C3=A9ment L=C3=A9ger Co-developed-by: Himanshu Chauhan Signed-off-by: Himanshu Chauhan Co-developed-by: Zhanpeng Zhang Signed-off-by: Zhanpeng Zhang --- MAINTAINERS | 12 ++ arch/riscv/include/asm/asm.h | 14 +- arch/riscv/include/asm/scs.h | 7 + arch/riscv/include/asm/sse.h | 82 +++++++++ arch/riscv/include/asm/thread_info.h | 1 + arch/riscv/kernel/Makefile | 1 + arch/riscv/kernel/asm-offsets.c | 14 ++ arch/riscv/kernel/entry.S | 14 ++ arch/riscv/kernel/sbi_sse.c | 246 +++++++++++++++++++++++++++ arch/riscv/kernel/sbi_sse_entry.S | 226 ++++++++++++++++++++++++ 10 files changed, 614 insertions(+), 3 deletions(-) create mode 100644 arch/riscv/include/asm/sse.h create mode 100644 arch/riscv/kernel/sbi_sse.c create mode 100644 arch/riscv/kernel/sbi_sse_entry.S diff --git a/MAINTAINERS b/MAINTAINERS index cc3cae2e378b..7a6af19f7e30 100644 --- a/MAINTAINERS +++ b/MAINTAINERS @@ -23595,6 +23595,18 @@ F: arch/riscv/boot/dts/spacemit/ N: spacemit K: spacemit =20 +RISC-V SUPERVISOR SOFTWARE EVENTS +M: Zhanpeng Zhang +M: Himanshu Chauhan +R: Yunhui Cui +L: linux-riscv@lists.infradead.org +S: Maintained +F: arch/riscv/include/asm/sse.h +F: arch/riscv/kernel/sbi_sse.c +F: arch/riscv/kernel/sbi_sse_entry.S +F: drivers/firmware/riscv/riscv_sbi_sse.c +F: include/linux/riscv_sbi_sse.h + RISC-V TENSTORRENT SoC SUPPORT M: Drew Fustini M: Joel Stanley diff --git a/arch/riscv/include/asm/asm.h b/arch/riscv/include/asm/asm.h index b8bf842d4c13..e1196caea02d 100644 --- a/arch/riscv/include/asm/asm.h +++ b/arch/riscv/include/asm/asm.h @@ -91,16 +91,24 @@ .endm =20 #ifdef CONFIG_SMP -.macro asm_per_cpu dst sym tmp - lw \tmp, TASK_TI_CPU_NUM(tp) - slli \tmp, \tmp, RISCV_LGPTR +.macro asm_per_cpu_with_cpu dst sym tmp cpu + slli \tmp, \cpu, RISCV_LGPTR la \dst, __per_cpu_offset add \dst, \dst, \tmp REG_L \tmp, 0(\dst) la \dst, \sym add \dst, \dst, \tmp .endm + +.macro asm_per_cpu dst sym tmp + lw \tmp, TASK_TI_CPU_NUM(tp) + asm_per_cpu_with_cpu \dst \sym \tmp \tmp +.endm #else /* CONFIG_SMP */ +.macro asm_per_cpu_with_cpu dst sym tmp cpu + la \dst, \sym +.endm + .macro asm_per_cpu dst sym tmp la \dst, \sym .endm diff --git a/arch/riscv/include/asm/scs.h b/arch/riscv/include/asm/scs.h index 023a412fe38d..0d70a35bc01a 100644 --- a/arch/riscv/include/asm/scs.h +++ b/arch/riscv/include/asm/scs.h @@ -17,6 +17,11 @@ load_per_cpu gp, irq_shadow_call_stack_ptr, \tmp .endm =20 +/* Load the per-CPU IRQ shadow call stack to gp. */ +.macro scs_load_sse_stack reg_evt + REG_L gp, SSE_REG_EVT_SHADOW_STACK(\reg_evt) +.endm + /* Load task_scs_sp(current) to gp. */ .macro scs_load_current REG_L gp, TASK_TI_SCS_SP(tp) @@ -40,6 +45,8 @@ .endm .macro scs_load_irq_stack tmp .endm +.macro scs_load_sse_stack reg_evt +.endm .macro scs_load_current .endm .macro scs_load_current_if_task_changed prev diff --git a/arch/riscv/include/asm/sse.h b/arch/riscv/include/asm/sse.h new file mode 100644 index 000000000000..cbd8618c6e00 --- /dev/null +++ b/arch/riscv/include/asm/sse.h @@ -0,0 +1,82 @@ +/* SPDX-License-Identifier: GPL-2.0-only */ +/* + * Copyright (C) 2024 Rivos Inc. + */ +#ifndef __ASM_SSE_H +#define __ASM_SSE_H + +#include +#include + +#include + +static inline bool riscv_sse_available(void) +{ +#ifdef CONFIG_RISCV_SBI + return sbi_probe_extension(SBI_EXT_SSE) > 0; +#else + return false; +#endif +} + +static inline void riscv_sse_mask_current_hart(void) +{ +#ifdef CONFIG_RISCV_SBI + struct sbiret ret; + + if (!riscv_sse_available()) + return; + + ret =3D sbi_ecall(SBI_EXT_SSE, SBI_SSE_HART_MASK, 0, 0, 0, 0, 0, 0); + if (ret.error && ret.error !=3D SBI_ERR_ALREADY_STOPPED) + pr_emerg("SSE hart mask failed: %ld\n", ret.error); +#endif +} + +#ifdef CONFIG_RISCV_SBI_SSE + +struct sse_event_interrupted_state { + unsigned long a6; + unsigned long a7; +}; + +struct sse_event_arch_data { + void *stack; + void *shadow_stack; + unsigned long tmp; + struct sse_event_interrupted_state *interrupted; + phys_addr_t interrupted_phys; + u32 evt_id; + unsigned long hart_id; + unsigned int cpu_id; +}; + +struct riscv_sse_interrupted_context { + struct pt_regs *regs; + unsigned long hstatus; +}; + +static inline bool sse_event_is_global(u32 evt) +{ + return !!(evt & SBI_SSE_EVENT_GLOBAL); +} + +void arch_sse_event_update_cpu(struct sse_event_arch_data *arch_evt, int c= pu); +int arch_sse_init_event(struct sse_event_arch_data *arch_evt, u32 evt_id, + int cpu); +void arch_sse_free_event(struct sse_event_arch_data *arch_evt); +int arch_sse_register_event(struct sse_event_arch_data *arch_evt); +void arch_sse_init_cpu(void); + +void sse_handle_event(struct sse_event_arch_data *arch_evt, + struct pt_regs *regs); +asmlinkage void handle_sse(void); +asmlinkage void noinstr do_sse(struct sse_event_arch_data *arch_evt, + struct pt_regs *regs, unsigned long hstatus); + +const struct riscv_sse_interrupted_context * +riscv_sse_get_interrupted_context(void); + +#endif + +#endif diff --git a/arch/riscv/include/asm/thread_info.h b/arch/riscv/include/asm/= thread_info.h index 55019fdfa9ec..d14b45610c73 100644 --- a/arch/riscv/include/asm/thread_info.h +++ b/arch/riscv/include/asm/thread_info.h @@ -36,6 +36,7 @@ #define OVERFLOW_STACK_SIZE SZ_4K =20 #define IRQ_STACK_SIZE THREAD_SIZE +#define SSE_STACK_SIZE THREAD_SIZE =20 #ifndef __ASSEMBLER__ =20 diff --git a/arch/riscv/kernel/Makefile b/arch/riscv/kernel/Makefile index ebe1c3588177..1e34f87e97c0 100644 --- a/arch/riscv/kernel/Makefile +++ b/arch/riscv/kernel/Makefile @@ -101,6 +101,7 @@ obj-$(CONFIG_DYNAMIC_FTRACE) +=3D mcount-dyn.o obj-$(CONFIG_PERF_EVENTS) +=3D perf_callchain.o obj-$(CONFIG_HAVE_PERF_REGS) +=3D perf_regs.o obj-$(CONFIG_RISCV_SBI) +=3D sbi.o sbi_ecall.o +obj-$(CONFIG_RISCV_SBI_SSE) +=3D sbi_sse.o sbi_sse_entry.o ifeq ($(CONFIG_RISCV_SBI), y) obj-$(CONFIG_SMP) +=3D sbi-ipi.o obj-$(CONFIG_SMP) +=3D cpu_ops_sbi.o diff --git a/arch/riscv/kernel/asm-offsets.c b/arch/riscv/kernel/asm-offset= s.c index a75f0cfea1e9..15363703cdd6 100644 --- a/arch/riscv/kernel/asm-offsets.c +++ b/arch/riscv/kernel/asm-offsets.c @@ -15,6 +15,8 @@ #include #include #include +#include +#include #include =20 void asm_offsets(void); @@ -533,6 +535,18 @@ void asm_offsets(void) DEFINE(FREGS_A6, offsetof(struct __arch_ftrace_regs, a6)); DEFINE(FREGS_A7, offsetof(struct __arch_ftrace_regs, a7)); #endif + +#ifdef CONFIG_RISCV_SBI_SSE + OFFSET(SSE_REG_EVT_STACK, sse_event_arch_data, stack); + OFFSET(SSE_REG_EVT_SHADOW_STACK, sse_event_arch_data, shadow_stack); + OFFSET(SSE_REG_EVT_TMP, sse_event_arch_data, tmp); + OFFSET(SSE_REG_HART_ID, sse_event_arch_data, hart_id); + OFFSET(SSE_REG_CPU_ID, sse_event_arch_data, cpu_id); + + DEFINE(SBI_EXT_SSE, SBI_EXT_SSE); + DEFINE(SBI_SSE_EVENT_COMPLETE, SBI_SSE_EVENT_COMPLETE); + DEFINE(ASM_NR_CPUS, CONFIG_NR_CPUS); +#endif #ifdef CONFIG_RISCV_SBI DEFINE(SBI_EXT_FWFT, SBI_EXT_FWFT); DEFINE(SBI_EXT_FWFT_SET, SBI_EXT_FWFT_SET); diff --git a/arch/riscv/kernel/entry.S b/arch/riscv/kernel/entry.S index d799c4e56f80..0b79fa7241ea 100644 --- a/arch/riscv/kernel/entry.S +++ b/arch/riscv/kernel/entry.S @@ -424,6 +424,15 @@ SYM_FUNC_END(call_on_irq_stack) * arguments are passed to schedule_tail. */ SYM_FUNC_START(__switch_to) +#ifdef CONFIG_RISCV_SBI_SSE + /* + * Mark the interval where tp changes from prev to next. SSE entry uses + * the interrupted tp while this per-CPU pointer is NULL. + */ + asm_per_cpu t0, __sbi_sse_entry_task, t1 + REG_S zero, 0(t0) +#endif + /* Save context into prev->thread */ li a4, TASK_THREAD_RA add a3, a0, a4 @@ -470,6 +479,11 @@ SYM_FUNC_START(__switch_to) REG_L s11, TASK_THREAD_S11_RA(a4) /* The offset of thread_info in task_struct is zero. */ move tp, a1 +#ifdef CONFIG_RISCV_SBI_SSE + /* Publish next only after tp contains its task_struct pointer. */ + asm_per_cpu t0, __sbi_sse_entry_task, t1 + REG_S tp, 0(t0) +#endif /* Switch to the next shadow call stack */ scs_load_current ret diff --git a/arch/riscv/kernel/sbi_sse.c b/arch/riscv/kernel/sbi_sse.c new file mode 100644 index 000000000000..7dd496e2bdb0 --- /dev/null +++ b/arch/riscv/kernel/sbi_sse.c @@ -0,0 +1,246 @@ +// SPDX-License-Identifier: GPL-2.0-or-later +/* + * Copyright (C) 2025 Rivos Inc. + */ +#include +#include +#include +#include +#include +#include + +#include +#include +#include +#include +#include +#include + +DEFINE_PER_CPU(struct task_struct *, __sbi_sse_entry_task); +static DEFINE_PER_CPU(struct riscv_sse_interrupted_context *, + riscv_sse_interrupted_context); + +const struct riscv_sse_interrupted_context * +riscv_sse_get_interrupted_context(void) +{ + return this_cpu_read(riscv_sse_interrupted_context); +} + +void __weak sse_handle_event(struct sse_event_arch_data *arch_evt, struct = pt_regs *regs) +{ +} + +void noinstr do_sse(struct sse_event_arch_data *arch_evt, + struct pt_regs *regs, unsigned long hstatus) +{ + struct riscv_sse_interrupted_context context =3D { regs, hstatus }; + struct riscv_sse_interrupted_context *previous; + struct sbiret sret; + + nmi_enter(); + instrumentation_begin(); + + /* Retrieve missing GPRs from SBI */ + sret =3D sbi_ecall(SBI_EXT_SSE, SBI_SSE_EVENT_ATTR_READ, arch_evt->evt_id, + SBI_SSE_ATTR_INTERRUPTED_A6, + (SBI_SSE_ATTR_INTERRUPTED_A7 - + SBI_SSE_ATTR_INTERRUPTED_A6) + 1, + (unsigned long)arch_evt->interrupted_phys, 0, 0); + if (sret.error) { + pr_warn("Failed to read interrupted registers for event %x: %ld\n", + arch_evt->evt_id, sret.error); + /* Let the client quiesce its source without using incomplete regs. */ + sse_handle_event(arch_evt, NULL); + goto out; + } + + memcpy(®s->a6, arch_evt->interrupted, + sizeof(*arch_evt->interrupted)); + + /* Make the interrupted frame visible while clients handle this event. */ + previous =3D this_cpu_read(riscv_sse_interrupted_context); + this_cpu_write(riscv_sse_interrupted_context, &context); + sse_handle_event(arch_evt, regs); + this_cpu_write(riscv_sse_interrupted_context, previous); + + if (memcmp(®s->a6, arch_evt->interrupted, + sizeof(*arch_evt->interrupted))) { + memcpy(arch_evt->interrupted, ®s->a6, + sizeof(*arch_evt->interrupted)); + sret =3D sbi_ecall(SBI_EXT_SSE, SBI_SSE_EVENT_ATTR_WRITE, + arch_evt->evt_id, SBI_SSE_ATTR_INTERRUPTED_A6, + (SBI_SSE_ATTR_INTERRUPTED_A7 - + SBI_SSE_ATTR_INTERRUPTED_A6) + 1, + (unsigned long)arch_evt->interrupted_phys, 0, 0); + /* + * If writeback fails, COMPLETE resumes with firmware's original + * a6/a7 rather than treating the shared buffer as committed. + */ + if (sret.error) + pr_warn("Failed to write interrupted registers for event %x: %ld\n", + arch_evt->evt_id, sret.error); + } + +out: + instrumentation_end(); + nmi_exit(); +} + +static void *alloc_to_stack_pointer(void *alloc) +{ + return alloc ? alloc + SSE_STACK_SIZE : NULL; +} + +static void *stack_pointer_to_alloc(void *stack) +{ + return stack ? stack - SSE_STACK_SIZE : NULL; +} + +static void arch_sse_flush_tlb_range(struct sse_event_arch_data *arch_evt, + unsigned long start, unsigned long size) +{ + unsigned long end =3D start + size; + + if (sse_event_is_global(arch_evt->evt_id)) + flush_tlb_kernel_range(start, end); + else + local_flush_tlb_kernel_range(start, end); +} + +static void arch_sse_shadow_stack_cpu_sync(struct sse_event_arch_data *arc= h_evt) +{ +#ifdef CONFIG_SHADOW_CALL_STACK + if (arch_evt->shadow_stack) + arch_sse_flush_tlb_range(arch_evt, + (unsigned long)arch_evt->shadow_stack, + SCS_SIZE); +#endif +} + +#ifdef CONFIG_VMAP_STACK +static void *sse_stack_alloc(unsigned int cpu) +{ + void *stack =3D arch_alloc_vmap_stack(SSE_STACK_SIZE, cpu_to_node(cpu)); + + return alloc_to_stack_pointer(stack); +} + +static void sse_stack_free(void *stack) +{ + vfree(stack_pointer_to_alloc(stack)); +} + +static void arch_sse_stack_cpu_sync(struct sse_event_arch_data *arch_evt) +{ + void *p_stack =3D arch_evt->stack; + unsigned long stack =3D (unsigned long)stack_pointer_to_alloc(p_stack); + + /* + * Flush the tlb to avoid taking any exception when accessing the + * vmapped stack inside the SSE handler + */ + arch_sse_flush_tlb_range(arch_evt, stack, SSE_STACK_SIZE); + + arch_sse_shadow_stack_cpu_sync(arch_evt); +} +#else /* CONFIG_VMAP_STACK */ +static void *sse_stack_alloc(unsigned int cpu) +{ + void *stack =3D kmalloc(SSE_STACK_SIZE, GFP_KERNEL); + + return alloc_to_stack_pointer(stack); +} + +static void sse_stack_free(void *stack) +{ + kfree(stack_pointer_to_alloc(stack)); +} + +static void arch_sse_stack_cpu_sync(struct sse_event_arch_data *arch_evt) +{ + arch_sse_shadow_stack_cpu_sync(arch_evt); +} +#endif /* CONFIG_VMAP_STACK */ + +static int sse_init_scs(int cpu, struct sse_event_arch_data *arch_evt) +{ + void *stack; + + if (!scs_is_enabled()) + return 0; + + stack =3D scs_alloc(cpu_to_node(cpu)); + if (!stack) + return -ENOMEM; + + arch_evt->shadow_stack =3D stack; + + return 0; +} + +void arch_sse_event_update_cpu(struct sse_event_arch_data *arch_evt, int c= pu) +{ + arch_evt->cpu_id =3D cpu; + arch_evt->hart_id =3D cpuid_to_hartid_map(cpu); +} + +void arch_sse_init_cpu(void) +{ + __this_cpu_write(__sbi_sse_entry_task, current); +} + +int arch_sse_init_event(struct sse_event_arch_data *arch_evt, u32 evt_id, + int cpu) +{ + void *stack; + + arch_evt->interrupted =3D kmalloc_obj(*arch_evt->interrupted, GFP_KERNEL); + if (!arch_evt->interrupted) + return -ENOMEM; + + arch_evt->evt_id =3D evt_id; + stack =3D sse_stack_alloc(cpu); + if (!stack) + goto err_free_interrupted; + + arch_evt->stack =3D stack; + + if (sse_init_scs(cpu, arch_evt)) { + sse_stack_free(arch_evt->stack); + goto err_free_interrupted; + } + + /* kmalloc keeps the two adjacent SBI attribute words contiguous. */ + arch_evt->interrupted_phys =3D virt_to_phys(arch_evt->interrupted); + + arch_sse_event_update_cpu(arch_evt, cpu); + + return 0; + +err_free_interrupted: + kfree(arch_evt->interrupted); + arch_evt->interrupted =3D NULL; + return -ENOMEM; +} + +void arch_sse_free_event(struct sse_event_arch_data *arch_evt) +{ + scs_free(arch_evt->shadow_stack); + sse_stack_free(arch_evt->stack); + kfree(arch_evt->interrupted); +} + +int arch_sse_register_event(struct sse_event_arch_data *arch_evt) +{ + struct sbiret sret; + + arch_sse_stack_cpu_sync(arch_evt); + + sret =3D sbi_ecall(SBI_EXT_SSE, SBI_SSE_EVENT_REGISTER, arch_evt->evt_id, + (unsigned long)handle_sse, (unsigned long)arch_evt, 0, + 0, 0); + if (sret.error =3D=3D SBI_ERR_NOT_SUPPORTED) + return -EOPNOTSUPP; + + return sbi_err_map_linux_errno(sret.error); +} diff --git a/arch/riscv/kernel/sbi_sse_entry.S b/arch/riscv/kernel/sbi_sse_= entry.S new file mode 100644 index 000000000000..e0e8efba12dd --- /dev/null +++ b/arch/riscv/kernel/sbi_sse_entry.S @@ -0,0 +1,226 @@ +/* SPDX-License-Identifier: GPL-2.0-or-later */ +/* + * Copyright (C) 2025 Rivos Inc. + */ + +#include +#include + +#include +#include +#include +#include +#include + +/* When entering handle_sse, the following registers are set: + * a6: contains the hartid + * a7: contains a sse_event_arch_data struct pointer + */ +SYM_CODE_START(handle_sse) + /* Save stack temporarily */ + REG_S sp, SSE_REG_EVT_TMP(a7) + /* Set entry stack */ + REG_L sp, SSE_REG_EVT_STACK(a7) + + addi sp, sp, -(PT_SIZE_ON_STACK) + REG_S ra, PT_RA(sp) + REG_S s0, PT_S0(sp) + REG_S s1, PT_S1(sp) + REG_S s2, PT_S2(sp) + REG_S s3, PT_S3(sp) + REG_S s4, PT_S4(sp) + REG_S s5, PT_S5(sp) + REG_S s6, PT_S6(sp) + REG_S s7, PT_S7(sp) + REG_S s8, PT_S8(sp) + REG_S s9, PT_S9(sp) + REG_S s10, PT_S10(sp) + REG_S s11, PT_S11(sp) + REG_S tp, PT_TP(sp) + REG_S t0, PT_T0(sp) + REG_S t1, PT_T1(sp) + REG_S t2, PT_T2(sp) + REG_S t3, PT_T3(sp) + REG_S t4, PT_T4(sp) + REG_S t5, PT_T5(sp) + REG_S t6, PT_T6(sp) + REG_S gp, PT_GP(sp) + REG_S a0, PT_A0(sp) + REG_S a1, PT_A1(sp) + REG_S a2, PT_A2(sp) + REG_S a3, PT_A3(sp) + REG_S a4, PT_A4(sp) + REG_S a5, PT_A5(sp) + + /* Retrieve entry sp */ + REG_L a4, SSE_REG_EVT_TMP(a7) + /* Save CSRs */ + csrr a0, CSR_EPC + csrr a1, CSR_SSTATUS + csrr a2, CSR_STVAL + csrr a3, CSR_SCAUSE + + REG_S a0, PT_EPC(sp) + REG_S a1, PT_STATUS(sp) + REG_S a2, PT_BADADDR(sp) + REG_S a3, PT_CAUSE(sp) + REG_S a4, PT_SP(sp) + + /* Disable user memory access and floating/vector computing */ + li t0, SR_SUM | SR_FS_VS + csrc CSR_STATUS, t0 + + load_global_pointer + scs_load_sse_stack a7 + +#ifdef CONFIG_SMP + REG_L t4, SSE_REG_HART_ID(a7) + lw t3, SSE_REG_CPU_ID(a7) + + bne t4, a6, .Lfind_hart_id_slowpath + +.Lcpu_id_found: +#else + mv t3, zero +#endif + + asm_per_cpu_with_cpu t2 __sbi_sse_entry_task t1 t3 + REG_L tp, 0(t2) + bnez tp, .Lcurrent_task_found + + /* __switch_to() marks its transition window with a NULL entry task. */ + REG_L tp, PT_TP(sp) + +.Lcurrent_task_found: + /* Nested exceptions temporarily replace these with the SSE stack. */ + REG_L s6, TASK_TI_KERNEL_SP(tp) + REG_L s7, TASK_TI_USER_SP(tp) + + mv a1, sp /* pt_regs on stack */ + + /* + * Run the SSE handler with the normal exception vector, but restore the + * interrupted stvec before completing the event. SSE can arrive while + * the kernel is using a temporary trap vector in a sensitive entry path. + */ + csrr s3, CSR_STVEC + la t0, handle_exception + csrw CSR_STVEC, t0 + + /* + * Preserve the full HS-mode virtualization state across the handler. + * hstatus is live supervisor state rather than an SSE interrupted + * attribute, and OpenSBI consumes hstatus.SPV during event completion. + * Saving the whole CSR keeps the handler episode transparent to KVM and + * avoids having to infer which hstatus bits may matter to a guest resume. + */ + li s5, 0 + ALTERNATIVE("nop", "csrr s5, hstatus", 0, RISCV_ISA_EXT_H, 1) + + /* + * Save sscratch for restoration since we might have interrupted the + * kernel in early exception path and thus, we don't know the content of + * sscratch. + */ + csrrw s4, CSR_SSCRATCH, x0 + + mv a0, a7 + mv a2, s5 + + call do_sse + + /* Leave no reference to the event stack in the interrupted task. */ + REG_S s7, TASK_TI_USER_SP(tp) + REG_S s6, TASK_TI_KERNEL_SP(tp) + + csrw CSR_SSCRATCH, s4 + ALTERNATIVE("nop", "csrw hstatus, s5", 0, RISCV_ISA_EXT_H, 1) + csrw CSR_STVEC, s3 + + REG_L a0, PT_STATUS(sp) + REG_L a1, PT_EPC(sp) + REG_L a2, PT_BADADDR(sp) + REG_L a3, PT_CAUSE(sp) + csrw CSR_SSTATUS, a0 + csrw CSR_EPC, a1 + csrw CSR_STVAL, a2 + csrw CSR_SCAUSE, a3 + + REG_L ra, PT_RA(sp) + REG_L s0, PT_S0(sp) + REG_L s1, PT_S1(sp) + REG_L s2, PT_S2(sp) + REG_L s3, PT_S3(sp) + REG_L s4, PT_S4(sp) + REG_L s5, PT_S5(sp) + REG_L s6, PT_S6(sp) + REG_L s7, PT_S7(sp) + REG_L s8, PT_S8(sp) + REG_L s9, PT_S9(sp) + REG_L s10, PT_S10(sp) + REG_L s11, PT_S11(sp) + REG_L tp, PT_TP(sp) + REG_L t0, PT_T0(sp) + REG_L t1, PT_T1(sp) + REG_L t2, PT_T2(sp) + REG_L t3, PT_T3(sp) + REG_L t4, PT_T4(sp) + REG_L t5, PT_T5(sp) + REG_L t6, PT_T6(sp) + REG_L gp, PT_GP(sp) + REG_L a0, PT_A0(sp) + REG_L a1, PT_A1(sp) + REG_L a2, PT_A2(sp) + REG_L a3, PT_A3(sp) + REG_L a4, PT_A4(sp) + REG_L a5, PT_A5(sp) + + REG_L sp, PT_SP(sp) + + li a7, SBI_EXT_SSE + li a6, SBI_SSE_EVENT_COMPLETE + ecall + + /* + * COMPLETE must resume the interrupted context and never return. Trap + * through the normal vector instead of falling into adjacent assembly. + */ + la t0, handle_exception + csrw CSR_STVEC, t0 + ebreak + /* The fatal trap must not return; execution should never reach here. */ + +#ifdef CONFIG_SMP +.Lfind_hart_id_slowpath: + + /* Restore current task struct from __sbi_sse_entry_task */ + li t1, ASM_NR_CPUS + /* Slowpath to find the CPU id associated to the hart id */ + la t0, __cpuid_to_hartid_map + li t3, 0 + +.Lhart_id_loop: + REG_L t2, 0(t0) + beq t2, a6, .Lcpu_id_found + + /* Increment pointer and CPU number */ + addi t3, t3, 1 + addi t0, t0, RISCV_SZPTR + bltu t3, t1, .Lhart_id_loop + + /* + * This should never happen since we expect the hart_id to match one + * of our CPU, but better be safe than sorry + */ + la tp, init_task + la a0, sse_hart_id_panic_string + la t0, panic + jalr t0 +#endif + +SYM_CODE_END(handle_sse) +ASM_NOKPROBE(handle_sse) + +SYM_DATA_START_LOCAL(sse_hart_id_panic_string) + .ascii "Unable to match hart_id with cpu\0" +SYM_DATA_END(sse_hart_id_panic_string) --=20 2.50.1 (Apple Git-155) From nobody Fri Sep 25 21:02:40 2026 Received: from mail-pj2-f34.google.com (mail-pj2-f34.google.com [74.125.227.162]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EAC01489FB2 for ; Mon, 21 Sep 2026 11:16:51 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.227.162 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789989414; cv=none; b=XaPe3pENxGAcs0vpJOYbdhZtDKtcQ7lh6RFNzbc5G080VT7572E78cG2+QE74xoWMxoITlWyn5bnnudOP3HA1YIQg+aI8UE2rgdckcCaCM1do5FOwo0ar/d5LM+v+tx7DaEbM8g/1W2QweHxLyfjEg488f5PrEi19FdXmP4bGyo= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789989414; c=relaxed/simple; bh=6j/r9yvub7FH1ktXyNHoVKCJiNGlNg0k7LU+GyyCF9o=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=T8iWJDI7j/674LUQOrBvVQbh4+FlCGyLwLPtBxvmJprV9ks6esi5nE/OixXeyDNiRygh0AmHTMcYqBnoPECVC69UhXCmZcqmdLBrmH3xgF5Q1paTPaak4qsJ8G+jHmnoTeVAoJNKL9YPRfQH2vzGpynlKD0T6Ge6owOO8cNEskU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com; spf=pass smtp.mailfrom=bytedance.com; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b=AhdgyaWm; arc=none smtp.client-ip=74.125.227.162 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=bytedance.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b="AhdgyaWm" Received: by mail-pj2-f34.google.com with SMTP id d9443c01a7336-2d747eb79f7so18938975ad.1 for ; Mon, 21 Sep 2026 04:16:51 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=bytedance.com; s=google; t=1789989411; x=1790594211; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=slvqYs+iaXZQUFiWP3raPQ3MYOooTUinSrFApU/wWdw=; b=AhdgyaWm8UkNABHKDWFTlNX2IN/lX1KTX3OCuASLm23yMrhtHPNFfgCi1GjFNjSHgo A90Yin1LZlYloPz5zikdtFLNObXT+C98pPryLlQkZevfXDy9NgDTZn05cZxuvbyvLf0k rUgVALDeB8Gti5pJZGKEHtLSOfIpdRTrF8LUfiNjP3I+vnQ6La/hXRAj8TJyJPs97rVL KgG2yUPBURITG9LPU2xTWmlLf1me+2pLrsYuV1QNYvK4QFMTmye0+66ngOmq/Y7qI0AL ksQXqyDQpX8ScwphU4WLvnXU3OqIFalINLVT/LAs8P8/wXxTgm5y0Mc8nJ+SGaUyyK7Y m4tA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789989411; x=1790594211; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=slvqYs+iaXZQUFiWP3raPQ3MYOooTUinSrFApU/wWdw=; b=bT51Axb/3BynZUk1FInDCyp7XN2U14RtPieb7l0tZpk5cyPdWYgtXSstzSW5+SFsWH eq2J/usauzIwOkkBHVLiiGjEUWGlJBdOP5nE9KD7lvXr7NE00O0Pa/x/LSRfB9wiwPWu cuw1DlG7AgVjz+6N4s9yFtJCdfhzq/dIZDJ9tDM2WLCsqugBghifbTPwCoDTTCBPORCN 1RV3Oks9wJEVwPpuicZAL+TVJFi+XpZ4SLBrnnHpt/sHw0PEgReOiG0E5/fVmR/Xfyyg mQXIUYr6VG56n8Gqm7Jer9IMB0Duw3Z3C1c/w3yDCk3honbUSZkgSNRgYe7N4uwMb4TD aSxA== X-Forwarded-Encrypted: i=1; AKwUvBzfYOBJYNuAG4iVzL3H26ZUvYMFlPzcDiRbN8Oye+scq9F5kh79KLLJDJ3gQMwLQ8Zqu8MPqQ5nbLHPqKs=@vger.kernel.org X-Gm-Message-State: AFuF++lFYaQkFjG2kwfJSfhUMMYu4FcvtT/U8Gb0DK3k7h+HnIGFx83h NIN8khbdkThog3jgEqia1kfZ1A7KyZcwFJvIaD9mNf4itn/3ZrbGEUR7QyqB6ppUp58= X-Gm-Gg: AYBFou1LhnoFi8JUlSydAtzmewrE4Xyea6U9KVpqd2BZAtqIKpUvtsFg46fX1yH3c0Z tF2vpygWUbPZWHS9kJRU1OIhDD/u7laK5zywwdYuNhdvPxkKRNxB1MoME7zP+SMAcOkpYPCJMvr edZB4Cv2mSg0xZoL4rzZTAbwash7LabAQRJz6RxMNLS5WYZsKmSdM4dT+iPaUegb4UjRZV8vUX6 hDPKR7BviMGq2GIoWlmZTdXgJS3rpFjLxXUCSbwLmWCnU4jIE+lgY4YahelZb8VYWUjQaUhgDQP 7l1zZR0EcMiBlrsFgsEM8Qb/83yxKTUULKl5Hnia8D3PjMfcY/Yrc4oy8kBOR9xUF6/9ZD/Z0ws XCxOWqqgtejvftW0eQJds/yTCo/ptxdAS1LNpTkNy5xL+my1hynGS/DHI5E3SeONGO1ULlKeEqx rEQcf/4v5K3rA2+W3iYlJ8ZWbYhgP0k/3ojoEa4F4dkOY08oMzLocK0PHAtaRtuHAbBIUwcrDjB VQ/fb1Yh4/pQq+fvGcvvjHLDdM1gJxMVguVrh+HaXv+5Q== X-Received: by 2002:a17:902:d2cb:b0:2df:45f7:6386 with SMTP id d9443c01a7336-2df45f76498mr38924045ad.15.1789989410710; Mon, 21 Sep 2026 04:16:50 -0700 (PDT) Received: from FJ7FR2JRQ3.bytedance.net ([240e:694:e20:401::8]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2ddc17ba651sm31831125ad.53.2026.09.21.04.16.26 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Mon, 21 Sep 2026 04:16:50 -0700 (PDT) From: Zhanpeng Zhang To: Paul Walmsley , Palmer Dabbelt , Albert Ou , Alexandre Ghiti , Himanshu Chauhan , Conor Dooley , Anup Patel Cc: =?UTF-8?q?Cl=C3=A9ment=20L=C3=A9ger?= , Yunhui Cui , Atish Patra , Peter Zijlstra , Ingo Molnar , Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Alexander Shishkin , Jiri Olsa , Ian Rogers , Adrian Hunter , James Clark , Will Deacon , Thomas Gleixner , Jonathan Corbet , Randy Dunlap , Shuah Khan , Shuah Khan , Yuanzhu , Yicong Yang , Susheng Yang , linux-riscv@lists.infradead.org, linux-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-arm-kernel@lists.infradead.org, Zhanpeng Zhang Subject: [PATCH v10 RESEND 3/9] riscv: sse: mask events during shutdown and kexec Date: Mon, 21 Sep 2026 19:15:00 +0800 Message-ID: <03472b45125793a768f3e5bbc168f167e673d4cc.1789974241.git.zhangzhanpeng.jasper@bytedance.com> X-Mailer: git-send-email 2.50.1 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" SSE delivery is independent of normal S-mode interrupts. Firmware may also retain an event registration until Linux explicitly unregisters it. A hart must therefore stop accepting SSE events before Linux stops servicing the registered handler. Mask SSE on the local hart before panic stop, CPU stop, restart, poweroff, and crash shutdown paths. This prevents firmware from entering Linux-owned handler state after the corresponding CPU or kernel context is no longer valid. A crash kernel cannot identify or take ownership of registrations inherited from the crashed kernel. Reject a later normal kexec while such SSE state may still exist, rather than transferring unknown firmware state to another kernel. Signed-off-by: Zhanpeng Zhang --- arch/riscv/kernel/machine_kexec.c | 11 +++++++++++ arch/riscv/kernel/reset.c | 18 ++++++++++++++++++ arch/riscv/kernel/smp.c | 17 +++++++++++++++++ 3 files changed, 46 insertions(+) diff --git a/arch/riscv/kernel/machine_kexec.c b/arch/riscv/kernel/machine_= kexec.c index 738df176ff6f..24e7affae70b 100644 --- a/arch/riscv/kernel/machine_kexec.c +++ b/arch/riscv/kernel/machine_kexec.c @@ -14,10 +14,13 @@ #include /* For set_memory_x() */ #include /* For unreachable() */ #include /* For cpu_down() */ +#include #include #include #include =20 +#include + /* * machine_kexec_prepare - Initialize kexec * @@ -36,6 +39,13 @@ machine_kexec_prepare(struct kimage *image) unsigned int control_code_buffer_sz =3D 0; int i =3D 0; =20 + /* A crash kernel cannot tear down registrations inherited from firmware.= */ + if (is_kdump_kernel() && image->type !=3D KEXEC_TYPE_CRASH && + riscv_sse_available()) { + pr_err("Normal kexec from a crash kernel is unsupported with SSE\n"); + return -EOPNOTSUPP; + } + /* Find the Flattened Device Tree and save its physical address */ for (i =3D 0; i < image->nr_segments; i++) { if (image->segment[i].memsz <=3D sizeof(fdt)) @@ -127,6 +137,7 @@ void machine_crash_shutdown(struct pt_regs *regs) { local_irq_disable(); + riscv_sse_mask_current_hart(); =20 /* shutdown non-crashing cpus */ crash_smp_send_stop(); diff --git a/arch/riscv/kernel/reset.c b/arch/riscv/kernel/reset.c index 14eb08a6db85..fdab37e7ae52 100644 --- a/arch/riscv/kernel/reset.c +++ b/arch/riscv/kernel/reset.c @@ -6,6 +6,20 @@ #include #include #include +#include + +#include + +#ifndef CONFIG_SMP +void __noreturn panic_smp_self_stop(void) +{ + riscv_sse_mask_current_hart(); + local_irq_disable(); + + for (;;) + cpu_relax(); +} +#endif =20 static void __noreturn default_power_off(void) { @@ -18,6 +32,8 @@ EXPORT_SYMBOL(pm_power_off); =20 void machine_restart(char *cmd) { + riscv_sse_mask_current_hart(); + /* * UpdateCapsule() depends on the system being reset via ResetSystem(). */ @@ -30,12 +46,14 @@ void machine_restart(char *cmd) =20 void machine_halt(void) { + riscv_sse_mask_current_hart(); do_kernel_power_off(); default_power_off(); } =20 void machine_power_off(void) { + riscv_sse_mask_current_hart(); do_kernel_power_off(); default_power_off(); } diff --git a/arch/riscv/kernel/smp.c b/arch/riscv/kernel/smp.c index fa66f9c97d74..7f0d4e7332f1 100644 --- a/arch/riscv/kernel/smp.c +++ b/arch/riscv/kernel/smp.c @@ -23,10 +23,13 @@ #include #include #include +#include =20 #include #include #include +#include +#include =20 enum ipi_message_type { IPI_RESCHEDULE, @@ -79,8 +82,18 @@ int riscv_hartid_to_cpuid(unsigned long hartid) return -ENOENT; } =20 +void __noreturn panic_smp_self_stop(void) +{ + riscv_sse_mask_current_hart(); + local_irq_disable(); + + for (;;) + cpu_relax(); +} + static void ipi_stop(void) { + riscv_sse_mask_current_hart(); set_cpu_online(smp_processor_id(), false); while (1) wait_for_interrupt(); @@ -91,6 +104,7 @@ static atomic_t waiting_for_crash_ipi =3D ATOMIC_INIT(0); =20 static inline void ipi_cpu_crash_stop(unsigned int cpu, struct pt_regs *re= gs) { + riscv_sse_mask_current_hart(); crash_save_cpu(regs, cpu); =20 atomic_dec(&waiting_for_crash_ipi); @@ -254,6 +268,8 @@ void smp_send_stop(void) { unsigned long timeout; =20 + riscv_sse_mask_current_hart(); + if (num_online_cpus() > 1) { cpumask_t mask; =20 @@ -301,6 +317,7 @@ void crash_smp_send_stop(void) return; =20 cpus_stopped =3D 1; + riscv_sse_mask_current_hart(); =20 /* * If this cpu is the only one alive at this point in time, online or --=20 2.50.1 (Apple Git-155) From nobody Fri Sep 25 21:02:40 2026 Received: from mail-pj2-f42.google.com (mail-pj2-f42.google.com [74.125.227.170]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6E71648C8D2 for ; Mon, 21 Sep 2026 11:17:16 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.227.170 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789989440; cv=none; b=Y4qhaUfJahDOyvPnlzg2PX+M/QCZ7+XNBw3n0bolZ/cHYEuUIn5LaZiJxve8hRqCGcC+tXTrI2mHy28O38a6jxgOJcaTTQQ59WXtuDT/j1HOqTfrwznI7FnQSZdPaYukO13nDk/E0+88mrATnqBlRrGELWbe//BdmlaFc+gNdhw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789989440; c=relaxed/simple; bh=TW7NE5mu4DhgBxnzqeoMNUpDP/0x2IV0CPFAlNMMMcU=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=gPOFCdhuXDp6igeEKB8qoc6jzNTf+T5w/lR5r/zAKwz/SHL4vyS8+4SUOUm4xbQJeSLtinhtWMbuXT0EuGFnj8jb9tL2SQ9XQeEcgL5Gn1mDcvs/r7pA4Gf5z6woAEbCK6S5Rbteb2wSZ0UGFT4XR5j6pyoCaCbWEgr3pi+/3S4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com; spf=pass smtp.mailfrom=bytedance.com; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b=bKskE4zD; arc=none smtp.client-ip=74.125.227.170 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=bytedance.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b="bKskE4zD" Received: by mail-pj2-f42.google.com with SMTP id d9443c01a7336-2df4aa80a73so5891945ad.3 for ; Mon, 21 Sep 2026 04:17:16 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=bytedance.com; s=google; t=1789989435; x=1790594235; darn=vger.kernel.org; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:from:to:cc:subject :date:message-id:reply-to:content-type; bh=NPc9pg6qzsYsf4qC29X3rATyfhrDb/9hBTEJ6abWqD8=; b=bKskE4zD/JoBu0lWV1DwDum0GlOIkaVLwbh99rreEAQy1AFxIzEvb6AS8BhvXGEuVx l636u3+sFJkLBoRRzX9CelJ4sj+QlFAptlJPZHKjqYsgQEtodh6St0gXzF1Ro0g0HFww ypXm1/38nxJG8v7tRBJyMP4OBAFIMN0RHFyezd8tS461GL4dGrsIKOKyIXYJQAzB8jEH tKXF0byAqHE2OITCcBtpGmuoYeMAo38njsEOWIMuUhbpaG3i29LD2OIopllEt66jH8kU YleJJXSVEBS4pCkZB63jPQGMD0+byLu/hb+2mrc3gw/xYUstFhV/J/Q5BxInENKrHCpf B7Ew== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789989435; x=1790594235; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=NPc9pg6qzsYsf4qC29X3rATyfhrDb/9hBTEJ6abWqD8=; b=x0miwhbY+DlviPJBN4v6I3uAlBJ/d4AplL3+Ww2jwxfQcCbPhiYVCIgadJ0J3dVfpm tKBlZpnF2/VGlcM/EBK4IvRKMFDwIlhBOjg9Z17rinZrymq7CyzlbIfibeeWpTvmZg68 NCOm3TOwXpFASgs7IYAqSWo0atfpa0hMJC5L1QJHUV7s54UHTzHjNgjWOdbL6JijwYpa F+/xGu4HWkqwFjRXRcMT9X2L5/cXKFKuyRdAo3II5aLXUudKDUAjn4aWntDtiuhoYQzS BPK6jswSHjiJ46mL/EEgDd0dyCvVhMY2E6u8Edc464uaCYp9CpkpsrJbTfSQ/A2jaLZj tf7A== X-Forwarded-Encrypted: i=1; AKwUvBzHJduTDY7GMUPSaWjnI9ZpF98pQO8c4WYKlu7Frcq9jV7bCZiPWR+uZmCDda4asgtqnrhozK1DPqKpOKU=@vger.kernel.org X-Gm-Message-State: AFuF++kDlg9DXANJ/03K763tZPxPZoS44TOHiGRNrGjUcMqe6OkNE8+h DcDnwAZ1E4oHijcyPOX2gktP2qe2/DuJyaEJzxbCJf6fYBPB1IQEzkjHx4ME1BJ5Dlk= X-Gm-Gg: AYBFou1EALojgVxf9ulRFBFx4CBkuCHWAhhUtglQ1Zst7RVh7KBzw/VneoGImM/gdHF FYPnUmjCimhimyZrnZFA0MzUZ7fE2R7TMSm+P04D3QNZDFodTuC+rtFe9QRHL5qUAd1dfq2M6Kb xnr5jn3+dsSEHcXbA05ZPcdrusHi0NCT3j6iqf48qGtUUT9FeTU+4UzcFrEkpJ7/ek2pdc6YN5o skAnlVxpVcAjE6UMdHRxoZK6RV5rqsTszgJ6txGU1l6ka4oUQ6sYsjEgc2uNK0zKsEyuRVWjPTO ZA9kOzbOP7a3WzTIJleX7ZrhDqXcvJTkvY8Z1q5rnXsHokRdANYVSliXNqohbiZfvu4pshyvweI BxoFrXpv5C+wFATmizVB4Z5VzJKQrf361Co9yvSgAYi78sMHGnvENyF732u9EIzEcyhWKO8mQ1N rw4GULfc0Nem/SrRFD2V8vLWnD6l7hhnOKSR/nII0W9MjM5xOjjfb6CptV9mTQILuVXQiOpXT2h pLEeRvLU4+/akESJzqWLA6+6AEnKgDiOl4= X-Received: by 2002:a17:902:f214:b0:2dd:c053:82f6 with SMTP id d9443c01a7336-2ddc0538455mr69471685ad.45.1789989435140; Mon, 21 Sep 2026 04:17:15 -0700 (PDT) Received: from FJ7FR2JRQ3.bytedance.net ([240e:694:e20:401::8]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2ddc17ba651sm31831125ad.53.2026.09.21.04.16.51 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Mon, 21 Sep 2026 04:17:14 -0700 (PDT) From: Zhanpeng Zhang To: Paul Walmsley , Palmer Dabbelt , Albert Ou , Alexandre Ghiti , Himanshu Chauhan , Conor Dooley , Anup Patel Cc: =?UTF-8?q?Cl=C3=A9ment=20L=C3=A9ger?= , Yunhui Cui , Atish Patra , Peter Zijlstra , Ingo Molnar , Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Alexander Shishkin , Jiri Olsa , Ian Rogers , Adrian Hunter , James Clark , Will Deacon , Thomas Gleixner , Jonathan Corbet , Randy Dunlap , Shuah Khan , Shuah Khan , Yuanzhu , Yicong Yang , Susheng Yang , linux-riscv@lists.infradead.org, linux-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-arm-kernel@lists.infradead.org, Zhanpeng Zhang , Conor Dooley Subject: [PATCH v10 RESEND 4/9] drivers: firmware: add riscv SSE support Date: Mon, 21 Sep 2026 19:15:01 +0800 Message-ID: X-Mailer: git-send-email 2.50.1 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable From: Cl=C3=A9ment L=C3=A9ger Add a driver-level interface for RISC-V SSE. Linux clients can register handlers, select a target CPU for global events, and enable, disable or unregister events. The architecture entry wrapper completes an event after the registered handler returns. PMU and GHES drivers can use this interface. Represent global events with one firmware registration and local events with one registration per possible CPU. Keep registration and enable state stable across CPU hotplug, and validate firmware-provided hart IDs before converting them to Linux CPU IDs. Use phys_addr_t for attribute buffers to match the physical addresses passed to firmware. Require MMU support because the current event-stack implementation relies on vmapped memory and TLB synchronization. Serialize client list updates with the SSE mutex and the CPU read lock. Normal CPU hotplug callbacks provide the matching write-side exclusion, and CPUHP state removal holds the SSE mutex. These rules avoid holding an additional spinlock across firmware calls. Local event operations select a per-CPU registration, so require callers to remain on the current CPU and use lockdep assertions to verify that contract. Because CPU hotplug callbacks run in a preemptible thread, disable preemption around the local operations they perform on the offlining or onlining hart. Propagate firmware failures from register, disable and unregister operations. Update Linux state only after successful firmware operations, and release an event only after all registrations are gone. Preserve the difference between SBI_ERR_NOT_SUPPORTED and a generic SBI failure so clients only select a fallback after firmware explicitly rejects an SSE operation. If setup fails partway through a local event, roll back only the CPU instances changed by that invocation. Keep non-fallback errors across CPUs. This lets a client distinguish an event rejected as unsupported by every failing hart from an unknown firmware failure. A failed registration rollback can leave firmware state without a client handle. Retain these events on a driver-owned cleanup list with a no-op handler. This lets CPU hotplug and shutdown retry cleanup without relying on client callback lifetime. Mask SSE before CPU teardown. If a normal CPU-offline teardown fails after partially changing local events, reconstruct their requested registration and enable state. Unmask the hart, then return the original error to abort the offline operation. Shutdown and CPUHP state removal cannot fail, so record incomplete cleanup and refuse an unsafe normal kexec instead. Close enable admission before reboot or kexec teardown. Protect the shutdown check and firmware enable operation with RCU, then drain existing enable calls before CPUHP removes registrations. This prevents a client from re-enabling an event between the teardown disable and unregister operations. A crash kernel cannot identify registrations inherited from the crashed kernel. Leave SSE masked, report the retained firmware state separately from an unavailable extension, and do not let clients select a delivery path that may still be owned by the old kernel. Signed-off-by: Cl=C3=A9ment L=C3=A9ger Co-developed-by: Himanshu Chauhan Signed-off-by: Himanshu Chauhan Co-developed-by: Zhanpeng Zhang Signed-off-by: Zhanpeng Zhang Acked-by: Conor Dooley --- MAINTAINERS | 7 + drivers/firmware/Kconfig | 1 + drivers/firmware/Makefile | 1 + drivers/firmware/riscv/Kconfig | 18 + drivers/firmware/riscv/Makefile | 3 + drivers/firmware/riscv/riscv_sbi_sse.c | 1219 ++++++++++++++++++++++++ include/linux/cpuhotplug.h | 1 + include/linux/riscv_sbi_sse.h | 89 ++ 8 files changed, 1339 insertions(+) create mode 100644 drivers/firmware/riscv/Kconfig create mode 100644 drivers/firmware/riscv/Makefile create mode 100644 drivers/firmware/riscv/riscv_sbi_sse.c create mode 100644 include/linux/riscv_sbi_sse.h diff --git a/MAINTAINERS b/MAINTAINERS index 7a6af19f7e30..0070a3ed6321 100644 --- a/MAINTAINERS +++ b/MAINTAINERS @@ -23501,6 +23501,13 @@ T: git git://git.kernel.org/pub/scm/linux/kernel/g= it/iommu/linux.git F: Documentation/devicetree/bindings/iommu/riscv,iommu.yaml F: drivers/iommu/riscv/ =20 +RISC-V FIRMWARE DRIVERS +M: Conor Dooley +L: linux-riscv@lists.infradead.org +S: Maintained +T: git git://git.kernel.org/pub/scm/linux/kernel/git/conor/linux.git +F: drivers/firmware/riscv/ + RISC-V MICROCHIP SUPPORT M: Conor Dooley M: Daire McNamara diff --git a/drivers/firmware/Kconfig b/drivers/firmware/Kconfig index b7cc11e4fbfa..cf0212576dfd 100644 --- a/drivers/firmware/Kconfig +++ b/drivers/firmware/Kconfig @@ -306,6 +306,7 @@ source "drivers/firmware/meson/Kconfig" source "drivers/firmware/microchip/Kconfig" source "drivers/firmware/psci/Kconfig" source "drivers/firmware/qcom/Kconfig" +source "drivers/firmware/riscv/Kconfig" source "drivers/firmware/samsung/Kconfig" source "drivers/firmware/smccc/Kconfig" source "drivers/firmware/tegra/Kconfig" diff --git a/drivers/firmware/Makefile b/drivers/firmware/Makefile index be46f1e1dc77..879953c316dc 100644 --- a/drivers/firmware/Makefile +++ b/drivers/firmware/Makefile @@ -35,6 +35,7 @@ obj-y +=3D efi/ obj-y +=3D imx/ obj-y +=3D psci/ obj-y +=3D qcom/ +obj-y +=3D riscv/ obj-y +=3D samsung/ obj-y +=3D smccc/ obj-y +=3D tegra/ diff --git a/drivers/firmware/riscv/Kconfig b/drivers/firmware/riscv/Kconfig new file mode 100644 index 000000000000..d15ff84e1258 --- /dev/null +++ b/drivers/firmware/riscv/Kconfig @@ -0,0 +1,18 @@ +# SPDX-License-Identifier: GPL-2.0-only +menu "Risc-V Specific firmware drivers" +depends on RISCV + +config RISCV_SBI_SSE + bool "Enable SBI Supervisor Software Events support" + depends on RISCV_SBI && MMU && !HIBERNATION + default y + help + The Supervisor Software Events support allows the SBI to deliver + NMI-like notifications to the supervisor mode software. When enabled, + this option provides support to register callbacks on specific SSE + events. + + Hibernation is not supported because firmware registrations do not + survive restoring a kernel image. + +endmenu diff --git a/drivers/firmware/riscv/Makefile b/drivers/firmware/riscv/Makef= ile new file mode 100644 index 000000000000..c8795d4bbb2e --- /dev/null +++ b/drivers/firmware/riscv/Makefile @@ -0,0 +1,3 @@ +# SPDX-License-Identifier: GPL-2.0 + +obj-$(CONFIG_RISCV_SBI_SSE) +=3D riscv_sbi_sse.o diff --git a/drivers/firmware/riscv/riscv_sbi_sse.c b/drivers/firmware/risc= v/riscv_sbi_sse.c new file mode 100644 index 000000000000..0cb738bab6c3 --- /dev/null +++ b/drivers/firmware/riscv/riscv_sbi_sse.c @@ -0,0 +1,1219 @@ +// SPDX-License-Identifier: GPL-2.0-or-later +/* + * Copyright (C) 2025 Rivos Inc. + */ + +#define pr_fmt(fmt) "sse: " fmt + +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include +#include + +struct sse_event { + struct list_head list; + u32 evt_id; + u32 priority; + sse_event_handler_fn __rcu *handler; + void *handler_arg; + /* Only valid for global events */ + unsigned int cpu; + /* + * Desired state requested by the client. Firmware state is tracked per + * instance because a failed transition can leave a partial result. + */ + bool enable_requested; + /* Registration failed, but firmware state still needs driver cleanup. */ + bool cleanup_pending; + + union { + struct sse_registered_event *global; + struct sse_registered_event __percpu *local; + }; +}; + +static bool sse_available __ro_after_init; +static bool sse_fw_state_retained __ro_after_init; +static bool sse_shutting_down; +static atomic_t sse_teardown_failed =3D ATOMIC_INIT(0); +/* + * Client-side updates hold sse_mutex and the CPU read lock. Normal CPU ho= tplug + * callbacks exclude them through the CPUHP write side. Initialization rep= lay + * runs before clients can register, and state removal holds sse_mutex. + */ +static LIST_HEAD(events); +static DEFINE_MUTEX(sse_mutex); + +/* + * A registration rollback can fail before the client receives an event + * handle. Keep the retained event independent of client-owned callback te= xt + * and data while the driver retries firmware cleanup. + */ +static int sse_cleanup_event_handler(u32 evt, void *arg, struct pt_regs *r= egs) +{ + return 0; +} + +struct sse_registered_event { + struct sse_event_arch_data arch; + struct sse_event *event; + unsigned long attr; + /* + * Actual firmware state for one global or per-CPU instance. A retry can + * then skip instances that already completed a partial transition. + */ + bool is_registered; + bool is_enabled; +}; + +void sse_handle_event(struct sse_event_arch_data *arch_event, + struct pt_regs *regs) +{ + sse_event_handler_fn *handler; + int ret; + struct sse_registered_event *reg_evt =3D + container_of(arch_event, struct sse_registered_event, arch); + struct sse_event *evt =3D reg_evt->event; + + rcu_read_lock(); + handler =3D rcu_dereference(evt->handler); + ret =3D handler(evt->evt_id, evt->handler_arg, regs); + rcu_read_unlock(); + if (ret) + pr_warn("event %x handler failed with error %d\n", evt->evt_id, ret); +} + +static struct sse_event *sse_event_get(u32 evt) +{ + struct sse_event *event; + + lockdep_assert_held(&sse_mutex); + + list_for_each_entry(event, &events, list) { + if (event->evt_id =3D=3D evt) + return event; + } + + return NULL; +} + +static phys_addr_t sse_event_get_attr_phys(struct sse_registered_event *re= g_evt) +{ + phys_addr_t phys; + void *addr =3D ®_evt->attr; + + if (sse_event_is_global(reg_evt->event->evt_id)) + phys =3D virt_to_phys(addr); + else + phys =3D per_cpu_ptr_to_phys(addr); + + return phys; +} + +static struct sse_registered_event *sse_get_reg_evt(struct sse_event *even= t) +{ + if (sse_event_is_global(event->evt_id)) + return event->global; + else + return per_cpu_ptr(event->local, smp_processor_id()); +} + +static int sse_err_map_linux_errno(long err) +{ + if (err =3D=3D SBI_ERR_NOT_SUPPORTED) + return -EOPNOTSUPP; + + return sbi_err_map_linux_errno(err); +} + +static int sse_sbi_event_func(struct sse_event *event, unsigned long func) +{ + struct sbiret ret; + u32 evt =3D event->evt_id; + struct sse_registered_event *reg_evt =3D sse_get_reg_evt(event); + + ret =3D sbi_ecall(SBI_EXT_SSE, func, evt, 0, 0, 0, 0, 0); + if (ret.error) { + pr_warn("Failed to execute func %lx, event %x, error %ld\n", + func, evt, ret.error); + return sse_err_map_linux_errno(ret.error); + } + + if (func =3D=3D SBI_SSE_EVENT_DISABLE) + reg_evt->is_enabled =3D false; + else if (func =3D=3D SBI_SSE_EVENT_ENABLE) + reg_evt->is_enabled =3D true; + + return 0; +} + +int sse_event_disable_local(struct sse_event *event) +{ + if (!sse_event_is_global(event->evt_id)) + lockdep_assert_preemption_disabled(); + + if (!sse_get_reg_evt(event)->is_enabled) + return 0; + + return sse_sbi_event_func(event, SBI_SSE_EVENT_DISABLE); +} +EXPORT_SYMBOL_GPL(sse_event_disable_local); + +int sse_event_enable_local(struct sse_event *event) +{ + struct sse_registered_event *reg_evt =3D sse_get_reg_evt(event); + int ret; + + if (!sse_event_is_global(event->evt_id)) + lockdep_assert_preemption_disabled(); + + rcu_read_lock(); + if (READ_ONCE(sse_shutting_down)) { + ret =3D -ESHUTDOWN; + goto out; + } + + if (!reg_evt->is_registered) { + ret =3D -EINVAL; + goto out; + } + + if (reg_evt->is_enabled) { + ret =3D 0; + goto out; + } + + ret =3D sse_sbi_event_func(event, SBI_SSE_EVENT_ENABLE); +out: + rcu_read_unlock(); + + return ret; +} +EXPORT_SYMBOL_GPL(sse_event_enable_local); + +static int sse_event_attr_get_no_lock(struct sse_registered_event *reg_evt, + unsigned long attr_id, unsigned long *val) +{ + struct sbiret sret; + u32 evt =3D reg_evt->event->evt_id; + phys_addr_t phys; + + phys =3D sse_event_get_attr_phys(reg_evt); + + sret =3D sbi_ecall(SBI_EXT_SSE, SBI_SSE_EVENT_ATTR_READ, evt, attr_id, 1, + (unsigned long)phys, 0, 0); + if (sret.error) { + pr_debug("Failed to get event %x attr %lx, error %ld\n", evt, + attr_id, sret.error); + return sse_err_map_linux_errno(sret.error); + } + + *val =3D reg_evt->attr; + + return 0; +} + +static int sse_event_attr_set_nolock(struct sse_registered_event *reg_evt, + unsigned long attr_id, unsigned long val) +{ + struct sbiret sret; + u32 evt =3D reg_evt->event->evt_id; + phys_addr_t phys; + + reg_evt->attr =3D val; + phys =3D sse_event_get_attr_phys(reg_evt); + + sret =3D sbi_ecall(SBI_EXT_SSE, SBI_SSE_EVENT_ATTR_WRITE, evt, attr_id, 1, + (unsigned long)phys, 0, 0); + if (sret.error) + pr_debug("Failed to set event %x attr %lx, error %ld\n", evt, + attr_id, sret.error); + + return sse_err_map_linux_errno(sret.error); +} + +static void sse_global_event_update_cpu(struct sse_event *event, + unsigned int cpu) +{ + struct sse_registered_event *reg_evt =3D event->global; + + event->cpu =3D cpu; + arch_sse_event_update_cpu(®_evt->arch, cpu); +} + +static int sse_event_set_target_cpu_nolock(struct sse_event *event, + unsigned int cpu) +{ + unsigned long hart_id, old_hart_id; + struct sse_registered_event *reg_evt =3D event->global; + u32 evt =3D event->evt_id; + unsigned int old_cpu; + bool was_enabled; + int ret; + + if (!sse_event_is_global(evt)) + return -EINVAL; + + if (cpu >=3D nr_cpu_ids || !cpu_online(cpu)) + return -EINVAL; + hart_id =3D cpuid_to_hartid_map(cpu); + old_cpu =3D event->cpu; + old_hart_id =3D cpuid_to_hartid_map(old_cpu); + + was_enabled =3D reg_evt->is_enabled; + if (was_enabled) { + ret =3D sse_event_disable_local(event); + if (ret) + return ret; + } + + ret =3D sse_event_attr_set_nolock(reg_evt, SBI_SSE_ATTR_PREFERRED_HART, + hart_id); + if (ret =3D=3D 0) + sse_global_event_update_cpu(event, cpu); + + if (was_enabled) { + int enable_ret; + + enable_ret =3D sse_event_enable_local(event); + if (enable_ret) { + int rollback_ret; + + /* + * The preferred hart was already changed. Restore both the + * firmware attribute and Linux's cached target before reporting + * the failed migration. + */ + rollback_ret =3D sse_event_attr_set_nolock(reg_evt, + SBI_SSE_ATTR_PREFERRED_HART, + old_hart_id); + if (!rollback_ret) { + sse_global_event_update_cpu(event, old_cpu); + rollback_ret =3D sse_event_enable_local(event); + } + if (!rollback_ret) + return enable_ret; + + pr_warn("Failed to restore global event %x to CPU %u: %d\n", + evt, old_cpu, rollback_ret); + event->enable_requested =3D false; + return enable_ret; + } + } + + return ret; +} + +int sse_event_set_target_cpu(struct sse_event *event, unsigned int cpu) +{ + int ret; + + if (cpu >=3D nr_cpu_ids) + return -EINVAL; + + scoped_guard(mutex, &sse_mutex) { + if (READ_ONCE(sse_shutting_down)) + return -ESHUTDOWN; + + scoped_guard(cpus_read_lock) { + if (!cpu_online(cpu)) + return -EINVAL; + + ret =3D sse_event_set_target_cpu_nolock(event, cpu); + } + } + + return ret; +} +EXPORT_SYMBOL_GPL(sse_event_set_target_cpu); + +static int sse_event_init_registered(unsigned int cpu, + struct sse_registered_event *reg_evt, + struct sse_event *event) +{ + reg_evt->event =3D event; + reg_evt->is_registered =3D false; + reg_evt->is_enabled =3D false; + + return arch_sse_init_event(®_evt->arch, event->evt_id, cpu); +} + +static void sse_event_free_registered(struct sse_registered_event *reg_evt) +{ + arch_sse_free_event(®_evt->arch); +} + +static int sse_event_alloc_global(struct sse_event *event) +{ + unsigned int cpu; + int err; + struct sse_registered_event *reg_evt; + + reg_evt =3D kzalloc_obj(*reg_evt, GFP_KERNEL); + if (!reg_evt) + return -ENOMEM; + + event->global =3D reg_evt; + cpu =3D cpumask_first(cpu_possible_mask); + if (cpu >=3D nr_cpu_ids) { + kfree(reg_evt); + return -ENODEV; + } + + err =3D sse_event_init_registered(cpu, reg_evt, event); + if (err) + kfree(reg_evt); + + return err; +} + +static int sse_event_alloc_local(struct sse_event *event) +{ + int err; + unsigned int cpu, err_cpu; + struct sse_registered_event *reg_evt; + struct sse_registered_event __percpu *reg_evts; + + reg_evts =3D alloc_percpu(struct sse_registered_event); + if (!reg_evts) + return -ENOMEM; + + event->local =3D reg_evts; + + for_each_possible_cpu(cpu) { + reg_evt =3D per_cpu_ptr(reg_evts, cpu); + err =3D sse_event_init_registered(cpu, reg_evt, event); + if (err) { + err_cpu =3D cpu; + goto err_free_per_cpu; + } + } + + return 0; + +err_free_per_cpu: + for_each_possible_cpu(cpu) { + if (cpu =3D=3D err_cpu) + break; + reg_evt =3D per_cpu_ptr(reg_evts, cpu); + sse_event_free_registered(reg_evt); + } + + free_percpu(reg_evts); + + return err; +} + +static struct sse_event *sse_event_alloc(u32 evt, u32 priority, + sse_event_handler_fn *handler, + void *arg) +{ + int err; + struct sse_event *event; + + event =3D kzalloc_obj(*event, GFP_KERNEL); + if (!event) + return ERR_PTR(-ENOMEM); + + event->evt_id =3D evt; + event->priority =3D priority; + event->handler_arg =3D arg; + RCU_INIT_POINTER(event->handler, handler); + + if (sse_event_is_global(evt)) + err =3D sse_event_alloc_global(event); + else + err =3D sse_event_alloc_local(event); + + if (err) { + kfree(event); + return ERR_PTR(err); + } + + return event; +} + +static int sse_sbi_register_event(struct sse_event *event, + struct sse_registered_event *reg_evt) +{ + int ret; + + if (reg_evt->is_registered) + return 0; + + ret =3D sse_event_attr_set_nolock(reg_evt, SBI_SSE_ATTR_PRIO, + event->priority); + if (ret) + return ret; + + ret =3D arch_sse_register_event(®_evt->arch); + if (!ret) + reg_evt->is_registered =3D true; + + return ret; +} + +static int sse_event_register_local(struct sse_event *event) +{ + int ret; + struct sse_registered_event *reg_evt; + + reg_evt =3D per_cpu_ptr(event->local, smp_processor_id()); + ret =3D sse_sbi_register_event(event, reg_evt); + if (ret) + pr_debug("Failed to register event %x: err %d\n", event->evt_id, + ret); + + return ret; +} + +static int sse_sbi_unregister_event(struct sse_event *event) +{ + struct sse_registered_event *reg_evt =3D sse_get_reg_evt(event); + int ret; + + if (!reg_evt->is_registered) + return 0; + + ret =3D sse_sbi_event_func(event, SBI_SSE_EVENT_UNREGISTER); + if (!ret) { + reg_evt->is_registered =3D false; + reg_evt->is_enabled =3D false; + } + + return ret; +} + +struct sse_per_cpu_evt { + struct sse_event *event; + unsigned long func; + atomic_t first_error; + atomic_t nonfallback_error; + cpumask_t changed; +}; + +static void sse_event_per_cpu_func(void *info) +{ + struct sse_per_cpu_evt *cpu_evt =3D info; + struct sse_registered_event *reg_evt; + bool changed; + int ret; + + reg_evt =3D sse_get_reg_evt(cpu_evt->event); + + if (cpu_evt->func =3D=3D SBI_SSE_EVENT_REGISTER) { + changed =3D !reg_evt->is_registered; + ret =3D sse_event_register_local(cpu_evt->event); + } else if (cpu_evt->func =3D=3D SBI_SSE_EVENT_UNREGISTER) { + changed =3D reg_evt->is_registered; + ret =3D sse_sbi_unregister_event(cpu_evt->event); + } else if (cpu_evt->func =3D=3D SBI_SSE_EVENT_ENABLE) { + changed =3D !reg_evt->is_enabled; + ret =3D sse_event_enable_local(cpu_evt->event); + } else if (cpu_evt->func =3D=3D SBI_SSE_EVENT_DISABLE) { + changed =3D reg_evt->is_enabled; + ret =3D sse_event_disable_local(cpu_evt->event); + } else { + changed =3D false; + ret =3D -EINVAL; + } + + if (ret) { + atomic_cmpxchg(&cpu_evt->first_error, 0, ret); + if (ret !=3D -EOPNOTSUPP) + atomic_cmpxchg(&cpu_evt->nonfallback_error, 0, ret); + } else if (changed) { + cpumask_set_cpu(smp_processor_id(), &cpu_evt->changed); + } +} + +static bool sse_event_is_registered(struct sse_event *event) +{ + unsigned int cpu; + + if (sse_event_is_global(event->evt_id)) + return event->global->is_registered; + + for_each_possible_cpu(cpu) { + if (per_cpu_ptr(event->local, cpu)->is_registered) + return true; + } + + return false; +} + +static bool sse_event_is_enabled(struct sse_event *event) +{ + unsigned int cpu; + + if (sse_event_is_global(event->evt_id)) + return event->global->is_enabled; + + for_each_possible_cpu(cpu) { + if (per_cpu_ptr(event->local, cpu)->is_enabled) + return true; + } + + return false; +} + +static void sse_event_free(struct sse_event *event) +{ + unsigned int cpu; + struct sse_registered_event *reg_evt; + + if (WARN_ON_ONCE(sse_event_is_registered(event))) + return; + + if (sse_event_is_global(event->evt_id)) { + sse_event_free_registered(event->global); + kfree(event->global); + } else { + for_each_possible_cpu(cpu) { + reg_evt =3D per_cpu_ptr(event->local, cpu); + sse_event_free_registered(reg_evt); + } + free_percpu(event->local); + } + + kfree(event); +} + +static struct sse_event *sse_register_failed(struct sse_event *event, int = ret) +{ + /* + * Keep failed rollback state visible to CPU hotplug and shutdown. The + * core-owned handler also makes the retained registration independent of + * the client whose registration request failed. + */ + if (sse_event_is_registered(event)) { + event->cleanup_pending =3D true; + rcu_assign_pointer(event->handler, sse_cleanup_event_handler); + synchronize_rcu(); + list_add(&event->list, &events); + pr_err("Event %x remains registered after rollback; cleanup retained\n", + event->evt_id); + ret =3D -EUCLEAN; + } else { + sse_event_free(event); + } + + return ERR_PTR(ret); +} + +void sse_event_cleanup(struct sse_event *event) +{ + guard(mutex)(&sse_mutex); + guard(cpus_read_lock)(); + + if (event->cleanup_pending) + return; + + /* + * Firmware may still enter the old callback after disable or unregister + * fails. Publish a core-owned callback, then wait before the client frees + * its callback data. + */ + event->cleanup_pending =3D true; + rcu_assign_pointer(event->handler, sse_cleanup_event_handler); + synchronize_rcu(); +} +EXPORT_SYMBOL_GPL(sse_event_cleanup); + +static void sse_release_cleanup_event(struct sse_event *event) +{ + if (!event->cleanup_pending || sse_event_is_registered(event)) + return; + + list_del(&event->list); + sse_event_free(event); +} + +static int sse_event_setup_all_cpus(struct sse_event *event, + unsigned long func, + unsigned long rollback_func) +{ + struct sse_per_cpu_evt cpu_evt; + int rollback_ret; + int ret; + + cpu_evt.event =3D event; + atomic_set(&cpu_evt.first_error, 0); + atomic_set(&cpu_evt.nonfallback_error, 0); + cpumask_clear(&cpu_evt.changed); + cpu_evt.func =3D func; + on_each_cpu(sse_event_per_cpu_func, &cpu_evt, 1); + /* IRQ fallback is safe only if every failing CPU reports unsupported. */ + ret =3D atomic_read(&cpu_evt.nonfallback_error); + if (!ret) + ret =3D atomic_read(&cpu_evt.first_error); + /* + * A previous attempt may already have changed some CPUs. Roll back only + * instances changed by this invocation. + */ + if (ret) { + cpu_evt.func =3D rollback_func; + atomic_set(&cpu_evt.first_error, 0); + atomic_set(&cpu_evt.nonfallback_error, 0); + on_each_cpu_mask(&cpu_evt.changed, sse_event_per_cpu_func, &cpu_evt, 1); + + rollback_ret =3D atomic_read(&cpu_evt.nonfallback_error); + if (!rollback_ret) + rollback_ret =3D atomic_read(&cpu_evt.first_error); + + /* A rollback failure leaves the firmware state uncertain. */ + return rollback_ret ?: ret; + } + + return 0; +} + +static int sse_event_teardown_all_cpus(struct sse_event *event, + unsigned long func) +{ + struct sse_per_cpu_evt cpu_evt; + + cpu_evt.event =3D event; + atomic_set(&cpu_evt.first_error, 0); + atomic_set(&cpu_evt.nonfallback_error, 0); + cpumask_clear(&cpu_evt.changed); + cpu_evt.func =3D func; + on_each_cpu(sse_event_per_cpu_func, &cpu_evt, 1); + + return atomic_read(&cpu_evt.first_error); +} + +int sse_event_enable(struct sse_event *event) +{ + int ret =3D 0; + + scoped_guard(mutex, &sse_mutex) { + if (READ_ONCE(sse_shutting_down)) + return -ESHUTDOWN; + + scoped_guard(cpus_read_lock) { + if (sse_event_is_global(event->evt_id)) { + ret =3D sse_event_enable_local(event); + } else { + ret =3D sse_event_setup_all_cpus(event, + SBI_SSE_EVENT_ENABLE, + SBI_SSE_EVENT_DISABLE); + } + event->enable_requested =3D !ret; + } + } + + return ret; +} +EXPORT_SYMBOL_GPL(sse_event_enable); + +static int sse_events_mask(void) +{ + struct sbiret ret; + + ret =3D sbi_ecall(SBI_EXT_SSE, SBI_SSE_HART_MASK, 0, 0, 0, 0, 0, 0); + if (ret.error =3D=3D SBI_ERR_ALREADY_STOPPED) + return 0; + + return sse_err_map_linux_errno(ret.error); +} + +static int sse_events_unmask(void) +{ + struct sbiret ret; + + ret =3D sbi_ecall(SBI_EXT_SSE, SBI_SSE_HART_UNMASK, 0, 0, 0, 0, 0, 0); + if (ret.error =3D=3D SBI_ERR_ALREADY_STARTED) + return 0; + + return sse_err_map_linux_errno(ret.error); +} + +static int sse_event_disable_nolock(struct sse_event *event) +{ + if (sse_event_is_global(event->evt_id)) + return sse_event_disable_local(event); + + return sse_event_teardown_all_cpus(event, SBI_SSE_EVENT_DISABLE); +} + +int sse_event_disable(struct sse_event *event) +{ + int ret =3D 0; + + scoped_guard(mutex, &sse_mutex) { + if (READ_ONCE(sse_shutting_down)) + return -ESHUTDOWN; + + scoped_guard(cpus_read_lock) { + if (!event->enable_requested && !sse_event_is_enabled(event)) + return 0; + + event->enable_requested =3D false; + ret =3D sse_event_disable_nolock(event); + if (!ret && sse_event_is_enabled(event)) + ret =3D -EIO; + } + } + + return ret; +} +EXPORT_SYMBOL_GPL(sse_event_disable); + +struct sse_event *sse_event_register(u32 evt, u32 priority, + sse_event_handler_fn *handler, void *arg) +{ + struct sse_event *event; + int cpu; + int ret =3D 0; + + if (sse_fw_state_retained) + return ERR_PTR(-EUCLEAN); + if (!sse_available) + return ERR_PTR(-EOPNOTSUPP); + + guard(mutex)(&sse_mutex); + if (READ_ONCE(sse_shutting_down)) + return ERR_PTR(-ESHUTDOWN); + + guard(cpus_read_lock)(); + + if (sse_event_get(evt)) + return ERR_PTR(-EEXIST); + + event =3D sse_event_alloc(evt, priority, handler, arg); + if (IS_ERR(event)) + return event; + + if (sse_event_is_global(evt)) { + unsigned long preferred_hart; + + ret =3D sse_event_attr_get_no_lock(event->global, + SBI_SSE_ATTR_PREFERRED_HART, + &preferred_hart); + if (ret) + return sse_register_failed(event, ret); + + cpu =3D riscv_hartid_to_cpuid(preferred_hart); + if (cpu < 0 || !cpu_online(cpu)) { + cpu =3D cpumask_first(cpu_online_mask); + if (cpu >=3D nr_cpu_ids) + return sse_register_failed(event, -ENODEV); + + ret =3D sse_event_set_target_cpu_nolock(event, cpu); + if (ret) + return sse_register_failed(event, ret); + } else { + sse_global_event_update_cpu(event, cpu); + } + + ret =3D sse_sbi_register_event(event, event->global); + if (ret) + return sse_register_failed(event, ret); + } else { + ret =3D sse_event_setup_all_cpus(event, SBI_SSE_EVENT_REGISTER, + SBI_SSE_EVENT_UNREGISTER); + if (ret) + return sse_register_failed(event, ret); + } + + list_add(&event->list, &events); + + return event; +} +EXPORT_SYMBOL_GPL(sse_event_register); + +static int sse_event_unregister_nolock(struct sse_event *event) +{ + if (sse_event_is_global(event->evt_id)) + return sse_sbi_unregister_event(event); + + return sse_event_teardown_all_cpus(event, SBI_SSE_EVENT_UNREGISTER); +} + +int sse_event_unregister(struct sse_event *event) +{ + int ret =3D 0; + + scoped_guard(mutex, &sse_mutex) { + if (READ_ONCE(sse_shutting_down)) + return -ESHUTDOWN; + + scoped_guard(cpus_read_lock) { + ret =3D sse_event_unregister_nolock(event); + if (ret) + return ret; + if (sse_event_is_registered(event)) + return -EBUSY; + + list_del(&event->list); + + sse_event_free(event); + } + } + + return ret; +} +EXPORT_SYMBOL_GPL(sse_event_unregister); + +static int sse_teardown_event(struct sse_event *event, unsigned int cpu); + +static int sse_cpu_online(unsigned int cpu) +{ + int ret, rollback_ret; + struct sse_event *event, *tmp; + struct sse_registered_event *reg_evt; + + arch_sse_init_cpu(); + + list_for_each_entry_safe(event, tmp, &events, list) { + if (sse_event_is_global(event->evt_id)) + continue; + if (event->cleanup_pending) { + scoped_guard(preempt) + ret =3D sse_teardown_event(event, cpu); + if (ret) + goto rollback; + sse_release_cleanup_event(event); + continue; + } + + /* + * The local SBI helpers act on the current hart and assert that + * preemption is disabled. Unlike the client-facing paths, which + * reach them through an IPI, the CPUHP thread runs preemptible. + */ + scoped_guard(preempt) { + ret =3D sse_event_register_local(event); + if (!ret) + ret =3D event->enable_requested ? + sse_event_enable_local(event) : + sse_event_disable_local(event); + } + if (ret) + goto rollback; + } + + /* Only unmask after every cached per-CPU state is reconstructed. */ + ret =3D sse_events_unmask(); + if (!ret) + return 0; + +rollback: + /* A failed startup callback is not followed by this state's teardown. */ + list_for_each_entry_safe(event, tmp, &events, list) { + if (sse_event_is_global(event->evt_id)) + continue; + + scoped_guard(preempt) { + reg_evt =3D sse_get_reg_evt(event); + rollback_ret =3D reg_evt->is_enabled ? + sse_event_disable_local(event) : 0; + if (!rollback_ret) + rollback_ret =3D reg_evt->is_registered ? + sse_sbi_unregister_event(event) : 0; + } + if (rollback_ret) { + pr_warn("Failed to roll back event %x on CPU %u: %d\n", + event->evt_id, cpu, rollback_ret); + atomic_set(&sse_teardown_failed, 1); + continue; + } + sse_release_cleanup_event(event); + } + + return ret; +} + +static int sse_teardown_event(struct sse_event *event, unsigned int cpu) +{ + struct sse_registered_event *reg_evt =3D sse_get_reg_evt(event); + int ret; + + if (reg_evt->is_enabled) { + ret =3D sse_event_disable_local(event); + if (ret) { + pr_warn("Failed to disable event %x on CPU %u: %d\n", + event->evt_id, cpu, ret); + return ret; + } + } + + ret =3D sse_sbi_unregister_event(event); + if (ret) + pr_warn("Failed to unregister event %x on CPU %u: %d\n", + event->evt_id, cpu, ret); + + return ret; +} + +static int sse_restore_local_events(unsigned int cpu) +{ + struct sse_event *event; + int first_error =3D 0; + int ret; + + list_for_each_entry(event, &events, list) { + if (sse_event_is_global(event->evt_id) || event->cleanup_pending) + continue; + + scoped_guard(preempt) { + ret =3D sse_event_register_local(event); + if (!ret) + ret =3D event->enable_requested ? + sse_event_enable_local(event) : + sse_event_disable_local(event); + } + if (ret) { + pr_warn("Failed to restore event %x on CPU %u: %d\n", + event->evt_id, cpu, ret); + if (!first_error) + first_error =3D ret; + } + } + + return first_error; +} + +static int sse_cpu_teardown(unsigned int cpu) +{ + /* Only a regular CPU-offline callback may abort the CPUHP operation. */ + bool regular_offline =3D READ_ONCE(sse_available) && + !READ_ONCE(sse_shutting_down); + unsigned int next_cpu; + struct sse_event *event, *tmp; + int first_error =3D 0; + int ret; + + /* Do not dismantle CPU state while firmware can still deliver SSE. */ + ret =3D sse_events_mask(); + if (ret) { + pr_warn("Failed to mask SSE on CPU %u during teardown: %d\n", + cpu, ret); + if (READ_ONCE(sse_shutting_down)) + atomic_set(&sse_teardown_failed, 1); + /* CPUHP installation rollback and state removal cannot fail. */ + return regular_offline ? ret : 0; + } + + list_for_each_entry_safe(event, tmp, &events, list) { + if (sse_event_is_global(event->evt_id)) + continue; + + scoped_guard(preempt) + ret =3D sse_teardown_event(event, cpu); + if (ret && !first_error) + first_error =3D ret; + sse_release_cleanup_event(event); + if (ret && regular_offline) + goto restore_cpu; + } + + list_for_each_entry_safe(event, tmp, &events, list) { + if (!sse_event_is_global(event->evt_id) || event->cpu !=3D cpu) + continue; + + /* + * cpuhp_remove_state() invokes teardown while every CPU remains + * online. Do not migrate to a CPU whose callback may have run. + */ + if (READ_ONCE(sse_shutting_down) || event->cleanup_pending) { + ret =3D sse_teardown_event(event, cpu); + } else { + next_cpu =3D cpumask_any_but(cpu_online_mask, cpu); + if (next_cpu >=3D nr_cpu_ids) { + ret =3D sse_teardown_event(event, cpu); + } else { + ret =3D sse_event_set_target_cpu_nolock(event, next_cpu); + if (ret) + pr_warn("Failed to migrate global event %x from CPU %u: %d\n", + event->evt_id, cpu, ret); + } + } + + if (ret && !first_error) + first_error =3D ret; + sse_release_cleanup_event(event); + if (ret && regular_offline) + goto restore_cpu; + } + + if (first_error) { + /* An offline CPU is not revisited when this CPUHP state is removed. */ + atomic_set(&sse_teardown_failed, 1); + } + + return 0; + +restore_cpu: + /* + * CPUHP leaves this CPU online when teardown returns an error. Restore + * every client-owned local event before making SSE delivery visible agai= n. + */ + ret =3D sse_restore_local_events(cpu); + if (!ret) + ret =3D sse_events_unmask(); + if (ret) { + atomic_set(&sse_teardown_failed, 1); + pr_warn("Failed to restore SSE after aborting CPU %u offline: %d\n", + cpu, ret); + } else { + pr_warn("Aborted CPU %u offline after SSE teardown failed: %d\n", + cpu, first_error); + } + + return first_error; +} + +static int sse_pm_notifier(struct notifier_block *nb, unsigned long action, + void *data) +{ + int ret; + + WARN_ON_ONCE(preemptible()); + + switch (action) { + case CPU_PM_ENTER: + ret =3D sse_events_mask(); + break; + case CPU_PM_EXIT: + case CPU_PM_ENTER_FAILED: + if (READ_ONCE(sse_shutting_down)) + return NOTIFY_OK; + ret =3D sse_events_unmask(); + break; + default: + return NOTIFY_DONE; + } + + if (ret) + return notifier_from_errno(ret); + + return NOTIFY_OK; +} + +static struct notifier_block sse_pm_nb =3D { + .notifier_call =3D sse_pm_notifier, +}; + +static int sse_panic_notifier(struct notifier_block *nb, unsigned long act= ion, + void *data) +{ + riscv_sse_mask_current_hart(); + + return NOTIFY_OK; +} + +static struct notifier_block sse_panic_nb =3D { + .notifier_call =3D sse_panic_notifier, + .priority =3D INT_MAX, +}; + +/* + * Mask all CPUs and unregister all events on reboot or kexec. + */ +static int sse_reboot_notifier(struct notifier_block *nb, unsigned long ac= tion, + void *data) +{ + int ret; + + scoped_guard(mutex, &sse_mutex) { + if (!sse_shutting_down) { + WRITE_ONCE(sse_shutting_down, true); + ret =3D cpu_pm_unregister_notifier(&sse_pm_nb); + if (ret) { + pr_warn("Failed to unregister CPU PM notifier: %d\n", + ret); + atomic_set(&sse_teardown_failed, 1); + } + /* Drain CPU PM callbacks and client enables before teardown. */ + synchronize_rcu(); + cpuhp_remove_state(CPUHP_AP_RISCV_SSE_ONLINE); + } + } + + /* Normal kexec preserves firmware state but discards old kernel memory. = */ + if (kexec_in_progress && atomic_read(&sse_teardown_failed)) + panic("SSE teardown failed; refusing unsafe kexec"); + + return NOTIFY_OK; +} + +static struct notifier_block sse_reboot_nb =3D { + .notifier_call =3D sse_reboot_notifier, +}; + +static int __init sse_init(void) +{ + int ret; + + /* + * A kdump kernel cannot identify registrations left by the crashed + * kernel. Keep them masked by not initializing SSE again. + */ + if (is_kdump_kernel() && riscv_sse_available()) { + sse_fw_state_retained =3D true; + pr_info("SSE remains disabled in the crash kernel\n"); + return -EOPNOTSUPP; + } + + if (sbi_probe_extension(SBI_EXT_SSE) <=3D 0) { + pr_info("Missing SBI SSE extension\n"); + return -EOPNOTSUPP; + } + pr_info("SBI SSE extension detected\n"); + + ret =3D cpu_pm_register_notifier(&sse_pm_nb); + if (ret) { + pr_warn("Failed to register CPU PM notifier...\n"); + return ret; + } + + ret =3D register_reboot_notifier(&sse_reboot_nb); + if (ret) { + pr_warn("Failed to register reboot notifier...\n"); + goto remove_cpupm; + } + + ret =3D atomic_notifier_chain_register(&panic_notifier_list, &sse_panic_n= b); + if (ret) { + pr_warn("Failed to register panic notifier...\n"); + goto remove_reboot; + } + + /* Tear down perf events before dismantling their SSE delivery path. */ + ret =3D cpuhp_setup_state(CPUHP_AP_RISCV_SSE_ONLINE, "riscv/sse:online", + sse_cpu_online, sse_cpu_teardown); + if (ret < 0) + goto remove_panic; + + sse_available =3D true; + + return 0; + +remove_panic: + atomic_notifier_chain_unregister(&panic_notifier_list, &sse_panic_nb); + +remove_reboot: + unregister_reboot_notifier(&sse_reboot_nb); + +remove_cpupm: + cpu_pm_unregister_notifier(&sse_pm_nb); + + return ret; +} +arch_initcall(sse_init); diff --git a/include/linux/cpuhotplug.h b/include/linux/cpuhotplug.h index feb32949aeea..d13af475508d 100644 --- a/include/linux/cpuhotplug.h +++ b/include/linux/cpuhotplug.h @@ -200,6 +200,7 @@ enum cpuhp_state { CPUHP_AP_ARM_MVEBU_SYNC_CLOCKS, CPUHP_AP_ARM_CORESIGHT_ONLINE, CPUHP_AP_X86_INTEL_EPB_ONLINE, + CPUHP_AP_RISCV_SSE_ONLINE, CPUHP_AP_PERF_ONLINE, CPUHP_AP_PERF_X86_ONLINE, CPUHP_AP_PERF_X86_UNCORE_ONLINE, diff --git a/include/linux/riscv_sbi_sse.h b/include/linux/riscv_sbi_sse.h new file mode 100644 index 000000000000..774a782e556d --- /dev/null +++ b/include/linux/riscv_sbi_sse.h @@ -0,0 +1,89 @@ +/* SPDX-License-Identifier: GPL-2.0-or-later */ +/* + * Copyright (C) 2025 Rivos Inc. + */ + +#ifndef __LINUX_RISCV_SBI_SSE_H +#define __LINUX_RISCV_SBI_SSE_H + +#include +#include +#include +#include + +struct sse_event; +struct pt_regs; + +typedef int (sse_event_handler_fn)(u32 event_num, void *arg, + struct pt_regs *regs); + +#ifdef CONFIG_RISCV_SBI_SSE + +/* + * The callback and its argument must remain valid until unregister succee= ds. + * The callback runs in NMI context and must not sleep. + * regs is NULL if firmware cannot provide a complete interrupted context. + */ +struct sse_event *sse_event_register(u32 event_num, u32 priority, + sse_event_handler_fn *handler, void *arg); + +int sse_event_unregister(struct sse_event *evt); + +/* + * Transfer a retained event to the SSE core for deferred cleanup. The cal= ler + * must not access the event or its callback data after this function retu= rns. + */ +void sse_event_cleanup(struct sse_event *evt); + +int sse_event_set_target_cpu(struct sse_event *sse_evt, unsigned int cpu); + +int sse_event_enable(struct sse_event *sse_evt); + +int sse_event_disable(struct sse_event *sse_evt); + +/* Local events require the caller to remain on the current CPU. */ +int sse_event_enable_local(struct sse_event *sse_evt); +int sse_event_disable_local(struct sse_event *sse_evt); + +#else +static inline struct sse_event *sse_event_register(u32 event_num, u32 prio= rity, + sse_event_handler_fn *handler, + void *arg) +{ + return ERR_PTR(-EOPNOTSUPP); +} + +static inline int sse_event_unregister(struct sse_event *evt) +{ + return -EOPNOTSUPP; +} + +static inline void sse_event_cleanup(struct sse_event *evt) { } + +static inline int sse_event_set_target_cpu(struct sse_event *sse_evt, + unsigned int cpu) +{ + return -EOPNOTSUPP; +} + +static inline int sse_event_enable(struct sse_event *sse_evt) +{ + return -EOPNOTSUPP; +} + +static inline int sse_event_disable(struct sse_event *sse_evt) +{ + return -EOPNOTSUPP; +} + +static inline int sse_event_enable_local(struct sse_event *sse_evt) +{ + return -EOPNOTSUPP; +} + +static inline int sse_event_disable_local(struct sse_event *sse_evt) +{ + return -EOPNOTSUPP; +} +#endif +#endif /* __LINUX_RISCV_SBI_SSE_H */ --=20 2.50.1 (Apple Git-155) From nobody Fri Sep 25 21:02:40 2026 Received: from mail-pj2-f12.google.com (mail-pj2-f12.google.com [74.125.227.140]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7388C48F846 for ; Mon, 21 Sep 2026 11:17:38 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.227.140 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789989461; cv=none; b=ouysK8EswLwYhC4Qx2H/SOdoAlcG1FUyn/D8sYyA88fsQoMHA/JEO7QZJg48dqO2JtNyiYwQ9yzi659rsC7ezD/W90tsg432dZXeVMC9C6oG1BbqIBp+pih3uG4Y/aPvNkP3gZoXTECq2cp8G+KrI0pq/h4ds8fPoNmGPnJvw9Q= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789989461; c=relaxed/simple; bh=hOV/NyTnPxdCBe6XO/1GKFUexB1rAl0f0lQQCm4ePvE=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Lo6iiRqes/wLmWCaFbFsfwGnMSNCzg/5h/1B1k8R1vNRfcMbY34JRCmmibB1W4fLBO481jxKaAJkkiePmSp8vhhQhsyHp8Z4FWKUY/Cx5XMTrt9XEaE17gfxILEj/e9x31npiiLzo376Us4cwLMQ9V/7QNp3LcJ9bEiADj/3jiQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com; spf=pass smtp.mailfrom=bytedance.com; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b=YtarrOsx; arc=none smtp.client-ip=74.125.227.140 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=bytedance.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b="YtarrOsx" Received: by mail-pj2-f12.google.com with SMTP id d9443c01a7336-2d90ba1d807so34014835ad.3 for ; Mon, 21 Sep 2026 04:17:38 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=bytedance.com; s=google; t=1789989458; x=1790594258; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=s0eHZBxWOds1UYdEnuOhccIJeDRWzv84kT8tp0myDIA=; b=YtarrOsx3K+S90zlvSL7PpvSj9CEc5Sqm4snKBOiCNTFS5AzbPaGp4ZP/jVS8pd87s UHXLz77rt0KEsrZpj54kYRm0kzLfZcmzJOEteryslGkPBaW+nmuUPYHZlCVBDsFX3DHm jzwwDEW1NLcDUCXxHm2w+2xynUZUP9eqZkmnSHacsOT9gzr00cjrOFR8NbqWnQH6Me45 nMwzMLuqKj/KwAjRCmAuT2PLY3CxRN/1WaQe27ocl5kzSx39e4IeEnBdWnJTmRYHpUYR VIZOKD5EpWmbXzunlK/y/jiS2zHVqy9irr2dlN8vw+CANSQ2NS446mHolOqQ+0ANaB9R +tFw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789989458; x=1790594258; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=s0eHZBxWOds1UYdEnuOhccIJeDRWzv84kT8tp0myDIA=; b=cdJTzPYsUyGr5bwARs6t1/jygMMwgyqkAxub4joH9wgyaJ5MXmIMGiPNM158T8hzh6 w4AQMuhEoLOis1rBt3DSXRON1W2sSVz5SqdQiAI2NlqMVulvSwiu9auMkE4d9E+kPzgP FW/T+F78wY8gOCj5qfH9+pAaVdKb1pBt1DZK8abenOBcOJOk/dEYT1DidNSvs5G/LpBH EPo2zKhLC0XsrzKEWdA6UCxTY3ATd/tTCk6m1+1Wp2h9WlXgioAbqhRfPKB6UkKcUB7Q 7Yb1Z0vKgDvboI/bXE/eUUmL9aDPZPOp7kDQup9HPyypPsLvivGN1dXnudKh5w1ONNG8 ND/g== X-Forwarded-Encrypted: i=1; AKwUvBy0Vj4UVEWWGR5Y7+GUa1gtN+Nu8dMnJ2Xyonrxl8m6pxQS+ZcTqW5j81mSANgIvE35lWmf4edSbZldUh8=@vger.kernel.org X-Gm-Message-State: AFuF++n9BQTMIbYBpKnM2yO38u/XMI7Vhg1tvNQBXumvOhcjSxqlacEg 9Unz5Wwzwx3scw7BpAk3voedJk3iVkFMYIDw3fNoqg5zbkQiYoA3jXqdNRZ3aazt84w= X-Gm-Gg: AYBFou1D966BDjc5WJuWDvlmQR7qYnhk1NPq2AMVPFAoDov/yGmKaIRd/Tr40uo5o49 UeDkUY+cXAXFAMljBjaV1iiGxIqW+CTPNULVaG7wuUShxAQqMdb9KBdYx1sef61ivrhU0UH3UOh XqHJokBsyYoUbfwM1B8KhGUciedwkdcFzeuW/UZ6/GSqLp0yZhMIWuyHO1MkoKUZmIE7qwI6puV zk0in/btS1gh1r5DdA6bPQ1duMJEUuNaO6Di6tyK9mYgWKy3dd/0B86UuLrr761TddkWlhmyXg5 rcfxOhFfSb4nD1tex66ws04I/qXdJKmTo1d5JaG9DA06AIVqQ5C/WgcvIUnfmzEkvsDvieXG3eE hc4WeeLn+m132ro0/LWjwkcSKyo8Jp9K4lYTq+cF0rQ+RJ0wq+ofvBF1A6O+/AEwOHflHQV6HTd OZFZeGl7i++YIqceCn34A/CM5DzAYVMCa5n4RKYrHT8L6fgUqRkDLLRginHzb6zoph5z7EHG682 WqRXHIkDidLTsjq4dqPl+sKGQvGsXyVmt6p X-Received: by 2002:a17:903:2349:b0:2dd:c100:a5de with SMTP id d9443c01a7336-2ddc100a706mr90374885ad.50.1789989457717; Mon, 21 Sep 2026 04:17:37 -0700 (PDT) Received: from FJ7FR2JRQ3.bytedance.net ([240e:694:e20:401::8]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2ddc17ba651sm31831125ad.53.2026.09.21.04.17.15 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Mon, 21 Sep 2026 04:17:37 -0700 (PDT) From: Zhanpeng Zhang To: Paul Walmsley , Palmer Dabbelt , Albert Ou , Alexandre Ghiti , Himanshu Chauhan , Conor Dooley , Anup Patel Cc: =?UTF-8?q?Cl=C3=A9ment=20L=C3=A9ger?= , Yunhui Cui , Atish Patra , Peter Zijlstra , Ingo Molnar , Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Alexander Shishkin , Jiri Olsa , Ian Rogers , Adrian Hunter , James Clark , Will Deacon , Thomas Gleixner , Jonathan Corbet , Randy Dunlap , Shuah Khan , Shuah Khan , Yuanzhu , Yicong Yang , Susheng Yang , linux-riscv@lists.infradead.org, linux-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-arm-kernel@lists.infradead.org, Zhanpeng Zhang Subject: [PATCH v10 RESEND 5/9] riscv: mm: avoid enabling interrupts for nofault page faults Date: Mon, 21 Sep 2026 19:15:02 +0800 Message-ID: <4bc18f96053a7859a3ff33747479b05a9e42a706.1789974241.git.zhangzhanpeng.jasper@bytedance.com> X-Mailer: git-send-email 2.50.1 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" RISC-V enables interrupts in handle_page_fault() before checking whether the fault occurred with fault handling disabled. A nofault access from an atomic context can therefore run tracepoints and open an interrupt window before reaching the exception-table fixup. Handle an exception-table entry before entering the generic fault path when fault handling is disabled. Also keep interrupts disabled until such a fault has been resolved. This makes RISC-V consistent with the expectation that an in-atomic nofault access does not enter the normal fault-handling path. It also removes one source of re-entry when perf sampling is delivered through an SBI Supervisor Software Event (SSE). Signed-off-by: Zhanpeng Zhang --- arch/riscv/mm/fault.c | 11 +++++++++-- 1 file changed, 9 insertions(+), 2 deletions(-) diff --git a/arch/riscv/mm/fault.c b/arch/riscv/mm/fault.c index 04ed6f8acae4..520495420462 100644 --- a/arch/riscv/mm/fault.c +++ b/arch/riscv/mm/fault.c @@ -294,6 +294,13 @@ void handle_page_fault(struct pt_regs *regs) if (kprobe_page_fault(regs, cause)) return; =20 + /* + * Nofault accesses must be resolved through the exception table before + * entering the generic fault path or enabling interrupts. + */ + if (unlikely(faulthandler_disabled()) && fixup_exception(regs)) + return; + if (user_mode(regs)) trace_page_fault_user(addr, regs, cause); else @@ -314,8 +321,8 @@ void handle_page_fault(struct pt_regs *regs) return; } =20 - /* Enable interrupts if they were enabled in the parent context. */ - if (!regs_irqs_disabled(regs)) + /* Do not open an interrupt window before a nofault fixup completes. */ + if (!regs_irqs_disabled(regs) && !faulthandler_disabled()) local_irq_enable(); =20 /* --=20 2.50.1 (Apple Git-155) From nobody Fri Sep 25 21:02:40 2026 Received: from mail-pj2-f12.google.com (mail-pj2-f12.google.com [74.125.227.140]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B199548A2D1 for ; Mon, 21 Sep 2026 11:18:00 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.227.140 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789989483; cv=none; b=JT28phq+aqnksE8eh7v87MWfmaYQlYb0lOe1L9610m2v8v7CYy0vo2q7euVPiKyvsCIaipGYmeRP7OmReXvv+qR9LL2XwODX4Ie3phakz0ZAlXqjCHI/Puk42ZiE/irFwGBqfFoJaU4gLiYZtkJqe+TxqaDu3sjIAl/FWRxGDzs= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789989483; c=relaxed/simple; bh=l9cbhsR6oIcRYl+I2g1jCIa/WhsXA/dARjMt3kDigl0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=BXA0NvaGtjd9xGQ+fKUYPh7KBpViq5LvDCoJcmSa+2YbT5i0b6H5ycWhGOH9KI/UhjMAtSVCEBmeFO0uyVfoEAYp8t8eiZvdQXle8ANETs7UuxWrrx8f6YgrfFQwVtemey5DfC0FAZsaUEL/Hm9zSSCZp37xM0pngeB9GhKSHg0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com; spf=pass smtp.mailfrom=bytedance.com; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b=jenZNvHq; arc=none smtp.client-ip=74.125.227.140 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=bytedance.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b="jenZNvHq" Received: by mail-pj2-f12.google.com with SMTP id d9443c01a7336-2d747ed1368so33830545ad.1 for ; Mon, 21 Sep 2026 04:18:00 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=bytedance.com; s=google; t=1789989480; x=1790594280; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=jS32Dx5IynGZ/Y/2aGGAGztAdcoBdPlUNApcd4RXZUY=; b=jenZNvHqVOq7BaieRqOS48emO4JyrhHQYVA7mvNKIVsRb9xdgaFAuGWgOJ0jtPM/fz CDWx2AzGibPiBkf061LwHxisZe3a5er5mhYD5O/SkFdRSrTQqb+AlSC1MeXgsrudt8tZ l6WDHrPlKNKPQs2rOrrbq26ULBR5NrniVkaLoyCqASd3DcQlJnmQzTJUtK1irRf9uykt yFJhseZWPbkOgJEc9Wk+LIi18irvlwcoDUgo+bYys+cXGDemJseATnaJzPjN5slYCA4b zG6l/cBs/OoyR3mN0q3LhlPWvP8uN29mfufGnxdQQYjEM3qnt1eaKONzCK26SMzNZ6IR xHkw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789989480; x=1790594280; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=jS32Dx5IynGZ/Y/2aGGAGztAdcoBdPlUNApcd4RXZUY=; b=gcCGJekeKPMToaO+gRJNV3B5uBBtmfg5XJI73oc13UQtoEdSVNb2iRCKHqTutg0Cqx 2gqibjl5u/7a6cLtAcn8Myg7gOzrNbu8jDEbETT05YOKJt/bBvusoyzuHHkGRYqBMxXQ ZX+mGFNfVMAhzhWw6BmUwh0N4/V3YRTLA4BxpKF7mzZ0iCMVd9By75YTROTfEEGp6dAS i3nPKtdv0R865DKVIN0nae83uNTr9ohqPYgE7unZ6n6tD1V78r3+Cd5H7lQx7mod/jrF XL8uT6K0LTyHQQylNeRPUgq+p8w13h9bNVL22L7oUeTo2TG1tLrMIKXXaI8D0IinZHXg MF8g== X-Forwarded-Encrypted: i=1; AKwUvBxWlJAumlE/VFLDP3GFbwYU+pVURjd/x+CefLxk04GwgMXEnVzSB3QhSddJs1wRStQM+aPX+x1ttHS1eOw=@vger.kernel.org X-Gm-Message-State: AFuF++kORNqhM0jumuqQSBYWGqjIFMSNWUF8yMoEIEiiNDD6Ab7TX/MG BR5XyQkfe+/mFUEqduVZwsi0r4iWUGRMbxITxeOmsp6AuFLi4F1Bo2EZMWL2HIZzFwk= X-Gm-Gg: AYBFou14Zs+cAd2ge91c78t2h7CBW/TLiWb7NRyppSqLge9al7Td7PRZ9RnYLl7AkuH SNH2jhbg8XxKotjQDE9Bsk1zg16c1j2lwBbMDwnagreGyps6wYI6R9XaNeacL8aJxvMeOlM8R3O nVj9Psgg5W081HzQ6RP8Yv7ZYw9a//KHuc8xSkuYBZK/2RhxolEeT9EStw9Ry2adag4r1xQ6Zvz J3J65ErOIEH5+zDz83FLy4wZdd0MtAhS1aMuPASTKPhau81wyhpHva84JIWB9nGlNA70et0m44G 560m8ZTPd5Dmv2OY5567PQSclH+io7r2aCxWyT8bzZ+b2lLhU/jAGYhtpvRvxBr87zrAJtHSa9B XTOvKLJJv5+D37X62OaKCdFttXBEGBuwhiQ+w1wUGllj2XUIsUydDvxFgrR3xzxNkOdGBYsFoFu yR/4zr1E86YdQsZSQfgmFLQblOvAP8RV5h17SGAE5yCJoUcMhoSPW1NMjW2vi2d1XqZ0InCCdbV QCUH4nLtXy3A4ZWupc3VMIRyekMdRuFo94= X-Received: by 2002:a17:903:1112:b0:2dd:c053:c20a with SMTP id d9443c01a7336-2ddc053c2a1mr103690755ad.38.1789989479599; Mon, 21 Sep 2026 04:17:59 -0700 (PDT) Received: from FJ7FR2JRQ3.bytedance.net ([240e:694:e20:401::8]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2ddc17ba651sm31831125ad.53.2026.09.21.04.17.38 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Mon, 21 Sep 2026 04:17:59 -0700 (PDT) From: Zhanpeng Zhang To: Paul Walmsley , Palmer Dabbelt , Albert Ou , Alexandre Ghiti , Himanshu Chauhan , Conor Dooley , Anup Patel Cc: =?UTF-8?q?Cl=C3=A9ment=20L=C3=A9ger?= , Yunhui Cui , Atish Patra , Peter Zijlstra , Ingo Molnar , Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Alexander Shishkin , Jiri Olsa , Ian Rogers , Adrian Hunter , James Clark , Will Deacon , Thomas Gleixner , Jonathan Corbet , Randy Dunlap , Shuah Khan , Shuah Khan , Yuanzhu , Yicong Yang , Susheng Yang , linux-riscv@lists.infradead.org, linux-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-arm-kernel@lists.infradead.org, Zhanpeng Zhang Subject: [PATCH v10 RESEND 6/9] perf: RISC-V: support callchains with SSE delivery Date: Mon, 21 Sep 2026 19:15:03 +0800 Message-ID: <270f3bcd3a9933c8be59492ae4d7c6ff2414d7d7.1789974241.git.zhangzhanpeng.jasper@bytedance.com> X-Mailer: git-send-email 2.50.1 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" PMU overflow delivery through SSE enters Linux with a synthetic supervisor context on a dedicated event stack. The perf unwinder must use the context interrupted by the overflow rather than treating the SSE handler frame as the sampled frame. Use the interrupted pt_regs published by the RISC-V SSE entry path. Walk a kernel callchain only when the interrupted PC, SP, and frame pointer are consistent with the current task stack. For sensitive entry windows and IRQ stacks whose bounds cannot be proven, retain the interrupted PC without following an unsafe frame chain. User callchains continue through the existing nofault RISC-V user unwinder. If hstatus.SPV says the interrupted context was a guest, do not interpret the guest stack through the host address space. DWARF callchains additionally copy a raw user stack. The generic arch_perf_out_copy_user() implementation can take an exception-table handled fault when a source page is not resident. Repeated nested faults from the SSE handler can corrupt the interrupted kernel context under load. Provide an SSE-specific RISC-V copy path. Verify that the active page table belongs to current, use fast-only GUP to acquire each resident source page without falling back to a faulting slow path, and copy through the kernel mapping while holding the page reference. Stop at the first unavailable page and preserve perf's existing truncated-user-stack semantics. Keep the generic in-atomic user copy unchanged outside an SSE handler. Susheng Yang reported this failure with perf callchain workloads. Reported-by: Susheng Yang Signed-off-by: Zhanpeng Zhang --- arch/riscv/include/asm/perf_event.h | 10 ++ arch/riscv/kernel/perf_callchain.c | 142 ++++++++++++++++++++++++++++ 2 files changed, 152 insertions(+) diff --git a/arch/riscv/include/asm/perf_event.h b/arch/riscv/include/asm/p= erf_event.h index bcc928fd3785..ddc404794751 100644 --- a/arch/riscv/include/asm/perf_event.h +++ b/arch/riscv/include/asm/perf_event.h @@ -18,6 +18,16 @@ (regs)->sp =3D current_stack_pointer; \ (regs)->status =3D SR_PP; \ } + +#ifdef CONFIG_RISCV_SBI_SSE +/* + * Raw user-stack sampling can run in the NMI-like SSE context. Route it + * through an implementation that does not fault on a non-resident page. + */ +unsigned long riscv_perf_out_copy_user(void *dst, const void *src, + unsigned long n); +#define arch_perf_out_copy_user riscv_perf_out_copy_user +#endif #endif =20 #endif /* _ASM_RISCV_PERF_EVENT_H */ diff --git a/arch/riscv/kernel/perf_callchain.c b/arch/riscv/kernel/perf_ca= llchain.c index b465bc9eb870..ec75689c7aec 100644 --- a/arch/riscv/kernel/perf_callchain.c +++ b/arch/riscv/kernel/perf_callchain.c @@ -1,9 +1,15 @@ // SPDX-License-Identifier: GPL-2.0 /* Copyright (C) 2019 Hangzhou C-SKY Microsystems co.,ltd. */ =20 +#include +#include #include +#include +#include #include =20 +#include +#include #include =20 static bool fill_callchain(void *entry, unsigned long pc) @@ -11,6 +17,128 @@ static bool fill_callchain(void *entry, unsigned long p= c) return perf_callchain_store(entry, pc) =3D=3D 0; } =20 +#ifdef CONFIG_RISCV_SBI_SSE +static bool sse_addr_on_task_stack(unsigned long addr, unsigned long size) +{ + unsigned long end =3D addr + size; + unsigned long stack; + + if (end < addr) + return false; + + stack =3D (unsigned long)task_stack_page(current); + if (addr >=3D stack && end <=3D stack + THREAD_SIZE) + return true; + + return false; +} + +static bool sse_kernel_regs_safe(struct pt_regs *regs) +{ + unsigned long fp =3D frame_pointer(regs); + unsigned long pc =3D instruction_pointer(regs); + unsigned long sp =3D user_stack_pointer(regs); + + if (!__kernel_text_address(pc)) + return false; + if (!sse_addr_on_task_stack(sp, sizeof(unsigned long))) + return false; + if (fp < sizeof(struct stackframe)) + return false; + + return sse_addr_on_task_stack(fp - sizeof(struct stackframe), + sizeof(struct stackframe)); +} + +static bool sse_callchain_is_guest(const struct riscv_sse_interrupted_cont= ext *context) +{ + return context && (context->hstatus & HSTATUS_SPV); +} + +static bool sse_callchain_kernel(struct perf_callchain_entry_ctx *entry, + struct pt_regs *regs) +{ + const struct riscv_sse_interrupted_context *context; + unsigned long pc; + + context =3D riscv_sse_get_interrupted_context(); + if (!context || context->regs !=3D regs) + return false; + + /* A guest stack cannot be walked using the host kernel address space. */ + if (sse_callchain_is_guest(context)) + return true; + + if (user_mode(regs)) + return true; + + if (sse_kernel_regs_safe(regs)) { + walk_stackframe(NULL, regs, fill_callchain, entry); + return true; + } + + /* + * Keep the sample useful for sensitive entry paths and IRQ stacks. The + * generic walker does not take explicit IRQ stack bounds, and its + * THREAD_SIZE alignment assumption fails for non-vmapped IRQ stacks. + * Conservatively avoid walking IRQ stacks in every configuration. + */ + pc =3D instruction_pointer(regs); + if (__kernel_text_address(pc)) + perf_callchain_store(entry, pc); + + return true; +} + +unsigned long riscv_perf_out_copy_user(void *dst, const void *src, + unsigned long n) +{ + unsigned long addr =3D (unsigned long)src; + unsigned long copied =3D 0; + + /* Keep the generic fast path unchanged outside an SSE handler. */ + if (!riscv_sse_get_interrupted_context()) { + unsigned long ret; + + pagefault_disable(); + ret =3D __copy_from_user_inatomic(dst, src, n); + pagefault_enable(); + return ret; + } + + if (!access_ok(src, n)) + return n; + + /* Do not sample user memory through an unrelated active page table. */ + if (!current->mm || + (csr_read(CSR_SATP) & SATP_PPN) !=3D virt_to_pfn(current->mm->pgd)) + return n; + + while (copied < n) { + unsigned long offset =3D offset_in_page(addr); + unsigned long chunk =3D min(n - copied, PAGE_SIZE - offset); + struct page *page; + + /* + * Fast-only GUP cannot fault. This follows perf_virt_to_phys(): + * local interrupts remain disabled throughout SSE processing, so a + * concurrent unmap cannot complete its TLB teardown before this + * temporary reference is put. + */ + if (!get_user_page_fast_only(addr, 0, &page)) + break; + + memcpy((char *)dst + copied, + (char *)page_address(page) + offset, chunk); + put_page(page); + addr +=3D chunk; + copied +=3D chunk; + } + + return n - copied; +} +#endif + /* * This will be called when the target is in user mode * This function will only be called when we use @@ -28,6 +156,15 @@ static bool fill_callchain(void *entry, unsigned long p= c) void perf_callchain_user(struct perf_callchain_entry_ctx *entry, struct pt_regs *regs) { +#ifdef CONFIG_RISCV_SBI_SSE + const struct riscv_sse_interrupted_context *context; + + context =3D riscv_sse_get_interrupted_context(); + /* A guest stack cannot be walked using the host address space. */ + if (sse_callchain_is_guest(context)) + return; +#endif + if (perf_guest_state()) { /* TODO: We don't support guest os callchain now */ return; @@ -39,6 +176,11 @@ void perf_callchain_user(struct perf_callchain_entry_ct= x *entry, void perf_callchain_kernel(struct perf_callchain_entry_ctx *entry, struct pt_regs *regs) { +#ifdef CONFIG_RISCV_SBI_SSE + if (sse_callchain_kernel(entry, regs)) + return; +#endif + if (perf_guest_state()) { /* TODO: We don't support guest os callchain now */ return; --=20 2.50.1 (Apple Git-155) From nobody Fri Sep 25 21:02:40 2026 Received: from mail-pj2-f13.google.com (mail-pj2-f13.google.com [74.125.227.141]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9556648820E for ; Mon, 21 Sep 2026 11:18:22 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.227.141 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789989506; cv=none; b=Gk5urhAiGa1F/4mvcx563+uYLDU6pX2ARsE7N9ORls+RjeGLLXfSwojICdKWuUK/5AUsPjH4fdZCk5W3p+iDd6prgT+8XkmiXVtXVhcKjQRDx6D8+BS46AsRCa0FLChC8CprsDrZQjQRanb+fB7F03j4B8Ay7N0Gn6Vf8hUgcBk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789989506; c=relaxed/simple; bh=W1froEPQpyXeZdY5eAcB9WwF2/pdxQAcaSD6w5f2l+I=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=dJNo+WLTtgY5XN05OMe0o9pPmH445SYjk0SA7MBWjpBq5zoJVK6xhgnrnElr9yIC3Yrr410AtmK9BBOyM9hED/005RK4sspenyyoUXAnBZuZPVLN0SpBL4er1mNG4buzV/V4aWPU472nYFyzOrKVFpzYPM64uyMUrKx5XTGbMag= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com; spf=pass smtp.mailfrom=bytedance.com; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b=Xvf7kxpk; arc=none smtp.client-ip=74.125.227.141 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=bytedance.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b="Xvf7kxpk" Received: by mail-pj2-f13.google.com with SMTP id d9443c01a7336-2d8fb334e72so25335235ad.1 for ; Mon, 21 Sep 2026 04:18:22 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=bytedance.com; s=google; t=1789989502; x=1790594302; darn=vger.kernel.org; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:from:to:cc:subject :date:message-id:reply-to:content-type; bh=jWpqNRZSz8bbv9viIXEDpbwojwBy/hlcpZ3EVGBHDIY=; b=Xvf7kxpkCE4G+8EKUVqCWEVmi+lpJgmFsmt2LyE1pgADoeb/AVB4WDG7FkZSNynm1u VXQaL217LFnq2Rmvw06ILBCxXWs6NQYSOK2htLZd0nKGBsj9uJR66vfCZM96keOTH8jp 2wYaIyjRMbUVAdQ/PEDKAII7BPdTN6sqvi809KbZotKqbXfOZ3oyXKJQ3dM3nou+MrUF s1i8tSVDNYoukPVg7TZwKclA1wVgPtiHEDwf6G16cXmwwvDM0VDcdoaWo82O2aAjf1j4 JltTuaXbvVHMFWYRnJoR1HvCr7BgEuMMTeU2Tn6mIXiYmLcVsQjqnMICYWiVrEFFPnIh ruXw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789989502; x=1790594302; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=jWpqNRZSz8bbv9viIXEDpbwojwBy/hlcpZ3EVGBHDIY=; b=H4o9/1OhkvETErQNpvfvaQu6BvYPXrjqsirDEW5z3lo4solOtFCgt0+XRVJX2WsPOG 1ukg+NPUBLAxNYPMKrz71XMzl6SECtdDm+IPyIRWOsh49QJEBn8DvEE8TNh0UidMfU97 vAuE1yvVEaxHGgDladoXzBsorfXZtKP9yONBp0ws4xCRfe21Djzvdv3vs6vwdtoijkR2 cpTg8d9R7Kx7C42vlb+Cs5S7Ce3Jz73N/I5k2E/V4TPb2Y4PU055JsqEEUWFqwkh0KfM IJUXev25eM/yzmgZS/IF9a6DtS46rS+fpNJMSpIK8tHKsjeHlOSgRBnIS7tQY7I6Rr5r 4BCQ== X-Forwarded-Encrypted: i=1; AKwUvBzEb8YZ73ljpxazwhdXb4U71OIJQe9YXwKrU02z1tyxFBv20YAXlQC8RznxXkdIEc+dC7ebza3PfPkYBbI=@vger.kernel.org X-Gm-Message-State: AFuF++l7n6esowRh54sM7ikImVLXnwa4S7qF8Djz8T+Z/ji+qVpoq/NO xnJfpaA9OiPuCZ9BL6BIoYaMl8usfYT0Udr3WT/wkxknHZm6iTeR20iTmpDgx2kH/JU= X-Gm-Gg: AYBFou1qPpVgIc7HYpDwzjSLzsWmjHqaxiPNZ8oqTNFwWCEzwhzXhJIfikzMyRCAU2V guuG7aXZ7xTH1auEJmyPNJfxifzycI0VIu3R5JiINHBO4XydY78MLIE8OPUlzXc1Oq67vD0b2Li K3iWxM0xYY89a4tsihcuyyMeVzhbyHJmLFd7x637CaWD9GdniuqiruVT6bHLYj5l77KU6KOKYSh 7OQL25kWatJLnbHSEu7RKsCOAR8HIn+lIdFCW48fAyPDh1BF2Q33fiXxnCc+8w0Ns27b9yC4CBN 0jQ/ZPIPhumqYptlOWRC6ozJ5ykGZ2LAiV3H3PTE+hXSHnNhH9L+eHmu/hWGM3H6gWzdl+lraZI sLmVUXFjGnH/geDjuLQwQBuUfSig+v9C/Ezym+2EwIGmAuUM5h9+nXM7oCiG2pDDLwPa4kSV/2t oi6OIk1RTfX4oTDwiVhGprhXwBjc6WJd6A0jYKcmiR8Jojg9AkdxCylmpICe1f5sMrngQewM4vg J4747NhlAk47NcTx9bJs78zVL/FP2Y6Rxs= X-Received: by 2002:a17:902:f710:b0:2dd:64:12ac with SMTP id d9443c01a7336-2ddb1b994bemr141830075ad.11.1789989501626; Mon, 21 Sep 2026 04:18:21 -0700 (PDT) Received: from FJ7FR2JRQ3.bytedance.net ([240e:694:e20:401::8]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2ddc17ba651sm31831125ad.53.2026.09.21.04.18.00 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Mon, 21 Sep 2026 04:18:21 -0700 (PDT) From: Zhanpeng Zhang To: Paul Walmsley , Palmer Dabbelt , Albert Ou , Alexandre Ghiti , Himanshu Chauhan , Conor Dooley , Anup Patel Cc: =?UTF-8?q?Cl=C3=A9ment=20L=C3=A9ger?= , Yunhui Cui , Atish Patra , Peter Zijlstra , Ingo Molnar , Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Alexander Shishkin , Jiri Olsa , Ian Rogers , Adrian Hunter , James Clark , Will Deacon , Thomas Gleixner , Jonathan Corbet , Randy Dunlap , Shuah Khan , Shuah Khan , Yuanzhu , Yicong Yang , Susheng Yang , linux-riscv@lists.infradead.org, linux-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-arm-kernel@lists.infradead.org, Zhanpeng Zhang Subject: [PATCH v10 RESEND 7/9] perf: RISC-V: add support for SSE event Date: Mon, 21 Sep 2026 19:15:04 +0800 Message-ID: <4f7828b3e30407de9df205264bb94edc767d318c.1789974241.git.zhangzhanpeng.jasper@bytedance.com> X-Mailer: git-send-email 2.50.1 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable From: Cl=C3=A9ment L=C3=A9ger Register a handler for the local PMU overflow SSE event so that RISC-V perf can receive overflows even when normal S-mode interrupts are masked. Reuse the existing overflow handler and pass it the interrupted pt_regs rebuilt by the architecture SSE entry path. Select the delivery mechanism once during PMU probe. Prefer SSE when its event can be registered and enabled. If the extension or PMU event is explicitly unsupported, use the ordinary PMU interrupt. Do not enable the interrupt after any other SSE setup failure or when a crash kernel may have inherited firmware state. Install the PMU enable and disable callbacks only after SSE delivery becomes active. Keep the local PMU SSE event disabled across CPU power management. On entry, the generic SSE notifier masks the hart before the lower-priority PMU notifier disables the event and stops the counters. On exit, the SSE notifier first unmasks the hart while the event remains disabled. The PMU notifier then restores counters and event userpage state before enabling the event. An unmask failure stops the notifier chain and leaves the counters stopped. Ordinary PMU interrupts retain their existing notifier ordering. An SSE overflow can arrive as soon as the event is enabled during probe. Publish the counter mask before SSE setup, so an early handler can stop the counter source even before perf starts admitting normal samples. After a real overflow, restart only counters whose perf state is still running. Honor a non-zero return from perf_event_overflow() and leave throttled events stopped. Guest attribution is not part of this version. Detect an interrupted guest from hstatus.SPV and skip its sample while still updating the period and counter state, rather than exposing guest state as a host sample. The perf PMU callbacks cannot return errors. If an SSE transition or interrupted-context read fails, latch the failure per CPU and stop its mapped events through the normal perf state transitions. Do not reset the firmware counter mapping behind perf, restart a failed event, or attempt a runtime switch to IRQ delivery. During cleanup, close callback admission and synchronously drain each CPU before disabling and unregistering the SSE event. If firmware refuses the cleanup, stop the counter source and transfer the event to the SSE core so later CPU hotplug or shutdown processing can retry without using freed PMU callback state. Signed-off-by: Cl=C3=A9ment L=C3=A9ger Co-developed-by: Himanshu Chauhan Signed-off-by: Himanshu Chauhan Co-developed-by: Zhanpeng Zhang Signed-off-by: Zhanpeng Zhang --- Documentation/arch/riscv/index.rst | 1 + Documentation/arch/riscv/pmu-sse.rst | 55 +++ MAINTAINERS | 1 + drivers/firmware/riscv/riscv_sbi_sse.c | 9 + drivers/perf/Kconfig | 11 + drivers/perf/riscv_pmu.c | 14 +- drivers/perf/riscv_pmu_sbi.c | 540 ++++++++++++++++++++----- include/linux/perf/riscv_pmu.h | 20 +- include/linux/riscv_sbi_sse.h | 6 + 9 files changed, 564 insertions(+), 93 deletions(-) create mode 100644 Documentation/arch/riscv/pmu-sse.rst diff --git a/Documentation/arch/riscv/index.rst b/Documentation/arch/riscv/= index.rst index ac535c52d509..5cb909c83a88 100644 --- a/Documentation/arch/riscv/index.rst +++ b/Documentation/arch/riscv/index.rst @@ -11,6 +11,7 @@ RISC-V architecture vm-layout hwprobe patch-acceptance + pmu-sse uabi vector cmodx diff --git a/Documentation/arch/riscv/pmu-sse.rst b/Documentation/arch/risc= v/pmu-sse.rst new file mode 100644 index 000000000000..b7458a9d0116 --- /dev/null +++ b/Documentation/arch/riscv/pmu-sse.rst @@ -0,0 +1,55 @@ +.. SPDX-License-Identifier: GPL-2.0 + +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D +RISC-V PMU overflow delivery through SSE +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +When ``CONFIG_RISCV_PMU_SBI_SSE`` is enabled and firmware provides the loc= al +PMU overflow event, the RISC-V SBI PMU driver uses Supervisor Software Eve= nts +(SSE) to deliver counter overflows. + +Delivery selection +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +The delivery mechanism is selected once while the PMU device is probed. T= he +driver first tries to register and enable the local PMU overflow SSE event= . It +tries the ordinary PMU interrupt only when the SSE extension or the local = PMU +event is explicitly unsupported. It does not change the delivery mechanism +after the PMU has been registered. + +Other SSE setup errors do not prove that firmware released the overflow ro= ute, +so the driver does not enable the ordinary PMU interrupt. The PMU remains +available for counting but not sampling. If setup retained an SSE event, = the +PMU keeps quiescing that event around perf scheduling changes without rear= ming +it. This preserves callback ownership without creating two possible deliv= ery +mechanisms for the same hardware overflow. + +Interrupted context +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +SSE enters Linux through a supervisor handler context constructed by firmw= are. +The entry contract preserves the interrupted GPRs except ``a6`` and ``a7``, +which Linux reads from the interrupted-register event attributes. Linux t= hen +reconstructs the interrupted ``pt_regs`` and publishes it while dispatchin= g the +event. + +Perf uses that context for register samples and callchains. Kernel stack +walking verifies that the interrupted frame belongs to the current task st= ack +before dereferencing it. For sensitive entry paths and IRQ stacks, Linux +records the interrupted PC without walking a stack whose bounds cannot be +proved. User callchains use the existing no-fault user unwinder. A DWARF +raw user-stack copy from an SSE handler must not take a page fault, so it +walks the current task's page tables with fast-only GUP, copies each resid= ent +page through its kernel mapping, and stops at the first non-resident page, +preserving perf's truncated-user-stack semantics. + +CPU power management +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +The PMU and SSE CPU power-management callbacks are ordered according to the +selected delivery mechanism. With SSE delivery, SSE events are masked bef= ore +the lower-priority PMU callback disables the local PMU event and stops the +counters on entry. On exit, the hart is unmasked while the local PMU even= t is +still disabled. The PMU callback then restores the counters and enables t= he +event, so an unmask failure leaves the counters stopped. The ordinary +interrupt path retains the existing PMU ordering. diff --git a/MAINTAINERS b/MAINTAINERS index 0070a3ed6321..a6fa10a0a94d 100644 --- a/MAINTAINERS +++ b/MAINTAINERS @@ -23563,6 +23563,7 @@ M: Atish Patra R: Anup Patel L: linux-riscv@lists.infradead.org S: Supported +F: Documentation/arch/riscv/pmu-sse.rst F: drivers/perf/riscv_pmu.c F: drivers/perf/riscv_pmu_legacy.c F: drivers/perf/riscv_pmu_sbi.c diff --git a/drivers/firmware/riscv/riscv_sbi_sse.c b/drivers/firmware/risc= v/riscv_sbi_sse.c index 0cb738bab6c3..69ce98380331 100644 --- a/drivers/firmware/riscv/riscv_sbi_sse.c +++ b/drivers/firmware/riscv/riscv_sbi_sse.c @@ -161,6 +161,15 @@ static int sse_sbi_event_func(struct sse_event *event,= unsigned long func) return 0; } =20 +bool sse_event_is_enabled_local(struct sse_event *event) +{ + if (!sse_event_is_global(event->evt_id)) + lockdep_assert_preemption_disabled(); + + return sse_get_reg_evt(event)->is_enabled; +} +EXPORT_SYMBOL_GPL(sse_event_is_enabled_local); + int sse_event_disable_local(struct sse_event *event) { if (!sse_event_is_global(event->evt_id)) diff --git a/drivers/perf/Kconfig b/drivers/perf/Kconfig index 245e7bb763b9..10b047856121 100644 --- a/drivers/perf/Kconfig +++ b/drivers/perf/Kconfig @@ -105,6 +105,17 @@ config RISCV_PMU_SBI full perf feature support i.e. counter overflow, privilege mode filtering, counter configuration. =20 +config RISCV_PMU_SBI_SSE + depends on RISCV_PMU_SBI && RISCV_SBI_SSE + bool "RISC-V PMU SSE events" + default n + help + Say y if you want to use SSE events to deliver PMU interrupts. This + provides a way to profile the kernel at any level by using NMI-like + SSE events. SSE events being really intrusive, this option allows + to select it only if needed. See Documentation/arch/riscv/pmu-sse.rst + for interrupted-context limitations. + config STARFIVE_STARLINK_PMU depends on ARCH_STARFIVE || COMPILE_TEST depends on 64BIT diff --git a/drivers/perf/riscv_pmu.c b/drivers/perf/riscv_pmu.c index 8e3cd0f35336..1f65c640b693 100644 --- a/drivers/perf/riscv_pmu.c +++ b/drivers/perf/riscv_pmu.c @@ -13,6 +13,7 @@ #include #include #include +#include #include #include =20 @@ -247,6 +248,11 @@ void riscv_pmu_start(struct perf_event *event, int fla= gs) if (flags & PERF_EF_RELOAD) WARN_ON_ONCE(!(event->hw.state & PERF_HES_UPTODATE)); =20 +#ifdef CONFIG_RISCV_PMU_SBI_SSE + if (unlikely(this_cpu_ptr(rvpmu->hw_events)->sse_failed)) + return; +#endif + hwc->state =3D 0; riscv_pmu_event_set_period(event); init_val =3D local64_read(&hwc->prev_count) & max_period; @@ -306,6 +312,7 @@ static int riscv_pmu_event_init(struct perf_event *even= t) struct hw_perf_event *hwc =3D &event->hw; struct riscv_pmu *rvpmu =3D to_riscv_pmu(event->pmu); int mapped_event; + int ret; u64 event_config =3D 0; uint64_t cmask; =20 @@ -331,8 +338,11 @@ static int riscv_pmu_event_init(struct perf_event *eve= nt) hwc->idx =3D -1; hwc->event_base =3D mapped_event; =20 - if (rvpmu->event_init) - rvpmu->event_init(event); + if (rvpmu->event_init) { + ret =3D rvpmu->event_init(event); + if (ret) + return ret; + } =20 if (!is_sampling_event(event)) { /* diff --git a/drivers/perf/riscv_pmu_sbi.c b/drivers/perf/riscv_pmu_sbi.c index 2991dd92def2..6948b592ccf1 100644 --- a/drivers/perf/riscv_pmu_sbi.c +++ b/drivers/perf/riscv_pmu_sbi.c @@ -16,13 +16,16 @@ #include #include #include +#include #include +#include #include #include #include =20 #include #include +#include #include #include #include @@ -95,7 +98,6 @@ static bool riscv_pmu_use_irq; static unsigned int riscv_pmu_irq_num; static unsigned int riscv_pmu_irq_mask; static unsigned int riscv_pmu_irq; - /* Cache the available counters in a bitmask */ static DECLARE_BITMAP(cmask, RISCV_MAX_COUNTERS); =20 @@ -925,7 +927,7 @@ static int pmu_sbi_get_ctrinfo(int nctr, unsigned long = *mask) return 0; } =20 -static inline void pmu_sbi_stop_all(struct riscv_pmu *pmu) +static inline void pmu_sbi_stop_all_mask(const unsigned long *ctr_mask) { int i; =20 @@ -934,14 +936,24 @@ static inline void pmu_sbi_stop_all(struct riscv_pmu = *pmu) * which may include counters that are not enabled yet. */ for (i =3D 0; i < BITS_TO_LONGS(RISCV_MAX_COUNTERS); i++) { - if (!pmu->cmask[i]) + if (!ctr_mask[i]) continue; sbi_ecall(SBI_EXT_PMU, SBI_EXT_PMU_COUNTER_STOP, - i * BITS_PER_LONG, pmu->cmask[i], + i * BITS_PER_LONG, ctr_mask[i], SBI_PMU_STOP_FLAG_RESET, 0, 0, 0); } } =20 +static inline void pmu_sbi_stop_all(struct riscv_pmu *pmu) +{ + pmu_sbi_stop_all_mask(pmu->cmask); +} + +static void pmu_sbi_stop_all_cpu(void *info) +{ + pmu_sbi_stop_all(info); +} + static inline void pmu_sbi_stop_hw_ctrs(struct riscv_pmu *pmu) { struct cpu_hw_events *cpu_hw_evt =3D this_cpu_ptr(pmu->hw_events); @@ -989,7 +1001,7 @@ static inline void pmu_sbi_stop_hw_ctrs(struct riscv_p= mu *pmu) static inline void pmu_sbi_start_ovf_ctrs_sbi(struct cpu_hw_events *cpu_hw= _evt, u64 ctr_ovf_mask) { - int idx =3D 0, i; + int idx, i; struct perf_event *event; unsigned long flag =3D SBI_PMU_START_FLAG_SET_INIT_VALUE; unsigned long ctr_start_mask =3D 0; @@ -998,7 +1010,19 @@ static inline void pmu_sbi_start_ovf_ctrs_sbi(struct = cpu_hw_events *cpu_hw_evt, u64 init_val =3D 0; =20 for (i =3D 0; i < BITS_TO_LONGS(RISCV_MAX_COUNTERS); i++) { - ctr_start_mask =3D cpu_hw_evt->used_hw_ctrs[i] & ~ctr_ovf_mask; + unsigned long ctr_ovf_mask_word; + int lidx; + + ctr_ovf_mask_word =3D ctr_ovf_mask >> (i * BITS_PER_LONG); + ctr_start_mask =3D 0; + for_each_set_bit(idx, &cpu_hw_evt->used_hw_ctrs[i], BITS_PER_LONG) { + lidx =3D idx + i * BITS_PER_LONG; + event =3D cpu_hw_evt->events[lidx]; + if (event && !(event->hw.state & PERF_HES_STOPPED)) + ctr_start_mask |=3D BIT(idx); + } + ctr_start_mask &=3D ~ctr_ovf_mask_word; + /* Start all the counters that did not overflow in a single shot */ if (ctr_start_mask) { sbi_ecall(SBI_EXT_PMU, SBI_EXT_PMU_COUNTER_START, i * BITS_PER_LONG, @@ -1007,23 +1031,25 @@ static inline void pmu_sbi_start_ovf_ctrs_sbi(struc= t cpu_hw_events *cpu_hw_evt, } =20 /* Reinitialize and start all the counter that overflowed */ - while (ctr_ovf_mask) { - if (ctr_ovf_mask & 0x01) { - event =3D cpu_hw_evt->events[idx]; - hwc =3D &event->hw; - max_period =3D riscv_pmu_ctr_get_width_mask(event); - init_val =3D local64_read(&hwc->prev_count) & max_period; + for (idx =3D 0; idx < RISCV_MAX_COUNTERS; idx++) { + if (!(ctr_ovf_mask & BIT_ULL(idx))) + continue; + + event =3D cpu_hw_evt->events[idx]; + if (!event || event->hw.state & PERF_HES_STOPPED) + continue; + + hwc =3D &event->hw; + max_period =3D riscv_pmu_ctr_get_width_mask(event); + init_val =3D local64_read(&hwc->prev_count) & max_period; #if defined(CONFIG_32BIT) - sbi_ecall(SBI_EXT_PMU, SBI_EXT_PMU_COUNTER_START, idx, 1, - flag, init_val, init_val >> 32, 0); + sbi_ecall(SBI_EXT_PMU, SBI_EXT_PMU_COUNTER_START, idx, 1, + flag, init_val, init_val >> 32, 0); #else - sbi_ecall(SBI_EXT_PMU, SBI_EXT_PMU_COUNTER_START, idx, 1, - flag, init_val, 0, 0); + sbi_ecall(SBI_EXT_PMU, SBI_EXT_PMU_COUNTER_START, idx, 1, + flag, init_val, 0, 0); #endif - perf_event_update_userpage(event); - } - ctr_ovf_mask =3D ctr_ovf_mask >> 1; - idx++; + perf_event_update_userpage(event); } } =20 @@ -1052,31 +1078,111 @@ static inline void pmu_sbi_start_ovf_ctrs_snapshot= (struct cpu_hw_events *cpu_hw_ } =20 for (i =3D 0; i < BITS_TO_LONGS(RISCV_MAX_COUNTERS); i++) { + unsigned long ctr_start_mask =3D 0; + int lidx; + /* Restore the counter values to relative indices for used hw counters */ - for_each_set_bit(idx, &cpu_hw_evt->used_hw_ctrs[i], BITS_PER_LONG) - sdata->ctr_values[idx] =3D - cpu_hw_evt->snapshot_cval_shcopy[idx + i * BITS_PER_LONG]; + for_each_set_bit(idx, &cpu_hw_evt->used_hw_ctrs[i], BITS_PER_LONG) { + lidx =3D idx + i * BITS_PER_LONG; + event =3D cpu_hw_evt->events[lidx]; + if (event && !(event->hw.state & PERF_HES_STOPPED)) + ctr_start_mask |=3D BIT(idx); + + sdata->ctr_values[idx] =3D cpu_hw_evt->snapshot_cval_shcopy[lidx]; + } + /* Start all the counters in a single shot */ - sbi_ecall(SBI_EXT_PMU, SBI_EXT_PMU_COUNTER_START, idx * BITS_PER_LONG, - cpu_hw_evt->used_hw_ctrs[i], flag, 0, 0, 0); + if (ctr_start_mask) + sbi_ecall(SBI_EXT_PMU, SBI_EXT_PMU_COUNTER_START, + i * BITS_PER_LONG, ctr_start_mask, flag, 0, 0, 0); } } =20 +static bool pmu_sbi_sse_failed(struct cpu_hw_events *cpu_hw_evt) +{ +#ifdef CONFIG_RISCV_PMU_SBI_SSE + return READ_ONCE(cpu_hw_evt->sse_failed); +#else + return false; +#endif +} + static void pmu_sbi_start_overflow_mask(struct riscv_pmu *pmu, u64 ctr_ovf_mask) { struct cpu_hw_events *cpu_hw_evt =3D this_cpu_ptr(pmu->hw_events); =20 + if (unlikely(pmu_sbi_sse_failed(cpu_hw_evt))) + return; + if (sbi_pmu_snapshot_available()) pmu_sbi_start_ovf_ctrs_snapshot(cpu_hw_evt, ctr_ovf_mask); else pmu_sbi_start_ovf_ctrs_sbi(cpu_hw_evt, ctr_ovf_mask); } =20 -static irqreturn_t pmu_sbi_ovf_handler(int irq, void *dev) +#ifdef CONFIG_RISCV_PMU_SBI_SSE +/* + * A local SSE delivery failure makes the current PMU state unsafe to resu= me. + * Latch the failure before stopping mapped events so the SSE transition a= nd + * overflow restart paths cannot undo the fail-safe while they are quiesce= d. + */ +static void pmu_sbi_fail_sse(struct riscv_pmu *pmu, const char *op, int re= t) +{ + struct cpu_hw_events *cpu_hw_evt =3D this_cpu_ptr(pmu->hw_events); + struct perf_event *event; + int idx; + + if (READ_ONCE(cpu_hw_evt->sse_failed)) + return; + + WRITE_ONCE(cpu_hw_evt->sse_failed, true); + pr_err_ratelimited("failed to %s local PMU SSE event: %d; stopping counte= rs\n", + op, ret); + + for (idx =3D 0; idx < RISCV_MAX_COUNTERS; idx++) { + event =3D cpu_hw_evt->events[idx]; + if (event) + riscv_pmu_stop(event, PERF_EF_UPDATE); + } +} + +static void pmu_sbi_sse_disable(struct pmu *pmu) +{ + struct riscv_pmu *rvpmu =3D to_riscv_pmu(pmu); + struct cpu_hw_events *cpu_hw_evt =3D this_cpu_ptr(rvpmu->hw_events); + int ret; + + if (!READ_ONCE(rvpmu->sse_active) || + READ_ONCE(cpu_hw_evt->sse_failed)) + return; + + ret =3D sse_event_disable_local(rvpmu->sse_evt); + if (unlikely(ret)) + pmu_sbi_fail_sse(rvpmu, "disable", ret); +} + +static void pmu_sbi_sse_enable(struct pmu *pmu) +{ + struct riscv_pmu *rvpmu =3D to_riscv_pmu(pmu); + struct cpu_hw_events *cpu_hw_evt =3D this_cpu_ptr(rvpmu->hw_events); + int ret; + + if (!READ_ONCE(rvpmu->sse_active) || + READ_ONCE(cpu_hw_evt->sse_failed)) + return; + + ret =3D sse_event_enable_local(rvpmu->sse_evt); + if (unlikely(ret)) + pmu_sbi_fail_sse(rvpmu, "enable", ret); +} +#endif + +static irqreturn_t pmu_sbi_ovf_handler(struct cpu_hw_events *cpu_hw_evt, + struct pt_regs *regs, bool from_sse, + bool from_guest) { struct perf_sample_data data; - struct pt_regs *regs; struct hw_perf_event *hw_evt; union sbi_pmu_ctr_info *info; int lidx, hidx, fidx; @@ -1084,7 +1190,6 @@ static irqreturn_t pmu_sbi_ovf_handler(int irq, void = *dev) struct perf_event *event; u64 overflow; u64 overflowed_ctrs =3D 0; - struct cpu_hw_events *cpu_hw_evt =3D dev; u64 start_clock =3D sched_clock(); struct riscv_pmu_snapshot_data *sdata; =20 @@ -1093,21 +1198,32 @@ static irqreturn_t pmu_sbi_ovf_handler(int irq, voi= d *dev) =20 sdata =3D cpu_hw_evt->snapshot_addr; =20 - /* Firmware counter don't support overflow yet */ + /* + * SSE can arrive before perf installs an event. The early exits below + * must stop the PMU source before firmware completes the SSE. + */ fidx =3D find_first_bit(cpu_hw_evt->used_hw_ctrs, RISCV_MAX_COUNTERS); if (fidx =3D=3D RISCV_MAX_COUNTERS) { - csr_clear(CSR_SIP, BIT(riscv_pmu_irq_num)); + if (from_sse) + pmu_sbi_stop_all_mask(cmask); + else + csr_clear(CSR_SIP, BIT(riscv_pmu_irq_num)); return IRQ_NONE; } =20 event =3D cpu_hw_evt->events[fidx]; if (!event) { - ALT_SBI_PMU_OVF_CLEAR_PENDING(riscv_pmu_irq_mask); + if (from_sse) + pmu_sbi_stop_all_mask(cmask); + else + ALT_SBI_PMU_OVF_CLEAR_PENDING(riscv_pmu_irq_mask); return IRQ_NONE; } =20 pmu =3D to_riscv_pmu(event->pmu); pmu_sbi_stop_hw_ctrs(pmu); + if (unlikely(pmu_sbi_sse_failed(cpu_hw_evt))) + return IRQ_NONE; =20 /* Overflow status register should only be read after counter are stopped= */ if (sbi_pmu_snapshot_available()) @@ -1117,15 +1233,17 @@ static irqreturn_t pmu_sbi_ovf_handler(int irq, voi= d *dev) =20 /* * Overflow interrupt pending bit should only be cleared after stopping - * all the counters to avoid any race condition. + * all the counters to avoid any race condition. When using SSE, + * interrupt is cleared when stopping counters. */ - ALT_SBI_PMU_OVF_CLEAR_PENDING(riscv_pmu_irq_mask); + if (!from_sse) + ALT_SBI_PMU_OVF_CLEAR_PENDING(riscv_pmu_irq_mask); =20 /* No overflow bit is set */ - if (!overflow) + if (!overflow) { + pmu_sbi_start_overflow_mask(pmu, 0); return IRQ_NONE; - - regs =3D get_irq_regs(); + } =20 for_each_set_bit(lidx, cpu_hw_evt->used_hw_ctrs, RISCV_MAX_COUNTERS) { struct perf_event *event =3D cpu_hw_evt->events[lidx]; @@ -1150,6 +1268,12 @@ static irqreturn_t pmu_sbi_ovf_handler(int irq, void= *dev) if (!(overflow & BIT_ULL(hidx))) continue; =20 +#ifdef CONFIG_CPU_PM + /* Do not let CPU-PM resume override this overflow decision. */ + if (from_sse) + clear_bit(lidx, cpu_hw_evt->pm_resume_hw_ctrs); +#endif + /* * Keep a track of overflowed counters so that they can be started * with updated initial value. @@ -1162,6 +1286,14 @@ static irqreturn_t pmu_sbi_ovf_handler(int irq, void= *dev) hw_evt->state |=3D PERF_HES_UPTODATE; perf_sample_data_init(&data, 0, hw_evt->last_period); if (riscv_pmu_event_set_period(event)) { + int overflow_ret; + + /* Guest attribution and guest stack sampling are not supported yet. */ + if (from_guest) { + hw_evt->state =3D 0; + continue; + } + /* * Unlike other ISAs, RISC-V don't have to disable interrupts * to avoid throttling here. As per the specification, the @@ -1170,128 +1302,291 @@ static irqreturn_t pmu_sbi_ovf_handler(int irq, v= oid *dev) * TODO: We will need to stop the guest counters once * virtualization support is added. */ - perf_event_overflow(event, &data, regs); + overflow_ret =3D perf_event_overflow(event, &data, regs); + if (!overflow_ret) + hw_evt->state =3D 0; + } else { + hw_evt->state =3D 0; } - /* Reset the state as we are going to start the counter after the loop */ - hw_evt->state =3D 0; } =20 pmu_sbi_start_overflow_mask(pmu, overflowed_ctrs); + perf_sample_event_took(sched_clock() - start_clock); =20 return IRQ_HANDLED; } =20 -static int pmu_sbi_starting_cpu(unsigned int cpu, struct hlist_node *node) +static irqreturn_t pmu_sbi_ovf_irq_handler(int irq, void *dev) { - struct riscv_pmu *pmu =3D hlist_entry_safe(node, struct riscv_pmu, node); - struct cpu_hw_events *cpu_hw_evt =3D this_cpu_ptr(pmu->hw_events); - - /* - * We keep enabling userspace access to CYCLE, TIME and INSTRET via the - * legacy option but that will be removed in the future. - */ - if (sysctl_perf_user_access =3D=3D SYSCTL_LEGACY) - csr_write(CSR_SCOUNTEREN, 0x7); - else - csr_write(CSR_SCOUNTEREN, 0x2); + return pmu_sbi_ovf_handler(dev, get_irq_regs(), false, false); +} =20 - /* Stop all the counters so that they can be enabled from perf */ - pmu_sbi_stop_all(pmu); +#ifdef CONFIG_RISCV_PMU_SBI_SSE +static int pmu_sbi_ovf_sse_handler(u32 evt, void *arg, struct pt_regs *reg= s) +{ + const struct riscv_sse_interrupted_context *context; + struct riscv_pmu *pmu =3D arg; + struct cpu_hw_events *hw_event =3D raw_cpu_ptr(pmu->hw_events); + bool from_guest; + + if (unlikely(!READ_ONCE(pmu->sse_active))) { + pmu_sbi_stop_all(pmu); + return -EIO; + } =20 - if (riscv_pmu_use_irq) { - cpu_hw_evt->irq =3D riscv_pmu_irq; - ALT_SBI_PMU_OVF_CLEAR_PENDING(riscv_pmu_irq_mask); - enable_percpu_irq(riscv_pmu_irq, IRQ_TYPE_NONE); + if (unlikely(!regs)) { + pmu_sbi_fail_sse(pmu, "read interrupted context for", -EIO); + return -EIO; } =20 - if (sbi_pmu_snapshot_available()) - return pmu_sbi_snapshot_setup(pmu, cpu); + context =3D riscv_sse_get_interrupted_context(); + from_guest =3D context && context->regs =3D=3D regs && + (context->hstatus & HSTATUS_SPV); + pmu_sbi_ovf_handler(hw_event, regs, true, from_guest); =20 return 0; } =20 -static int pmu_sbi_dying_cpu(unsigned int cpu, struct hlist_node *node) +static int pmu_sbi_setup_sse(struct riscv_pmu *pmu) { - if (riscv_pmu_use_irq) { - disable_percpu_irq(riscv_pmu_irq); - } + int ret; + struct sse_event *evt; =20 - /* Disable all counters access for user mode now */ - csr_write(CSR_SCOUNTEREN, 0x0); + evt =3D sse_event_register(SBI_SSE_EVENT_LOCAL_PMU_OVERFLOW, 0, + pmu_sbi_ovf_sse_handler, pmu); + if (IS_ERR(evt)) + return PTR_ERR(evt); + pmu->sse_evt =3D evt; =20 - if (sbi_pmu_snapshot_available()) - return pmu_sbi_snapshot_disable(); + ret =3D sse_event_enable(evt); + if (ret) { + int cleanup_ret; =20 - return 0; + cleanup_ret =3D sse_event_disable(evt); + if (cleanup_ret) { + pr_warn("failed to disable SSE event after setup error: %d\n", + cleanup_ret); + return cleanup_ret; + } + + cleanup_ret =3D sse_event_unregister(evt); + + if (cleanup_ret) { + pr_warn("failed to unregister SSE event: %d\n", + cleanup_ret); + } else { + pmu->sse_evt =3D NULL; + } + return cleanup_ret ?: ret; + } + + WRITE_ONCE(pmu->sse_active, true); + pr_info("using SSE for PMU event delivery\n"); + + return ret; } =20 -static int pmu_sbi_setup_irqs(struct riscv_pmu *pmu, struct platform_devic= e *pdev) +static void pmu_sbi_cleanup_sse(struct riscv_pmu *pmu) { + struct sse_event *sse_evt; int ret; - struct cpu_hw_events __percpu *hw_events =3D pmu->hw_events; - struct irq_domain *domain =3D NULL; =20 + sse_evt =3D pmu->sse_evt; + if (!sse_evt) + return; + /* + * Close callback admission before draining each CPU. PMU callbacks run + * with local interrupts disabled, so the synchronous IPI cannot complete + * until a callback that observed the old state has returned. + */ + WRITE_ONCE(pmu->sse_active, false); + on_each_cpu(pmu_sbi_stop_all_cpu, pmu, 1); + + ret =3D sse_event_disable(sse_evt); + if (ret) { + pr_warn("failed to disable SSE event: %d\n", ret); + goto retain; + } + + ret =3D sse_event_unregister(sse_evt); + if (ret) { + pr_warn("failed to unregister SSE event: %d\n", ret); + goto retain; + } + + pmu->sse_evt =3D NULL; + return; + +retain: + sse_event_cleanup(sse_evt); + pmu->sse_evt =3D NULL; +} + +static bool pmu_sbi_sse_state_retained(struct riscv_pmu *pmu) +{ + return pmu->sse_evt; +} +#else +static int pmu_sbi_setup_sse(struct riscv_pmu *pmu) +{ + return -EOPNOTSUPP; +} + +static void pmu_sbi_cleanup_sse(struct riscv_pmu *pmu) {} + +static bool pmu_sbi_sse_state_retained(struct riscv_pmu *pmu) +{ + return false; +} +#endif + +static bool pmu_sbi_select_irq(void) +{ if (riscv_isa_extension_available(NULL, SSCOFPMF)) { riscv_pmu_irq_num =3D RV_IRQ_PMU; - riscv_pmu_use_irq =3D true; + return true; } else if (IS_ENABLED(CONFIG_ERRATA_THEAD_PMU) && riscv_cached_mvendorid(0) =3D=3D THEAD_VENDOR_ID && riscv_cached_marchid(0) =3D=3D 0 && riscv_cached_mimpid(0) =3D=3D 0) { riscv_pmu_irq_num =3D THEAD_C9XX_RV_IRQ_PMU; - riscv_pmu_use_irq =3D true; + return true; } else if (riscv_has_vendor_extension_unlikely(ANDES_VENDOR_ID, RISCV_ISA_VENDOR_EXT_XANDESPMU) && IS_ENABLED(CONFIG_ANDES_CUSTOM_PMU)) { riscv_pmu_irq_num =3D ANDES_SLI_CAUSE_BASE + ANDES_RV_IRQ_PMOVI; - riscv_pmu_use_irq =3D true; + return true; } =20 - riscv_pmu_irq_mask =3D BIT(riscv_pmu_irq_num % BITS_PER_LONG); + return false; +} + +static int pmu_sbi_setup_irq(struct riscv_pmu *pmu) +{ + struct cpu_hw_events __percpu *hw_events =3D pmu->hw_events; + struct irq_domain *domain; + int ret; =20 - if (!riscv_pmu_use_irq) + if (!pmu_sbi_select_irq()) return -EOPNOTSUPP; =20 + riscv_pmu_irq_mask =3D BIT(riscv_pmu_irq_num % BITS_PER_LONG); + domain =3D irq_find_matching_fwnode(riscv_get_intc_hwnode(), DOMAIN_BUS_ANY); if (!domain) { pr_err("Failed to find INTC IRQ root domain\n"); - ret =3D -ENODEV; - goto err; + return -ENODEV; } =20 riscv_pmu_irq =3D irq_create_mapping(domain, riscv_pmu_irq_num); if (!riscv_pmu_irq) { pr_err("Failed to map PMU interrupt for node\n"); - ret =3D -ENODEV; - goto err; + return -ENODEV; } =20 - ret =3D request_percpu_irq(riscv_pmu_irq, pmu_sbi_ovf_handler, "riscv-pmu= ", hw_events); + ret =3D request_percpu_irq(riscv_pmu_irq, pmu_sbi_ovf_irq_handler, + "riscv-pmu", hw_events); if (ret) { pr_err("registering percpu irq failed [%d]\n", ret); irq_dispose_mapping(riscv_pmu_irq); riscv_pmu_irq =3D 0; - goto err; + return ret; } =20 return 0; -err: - riscv_pmu_use_irq =3D false; - return ret; +} + +static int pmu_sbi_starting_cpu(unsigned int cpu, struct hlist_node *node) +{ + struct riscv_pmu *pmu =3D hlist_entry_safe(node, struct riscv_pmu, node); + struct cpu_hw_events *cpu_hw_evt =3D this_cpu_ptr(pmu->hw_events); + + /* + * We keep enabling userspace access to CYCLE, TIME and INSTRET via the + * legacy option but that will be removed in the future. + */ + if (sysctl_perf_user_access =3D=3D SYSCTL_LEGACY) + csr_write(CSR_SCOUNTEREN, 0x7); + else + csr_write(CSR_SCOUNTEREN, 0x2); + + /* Stop all the counters so that they can be enabled from perf */ + pmu_sbi_stop_all(pmu); + + if (riscv_pmu_use_irq) { + cpu_hw_evt->irq =3D riscv_pmu_irq; + ALT_SBI_PMU_OVF_CLEAR_PENDING(riscv_pmu_irq_mask); + enable_percpu_irq(riscv_pmu_irq, IRQ_TYPE_NONE); + } + + if (sbi_pmu_snapshot_available()) + return pmu_sbi_snapshot_setup(pmu, cpu); + + return 0; +} + +static int pmu_sbi_dying_cpu(unsigned int cpu, struct hlist_node *node) +{ + if (riscv_pmu_use_irq) + disable_percpu_irq(riscv_pmu_irq); + + /* Disable all counters access for user mode now */ + csr_write(CSR_SCOUNTEREN, 0x0); + + if (sbi_pmu_snapshot_available()) + return pmu_sbi_snapshot_disable(); + + return 0; +} + +static int pmu_sbi_setup_irqs(struct riscv_pmu *pmu, struct platform_devic= e *pdev) +{ + int irq_ret; + int sse_ret; + + /* Do not claim an IRQ route while a crash kernel leaves SSE state intact= . */ + if (is_kdump_kernel() && riscv_sse_available()) { + pr_warn("PMU delivery unavailable with retained crash-kernel SSE state\n= "); + return -EUCLEAN; + } + + sse_ret =3D pmu_sbi_setup_sse(pmu); + if (!sse_ret) { + riscv_pmu_use_irq =3D false; + return 0; + } + if (pmu_sbi_sse_state_retained(pmu)) { + pr_err("PMU-SSE setup failed with firmware state retained: %d\n", + sse_ret); + return -EUCLEAN; + } + /* Only an explicitly unsupported SSE path proves IRQ fallback is safe. */ + if (sse_ret !=3D -EOPNOTSUPP) { + pr_err("PMU-SSE setup failed: %d\n", sse_ret); + return sse_ret; + } + + irq_ret =3D pmu_sbi_setup_irq(pmu); + if (!irq_ret) { + riscv_pmu_use_irq =3D true; + return 0; + } + return irq_ret; } =20 #ifdef CONFIG_CPU_PM -static int riscv_pm_pmu_notify(struct notifier_block *b, unsigned long cmd, - void *v) +static int riscv_pm_pmu_update(struct riscv_pmu *rvpmu, unsigned long cmd) { - struct riscv_pmu *rvpmu =3D container_of(b, struct riscv_pmu, riscv_pm_nb= ); struct cpu_hw_events *cpuc =3D this_cpu_ptr(rvpmu->hw_events); bool enabled =3D !bitmap_empty(cpuc->used_hw_ctrs, RISCV_MAX_COUNTERS); struct perf_event *event; int idx; =20 + if (cmd =3D=3D CPU_PM_ENTER) + bitmap_zero(cpuc->pm_resume_hw_ctrs, RISCV_MAX_COUNTERS); + if (!enabled) return NOTIFY_OK; =20 @@ -1302,6 +1597,8 @@ static int riscv_pm_pmu_notify(struct notifier_block = *b, unsigned long cmd, =20 switch (cmd) { case CPU_PM_ENTER: + if (!(event->hw.state & PERF_HES_STOPPED)) + set_bit(idx, cpuc->pm_resume_hw_ctrs); /* * Stop and update the counter */ @@ -1309,6 +1606,8 @@ static int riscv_pm_pmu_notify(struct notifier_block = *b, unsigned long cmd, break; case CPU_PM_EXIT: case CPU_PM_ENTER_FAILED: + if (!test_and_clear_bit(idx, cpuc->pm_resume_hw_ctrs)) + break; /* * Restore and enable the counter. */ @@ -1322,9 +1621,59 @@ static int riscv_pm_pmu_notify(struct notifier_block= *b, unsigned long cmd, return NOTIFY_OK; } =20 +static int riscv_pm_pmu_notify(struct notifier_block *b, + unsigned long cmd, void *v) +{ + struct riscv_pmu *rvpmu =3D container_of(b, struct riscv_pmu, + riscv_pm_nb); + +#ifdef CONFIG_RISCV_PMU_SBI_SSE + struct cpu_hw_events *cpuc =3D this_cpu_ptr(rvpmu->hw_events); + int ret; + + if (!riscv_pmu_use_irq && READ_ONCE(rvpmu->sse_active)) { + switch (cmd) { + case CPU_PM_ENTER: + cpuc->pm_resume_sse =3D + sse_event_is_enabled_local(rvpmu->sse_evt); + if (cpuc->pm_resume_sse) { + ret =3D sse_event_disable_local(rvpmu->sse_evt); + if (ret) { + cpuc->pm_resume_sse =3D false; + pmu_sbi_fail_sse(rvpmu, "disable for CPU PM", + ret); + return notifier_from_errno(ret); + } + } + break; + case CPU_PM_EXIT: + case CPU_PM_ENTER_FAILED: + ret =3D riscv_pm_pmu_update(rvpmu, cmd); + if (!cpuc->pm_resume_sse) + return ret; + + cpuc->pm_resume_sse =3D false; + ret =3D sse_event_enable_local(rvpmu->sse_evt); + if (ret) { + pmu_sbi_fail_sse(rvpmu, "enable after CPU PM", ret); + return notifier_from_errno(ret); + } + return NOTIFY_OK; + default: + break; + } + } +#endif + + return riscv_pm_pmu_update(rvpmu, cmd); +} + static int riscv_pm_pmu_register(struct riscv_pmu *pmu) { pmu->riscv_pm_nb.notifier_call =3D riscv_pm_pmu_notify; + /* Keep PMU-SSE disabled until counters and userpage state are restored. = */ + pmu->riscv_pm_nb.priority =3D riscv_pmu_use_irq ? 1 : -1; + return cpu_pm_register_notifier(&pmu->riscv_pm_nb); } =20 @@ -1339,6 +1688,8 @@ static inline void riscv_pm_pmu_unregister(struct ris= cv_pmu *pmu) { } =20 static void riscv_pmu_destroy(struct riscv_pmu *pmu) { + pmu_sbi_cleanup_sse(pmu); + if (sbi_v2_available) { if (sbi_pmu_snapshot_available()) { pmu_sbi_snapshot_disable(); @@ -1350,7 +1701,7 @@ static void riscv_pmu_destroy(struct riscv_pmu *pmu) cpuhp_state_remove_instance(CPUHP_AP_PERF_RISCV_STARTING, &pmu->node); } =20 -static void pmu_sbi_event_init(struct perf_event *event) +static int pmu_sbi_event_init(struct perf_event *event) { /* * The permissions are set at event_init so that we do not depend @@ -1362,6 +1713,8 @@ static void pmu_sbi_event_init(struct perf_event *eve= nt) event->hw.flags |=3D PERF_EVENT_FLAG_USER_ACCESS; else event->hw.flags |=3D PERF_EVENT_FLAG_LEGACY; + + return 0; } =20 static void pmu_sbi_event_mapped(struct perf_event *event, struct mm_struc= t *mm) @@ -1491,6 +1844,7 @@ static int pmu_sbi_device_probe(struct platform_devic= e *pdev) /* cache all the information about counters now */ if (pmu_sbi_get_ctrinfo(num_counters, cmask)) goto out_free; + bitmap_copy(pmu->cmask, cmask, RISCV_MAX_COUNTERS); =20 ret =3D pmu_sbi_setup_irqs(pmu, pdev); if (ret < 0) { @@ -1500,9 +1854,15 @@ static int pmu_sbi_device_probe(struct platform_devi= ce *pdev) } irq_requested =3D (ret =3D=3D 0); =20 +#ifdef CONFIG_RISCV_PMU_SBI_SSE + if (pmu->sse_active) { + pmu->pmu.pmu_enable =3D pmu_sbi_sse_enable; + pmu->pmu.pmu_disable =3D pmu_sbi_sse_disable; + } +#endif + pmu->pmu.attr_groups =3D riscv_pmu_attr_groups; pmu->pmu.parent =3D &pdev->dev; - bitmap_copy(pmu->cmask, cmask, RISCV_MAX_COUNTERS); pmu->ctr_start =3D pmu_sbi_ctr_start; pmu->ctr_stop =3D pmu_sbi_ctr_stop; pmu->event_map =3D pmu_sbi_event_map; diff --git a/include/linux/perf/riscv_pmu.h b/include/linux/perf/riscv_pmu.h index ecaa40370830..b06abc705100 100644 --- a/include/linux/perf/riscv_pmu.h +++ b/include/linux/perf/riscv_pmu.h @@ -28,6 +28,8 @@ =20 #define RISCV_PMU_CONFIG1_GUEST_EVENTS 0x1 =20 +struct sse_event; + struct cpu_hw_events { /* currently enabled events */ int n_events; @@ -39,6 +41,18 @@ struct cpu_hw_events { DECLARE_BITMAP(used_hw_ctrs, RISCV_MAX_COUNTERS); /* currently enabled firmware counters */ DECLARE_BITMAP(used_fw_ctrs, RISCV_MAX_COUNTERS); +#ifdef CONFIG_RISCV_PMU_SBI_SSE + /* Keep counters stopped after an unrecoverable SSE transition failure. */ + bool sse_failed; +#endif +#ifdef CONFIG_CPU_PM + /* Counters stopped by CPU PM and still waiting to be restored. */ + DECLARE_BITMAP(pm_resume_hw_ctrs, RISCV_MAX_COUNTERS); +#ifdef CONFIG_RISCV_PMU_SBI_SSE + /* Restore the local PMU SSE event after counters and userpage state. */ + bool pm_resume_sse; +#endif +#endif /* The virtual address of the shared memory where counter snapshot will b= e taken */ void *snapshot_addr; /* The physical address of the shared memory where counter snapshot will = be taken */ @@ -54,6 +68,10 @@ struct riscv_pmu { char *name; =20 irqreturn_t (*handle_irq)(int irq_num, void *dev); +#ifdef CONFIG_RISCV_PMU_SBI_SSE + struct sse_event *sse_evt; + bool sse_active; +#endif =20 DECLARE_BITMAP(cmask, RISCV_MAX_COUNTERS); u64 (*ctr_read)(struct perf_event *event); @@ -63,7 +81,7 @@ struct riscv_pmu { void (*ctr_start)(struct perf_event *event, u64 init_val); void (*ctr_stop)(struct perf_event *event, unsigned long flag); int (*event_map)(struct perf_event *event, u64 *config); - void (*event_init)(struct perf_event *event); + int (*event_init)(struct perf_event *event); void (*event_mapped)(struct perf_event *event, struct mm_struct *mm); void (*event_unmapped)(struct perf_event *event, struct mm_struct *mm); uint8_t (*csr_index)(struct perf_event *event); diff --git a/include/linux/riscv_sbi_sse.h b/include/linux/riscv_sbi_sse.h index 774a782e556d..f6b3881c8884 100644 --- a/include/linux/riscv_sbi_sse.h +++ b/include/linux/riscv_sbi_sse.h @@ -42,6 +42,7 @@ int sse_event_enable(struct sse_event *sse_evt); int sse_event_disable(struct sse_event *sse_evt); =20 /* Local events require the caller to remain on the current CPU. */ +bool sse_event_is_enabled_local(struct sse_event *sse_evt); int sse_event_enable_local(struct sse_event *sse_evt); int sse_event_disable_local(struct sse_event *sse_evt); =20 @@ -76,6 +77,11 @@ static inline int sse_event_disable(struct sse_event *ss= e_evt) return -EOPNOTSUPP; } =20 +static inline bool sse_event_is_enabled_local(struct sse_event *sse_evt) +{ + return false; +} + static inline int sse_event_enable_local(struct sse_event *sse_evt) { return -EOPNOTSUPP; --=20 2.50.1 (Apple Git-155) From nobody Fri Sep 25 21:02:40 2026 Received: from mail-pl1-f169.google.com (mail-pl1-f169.google.com [209.85.214.169]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1E8A648A8BE for ; Mon, 21 Sep 2026 11:18:45 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.169 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789989529; cv=none; b=ahcCtGiJeMkJW0GjEwbwTVcFa+YrB8qQ2nNHWkV0tmF9IbE4gClAFYEHO1oiIngZYH3Mg256PCHcp4KuJQ5ViAKROaDbOoPxkPn5ayXWlO1/HlQM+EtUkRIrTUIAHFBRJKCEwoKGWY109QGUJSUo1s3YU9trrORPTVrITJNK07E= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789989529; c=relaxed/simple; bh=bgt1nLiNXMSrRAilvyjGW4f364m8jNeUp5jALakBMhQ=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=TI8ng++3WxTK/rGwMUYOFGyJ2pwed9VT94iL+ARoVjXr/2FAHkWvNgGn0DP/Xl+p/YTkVPH07MBI0buzS3WRwpmYfMu+cbWBFgV5ouOYVvh9gN6d9KifnOKsVuTMVDZlebY9O13p8AjAEs6/8Mvi4H0P44VVcHhi/DwSxrbaKfU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com; spf=pass smtp.mailfrom=bytedance.com; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b=i1Veb3sp; arc=none smtp.client-ip=209.85.214.169 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=bytedance.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b="i1Veb3sp" Received: by mail-pl1-f169.google.com with SMTP id d9443c01a7336-2caced6038eso24883835ad.0 for ; Mon, 21 Sep 2026 04:18:45 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=bytedance.com; s=google; t=1789989525; x=1790594325; darn=vger.kernel.org; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:from:to:cc:subject :date:message-id:reply-to:content-type; bh=1kyfwVGJJ+MGlp30hz4Tcdovm8ViU3NJsmQbY3BeN+A=; b=i1Veb3spC+wbXfUUONNYMD1T8Wj1/dXbVaTWK3p/fxVYXU4WIK/f9bREC8VRIeXcRr z6GiZTBYxkU/FT4JnLiAd2r/fJWrs0XGIuF82XKqBNNpnkL1vGeh+Mo68aiiMxvZ5aSI XRCuFRSFFNPXgY0Q61E7I41WvxxGv8sHNTNabGkadVX/VIOKwhsgePi+eecgqyb9qoqs HITIhXWP2oOcxC2VAFk7AoMLvNXdrawVERs7swSZxiHkVDPoygTVqKqv06QMv9GiyhnR XPORhH8jahA8XD4MiiJS1ok8ui9OhAKhLnBbfmdo9BeX8WNOWvx7GqjzOSHUnWAIFnp0 EGZQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789989525; x=1790594325; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:date:subject:cc:to:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=1kyfwVGJJ+MGlp30hz4Tcdovm8ViU3NJsmQbY3BeN+A=; b=TCjHGBkUfxfDvjSJDfIVU0GxiEyb8JaQurT0OHMUdlontdZ3CI4xrtuGSB0tcZBrEn QfAwXoqIQ026zcVMXJ21bcex4GThMCC2G+Moo4jUSa3400KKcvc9kXvEATufz0JR60Cv TXOpBq97dEo3njAzbgpbRl9749dFHzgkw0Z1lQraD9ugD+FhDfzZtJ4eVLoY6kUu6i3W gB8K8XtZTwe2TrYG0O6w4K7QO3VNgLfG2OWRmzIus37pmHkWIKJ7FkAc7cSCY4JcR5W3 BadkVp3FaogAgN5NM2JLYQo5A6eq7DxWABNYLmJyvXBNOp0sYvKiBz9RdEBUH0+BZP47 j7fw== X-Forwarded-Encrypted: i=1; AKwUvBxsHaJ7wJoCBvw9vBvfVHCNRtY91IueAlaSZkBWrXkKRUn2N7PmFcCIeXXuSCZ2VLbDo439vyrYQ9vUs7I=@vger.kernel.org X-Gm-Message-State: AFuF++kDUnFP00NfyOfkyzFC11EPX/fPpU9x0Xzn/uZAeECco9BOmlaR sHwmAR99JPcrGYkrYbMp54NX7Sh/frwY5z25Ptgwo+EiiU+zu8TO1SJ5N7m1XTqykm0= X-Gm-Gg: AYBFou0tfmjM/+9x/FhPYK+V3MLYI4TT5IYIxe01mRWc9jMsY5XOoz1v6rTion2CGSD T65OPVvZJOHdMx+wrI6rjJH9egBlybjcuJd9FwwFbVdtRsA0+nYCmj1vjnPROfITcnGsQUmJ3il 2MO0BPvpI2B7PLuIkHm1IqHTv6/ayqiBkkcyGxKullcagkP5oqnFPWD/f+4Beh0XyKrtYGaM3HG 1Sb/S8nTbiT7xAXXpuP13n853h42DYW2/vTN120wCRuo3D0nLEkdLmxbDQA1KXX0yv5WbmbmWLJ 9a7pKGIySK7GMbP7eU3LhWWZZxinv47Ik7aetBb8yh0W/aRl4JmmfzqVK6B9MxAIJcHPGI2XRQ3 yTF4ri9KAY4bVGQ3vPqjbuqRHl8fspsqS8OQ5cpAh9umEr4xJVLeGNYTX9iQKwsZJkLX58U4qdW AJl7yh/6rn/F0oDwdkZrk0lU6AJSivcUb5qvkZ/kb811ItlKDVq9BptwirK+p5T7E0EYGXP4YkT YtgLiL7orXeAi8ujIQJfJkZkJti8lyJA7yyv0j+5DL7VA== X-Received: by 2002:a17:902:ebc9:b0:2d8:d4d2:d138 with SMTP id d9443c01a7336-2dd9ca272c6mr161124045ad.20.1789989524877; Mon, 21 Sep 2026 04:18:44 -0700 (PDT) Received: from FJ7FR2JRQ3.bytedance.net ([240e:694:e20:401::8]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2ddc17ba651sm31831125ad.53.2026.09.21.04.18.22 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Mon, 21 Sep 2026 04:18:44 -0700 (PDT) From: Zhanpeng Zhang To: Paul Walmsley , Palmer Dabbelt , Albert Ou , Alexandre Ghiti , Himanshu Chauhan , Conor Dooley , Anup Patel Cc: =?UTF-8?q?Cl=C3=A9ment=20L=C3=A9ger?= , Yunhui Cui , Atish Patra , Peter Zijlstra , Ingo Molnar , Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Alexander Shishkin , Jiri Olsa , Ian Rogers , Adrian Hunter , James Clark , Will Deacon , Thomas Gleixner , Jonathan Corbet , Randy Dunlap , Shuah Khan , Shuah Khan , Yuanzhu , Yicong Yang , Susheng Yang , linux-riscv@lists.infradead.org, linux-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-arm-kernel@lists.infradead.org, Zhanpeng Zhang Subject: [PATCH v10 RESEND 8/9] selftests/riscv: add SSE test module Date: Mon, 21 Sep 2026 19:15:05 +0800 Message-ID: X-Mailer: git-send-email 2.50.1 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable From: Cl=C3=A9ment L=C3=A9ger Add an SSE selftest module and runner. Loading the module executes smoke tests for the SSE framework, and the runner reports any error emitted by the module. Add stress=3D{0,1,2} modes to exercise repeated handler entry and completion, single and multiple read-only SBI calls from a handler, and self re-injection. Check the SBI return values so handler execution alone cannot produce a false pass, and avoid touching an unreserved PMU counter owned by another user. Harden the test against false failures and leaks by using TEST_PROGS for the runner, using phys_addr_t for SBI attribute buffers, adding timeouts to busy waits, pinning each priority chain with migrate_disable(), and holding the CPU read lock while a fast-test target is selected, injected, and completed. Unregister all registered events on error, check teardown failures, and let kthread_stop() drive monitor-thread exit. Track handler progress across fast, priority, and stress paths. If firmware reports injectable events but the test cannot acquire or handle any of them, report SKIP instead of silently passing a capability-only run. Signed-off-by: Cl=C3=A9ment L=C3=A9ger Co-developed-by: Himanshu Chauhan Signed-off-by: Himanshu Chauhan Co-developed-by: Zhanpeng Zhang Signed-off-by: Zhanpeng Zhang --- MAINTAINERS | 2 + tools/testing/selftests/riscv/Makefile | 2 +- tools/testing/selftests/riscv/sse/Makefile | 5 + .../selftests/riscv/sse/module/Makefile | 22 + .../riscv/sse/module/riscv_sse_test.c | 1154 +++++++++++++++++ .../selftests/riscv/sse/run_sse_test.sh | 59 + 6 files changed, 1243 insertions(+), 1 deletion(-) create mode 100644 tools/testing/selftests/riscv/sse/Makefile create mode 100644 tools/testing/selftests/riscv/sse/module/Makefile create mode 100644 tools/testing/selftests/riscv/sse/module/riscv_sse_test= .c create mode 100644 tools/testing/selftests/riscv/sse/run_sse_test.sh diff --git a/MAINTAINERS b/MAINTAINERS index a6fa10a0a94d..275e54d995ab 100644 --- a/MAINTAINERS +++ b/MAINTAINERS @@ -23489,6 +23489,7 @@ C: irc://irc.libera.chat/riscv P: Documentation/arch/riscv/patch-acceptance.rst T: git git://git.kernel.org/pub/scm/linux/kernel/git/riscv/linux.git F: arch/riscv/ +F: tools/testing/selftests/riscv/ N: riscv K: riscv =20 @@ -23614,6 +23615,7 @@ F: arch/riscv/kernel/sbi_sse.c F: arch/riscv/kernel/sbi_sse_entry.S F: drivers/firmware/riscv/riscv_sbi_sse.c F: include/linux/riscv_sbi_sse.h +F: tools/testing/selftests/riscv/sse/ =20 RISC-V TENSTORRENT SoC SUPPORT M: Drew Fustini diff --git a/tools/testing/selftests/riscv/Makefile b/tools/testing/selftes= ts/riscv/Makefile index 5671b4405a12..43c7c8f97676 100644 --- a/tools/testing/selftests/riscv/Makefile +++ b/tools/testing/selftests/riscv/Makefile @@ -5,7 +5,7 @@ ARCH ?=3D $(shell uname -m 2>/dev/null || echo not) =20 ifneq (,$(filter $(ARCH),riscv)) -RISCV_SUBTARGETS ?=3D abi hwprobe mm sigreturn vector cfi +RISCV_SUBTARGETS ?=3D abi hwprobe mm sigreturn vector cfi sse else RISCV_SUBTARGETS :=3D endif diff --git a/tools/testing/selftests/riscv/sse/Makefile b/tools/testing/sel= ftests/riscv/sse/Makefile new file mode 100644 index 000000000000..7e2677fdce09 --- /dev/null +++ b/tools/testing/selftests/riscv/sse/Makefile @@ -0,0 +1,5 @@ +TEST_GEN_MODS_DIR :=3D module + +TEST_PROGS :=3D run_sse_test.sh + +include ../../lib.mk diff --git a/tools/testing/selftests/riscv/sse/module/Makefile b/tools/test= ing/selftests/riscv/sse/module/Makefile new file mode 100644 index 000000000000..eac4b1c6228b --- /dev/null +++ b/tools/testing/selftests/riscv/sse/module/Makefile @@ -0,0 +1,22 @@ +ifneq ($(CONFIG_RISCV_SBI_SSE),) +obj-m +=3D riscv_sse_test.o +endif + +ifndef KERNELRELEASE + +TESTMODS_DIR :=3D $(realpath $(dir $(abspath $(lastword $(MAKEFILE_LIST)))= )) +KDIR ?=3D /lib/modules/$(shell uname -r)/build + +# Ensure that KDIR exists, otherwise skip the compilation +modules: +ifneq ("$(wildcard $(KDIR))", "") + $(Q)$(MAKE) -C $(KDIR) modules KBUILD_EXTMOD=3D$(TESTMODS_DIR) +endif + +# Ensure that KDIR exists, otherwise skip the clean target +clean: +ifneq ("$(wildcard $(KDIR))", "") + $(Q)$(MAKE) -C $(KDIR) clean KBUILD_EXTMOD=3D$(TESTMODS_DIR) +endif + +endif diff --git a/tools/testing/selftests/riscv/sse/module/riscv_sse_test.c b/to= ols/testing/selftests/riscv/sse/module/riscv_sse_test.c new file mode 100644 index 000000000000..cc5c2e46f2fd --- /dev/null +++ b/tools/testing/selftests/riscv/sse/module/riscv_sse_test.c @@ -0,0 +1,1154 @@ +// SPDX-License-Identifier: GPL-2.0-or-later +/* + * Copyright (C) 2025 Rivos Inc. + */ + +#define pr_fmt(fmt) "riscv_sse_test: " fmt + +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include +#include + +#define RUN_LOOP_COUNT 1000 +#define SSE_FAILED_PREFIX "FAILED: " +#define SSE_SKIP_PREFIX "SKIP: " +#define STRESS_DURATION_MS 3000 +#define STRESS_INJECT_NS 10000 +#define STRESS_REINJECT_DEPTH 10 +#define sse_err(...) pr_err(SSE_FAILED_PREFIX __VA_ARGS__) +#define sse_skip(...) pr_info(SSE_SKIP_PREFIX __VA_ARGS__) + +enum sse_stress_mode { + SSE_STRESS_OFF, + SSE_STRESS_AFTER_SMOKE, + SSE_STRESS_ONLY, +}; + +static int stress; +module_param(stress, int, 0444); +MODULE_PARM_DESC(stress, "Stress mode: 0=3Doff, 1=3Dafter smoke, 2=3Dstres= s only"); + +static char *run_id =3D "unknown"; +module_param(run_id, charp, 0444); +MODULE_PARM_DESC(run_id, "Unique identifier used to delimit one test run"); + +/* Do not report PASS for a capability-only run that handled no event. */ +static atomic_t sse_test_handler_count =3D ATOMIC_INIT(0); +static bool sse_stress_event_can_inject; + +struct sse_event_desc { + u32 evt_id; + const char *name; + bool can_inject; +}; + +static struct sse_event_desc sse_event_descs[] =3D { + { + .evt_id =3D SBI_SSE_EVENT_LOCAL_HIGH_PRIO_RAS, + .name =3D "local_high_prio_ras", + }, + { + .evt_id =3D SBI_SSE_EVENT_LOCAL_DOUBLE_TRAP, + .name =3D "local_double_trap", + }, + { + .evt_id =3D SBI_SSE_EVENT_GLOBAL_HIGH_PRIO_RAS, + .name =3D "global_high_prio_ras", + }, + { + .evt_id =3D SBI_SSE_EVENT_LOCAL_PMU_OVERFLOW, + .name =3D "local_pmu_overflow", + }, + { + .evt_id =3D SBI_SSE_EVENT_LOCAL_LOW_PRIO_RAS, + .name =3D "local_low_prio_ras", + }, + { + .evt_id =3D SBI_SSE_EVENT_GLOBAL_LOW_PRIO_RAS, + .name =3D "global_low_prio_ras", + }, + { + .evt_id =3D SBI_SSE_EVENT_LOCAL_SOFTWARE_INJECTED, + .name =3D "local_software_injected", + }, + { + .evt_id =3D SBI_SSE_EVENT_GLOBAL_SOFTWARE_INJECTED, + .name =3D "global_software_injected", + } +}; + +static DEFINE_MUTEX(sse_test_cleanup_lock); +/* Firmware permits only one registration for each event ID. */ +static struct sse_event *sse_test_cleanup_events[ARRAY_SIZE(sse_event_desc= s)]; + +static void sse_test_cleanup_workfn(struct work_struct *work); +static DECLARE_DELAYED_WORK(sse_test_cleanup_work, sse_test_cleanup_workfn= ); + +static void sse_test_queue_cleanup(struct sse_event *event) +{ + int i, free_slot =3D -1; + + mutex_lock(&sse_test_cleanup_lock); + for (i =3D 0; i < ARRAY_SIZE(sse_test_cleanup_events); i++) { + if (sse_test_cleanup_events[i] =3D=3D event) + goto out_schedule; + if (!sse_test_cleanup_events[i] && free_slot < 0) + free_slot =3D i; + } + + if (WARN_ON_ONCE(free_slot < 0)) + goto out_unlock; + + sse_test_cleanup_events[free_slot] =3D event; + +out_schedule: + mod_delayed_work(system_wq, &sse_test_cleanup_work, + msecs_to_jiffies(100)); +out_unlock: + mutex_unlock(&sse_test_cleanup_lock); +} + +static struct sse_event_desc *sse_get_evt_desc(u32 evt) +{ + int i; + + for (i =3D 0; i < ARRAY_SIZE(sse_event_descs); i++) { + if (sse_event_descs[i].evt_id =3D=3D evt) + return &sse_event_descs[i]; + } + + return NULL; +} + +static const char *sse_evt_name(u32 evt) +{ + struct sse_event_desc *desc =3D sse_get_evt_desc(evt); + + return desc ? desc->name : NULL; +} + +static bool sse_test_can_inject_event(u32 evt) +{ + struct sse_event_desc *desc =3D sse_get_evt_desc(evt); + + return desc ? desc->can_inject : false; +} + +/* + * Firmware can invoke the callback until unregister succeeds. Pin the mod= ule + * so its handler text cannot disappear first. + */ +static struct sse_event *sse_test_event_register(u32 evt, u32 priority, + sse_event_handler_fn *handler, + void *arg) +{ + struct sse_event *event; + + event =3D sse_event_register(evt, priority, handler, arg); + if (!IS_ERR(event)) + __module_get(THIS_MODULE); + + return event; +} + +static int sse_test_event_unregister(struct sse_event *event) +{ + int ret; + + ret =3D sse_event_unregister(event); + if (!ret) + module_put(THIS_MODULE); + else + sse_test_queue_cleanup(event); + + return ret; +} + +static void sse_test_cleanup_workfn(struct work_struct *work) +{ + bool retry =3D false; + int i, ret; + + mutex_lock(&sse_test_cleanup_lock); + for (i =3D 0; i < ARRAY_SIZE(sse_test_cleanup_events); i++) { + struct sse_event *event =3D sse_test_cleanup_events[i]; + + if (!event) + continue; + + ret =3D sse_event_disable(event); + if (!ret) + ret =3D sse_event_unregister(event); + if (ret) { + retry =3D true; + continue; + } + + sse_test_cleanup_events[i] =3D NULL; + module_put(THIS_MODULE); + } + mutex_unlock(&sse_test_cleanup_lock); + + if (retry) + mod_delayed_work(system_wq, &sse_test_cleanup_work, + msecs_to_jiffies(100)); +} + +static struct sbiret sbi_sse_ecall(int fid, unsigned long arg0, unsigned l= ong arg1) +{ + return sbi_ecall(SBI_EXT_SSE, fid, arg0, arg1, 0, 0, 0, 0); +} + +static int sse_event_attr_read(u32 evt, unsigned long attr_id, + unsigned long *attr_buf) +{ + struct sbiret sret; + phys_addr_t phys; + + phys =3D virt_to_phys(attr_buf); + + sret =3D sbi_ecall(SBI_EXT_SSE, SBI_SSE_EVENT_ATTR_READ, evt, attr_id, 1, + (unsigned long)phys, 0, 0); + if (sret.error) + return sbi_err_map_linux_errno(sret.error); + + return 0; +} + +static int sse_event_attr_get(u32 evt, unsigned long attr_id, + unsigned long *val) +{ + unsigned long *attr_buf; + int ret; + + attr_buf =3D kmalloc_obj(*attr_buf, GFP_KERNEL); + if (!attr_buf) + return -ENOMEM; + + ret =3D sse_event_attr_read(evt, attr_id, attr_buf); + if (!ret) + *val =3D *attr_buf; + kfree(attr_buf); + + return ret; +} + +static int sse_test_signal(u32 evt, unsigned int cpu) +{ + unsigned long hart_id =3D cpuid_to_hartid_map(cpu); + struct sbiret ret; + + ret =3D sbi_sse_ecall(SBI_SSE_EVENT_INJECT, evt, hart_id); + if (ret.error) { + sse_err("Failed to signal event %x, error %ld\n", evt, ret.error); + return sbi_err_map_linux_errno(ret.error); + } + + return 0; +} + +static int sse_test_wait_not_running(u32 evt) +{ + unsigned long timeout =3D jiffies + HZ; + unsigned long status; + int ret; + + do { + ret =3D sse_event_attr_get(evt, SBI_SSE_ATTR_STATUS, &status); + if (ret) { + sse_err("Failed to get status for evt %x, error %d\n", evt, ret); + return ret; + } + status &=3D SBI_SSE_ATTR_STATUS_STATE_MASK; + cpu_relax(); + } while (status =3D=3D SBI_SSE_STATE_RUNNING && time_before(jiffies, time= out)); + + if (status =3D=3D SBI_SSE_STATE_RUNNING) { + sse_err("Timed out waiting for event %x to leave RUNNING state\n", evt); + return -ETIMEDOUT; + } + + return 0; +} + +struct sse_test_wait_status { + u32 evt; + unsigned long *attr_buf; + unsigned long status; + int ret; +}; + +static void sse_test_read_status_local(void *info) +{ + struct sse_test_wait_status *wait =3D info; + + wait->ret =3D sse_event_attr_read(wait->evt, SBI_SSE_ATTR_STATUS, + wait->attr_buf); + if (!wait->ret) + wait->status =3D *wait->attr_buf & SBI_SSE_ATTR_STATUS_STATE_MASK; +} + +static int sse_test_wait_not_running_on_cpu(u32 evt, unsigned int cpu) +{ + struct sse_test_wait_status wait =3D { .evt =3D evt }; + unsigned long timeout; + int ret =3D 0; + + if (sse_event_is_global(evt)) + return sse_test_wait_not_running(evt); + + wait.attr_buf =3D kmalloc_obj(*wait.attr_buf, GFP_KERNEL); + if (!wait.attr_buf) + return -ENOMEM; + + timeout =3D jiffies + HZ; + do { + ret =3D smp_call_function_single(cpu, sse_test_read_status_local, + &wait, true); + if (ret || wait.ret) { + ret =3D ret ?: wait.ret; + break; + } + if (wait.status !=3D SBI_SSE_STATE_RUNNING) + break; + usleep_range(100, 200); + } while (time_before(jiffies, timeout)); + + if (!ret && wait.status =3D=3D SBI_SSE_STATE_RUNNING) { + sse_err("Timed out waiting for event %x on CPU %u\n", evt, cpu); + ret =3D -ETIMEDOUT; + } + + kfree(wait.attr_buf); + + return ret; +} + +static int sse_test_inject_event(struct sse_event *event, u32 evt, unsigne= d int cpu) +{ + int res; + + if (sse_event_is_global(evt)) { + /* + * Due to the fact the completion might happen faster than + * the call to SBI_SSE_COMPLETE in the handler, if the event was + * running on another CPU, we need to wait for the event status + * to be !RUNNING. + */ + res =3D sse_test_wait_not_running(evt); + if (res) + return res; + + res =3D sse_event_set_target_cpu(event, cpu); + if (res) { + sse_err("Failed to set cpu for evt %x, error %d\n", evt, res); + return res; + } + } + + return sse_test_signal(evt, cpu); +} + +struct fast_test_arg { + u32 evt; + int cpu; + bool args_ready; + bool completion; +}; + +/* A failed unregister may leave firmware holding this handler argument. */ +static struct fast_test_arg fast_test_arg; + +static int sse_test_handler(u32 evt, void *arg, struct pt_regs *regs) +{ + int ret =3D 0; + struct fast_test_arg *targ =3D arg; + u32 test_evt; + int cpu; + + atomic_inc(&sse_test_handler_count); + + /* Pairs with the argument publication in sse_run_fast_test_cpu(). */ + if (!smp_load_acquire(&targ->args_ready)) { + sse_err("Received SSE event %x before its test arguments were published\= n", + evt); + ret =3D -EINVAL; + goto complete; + } + + test_evt =3D READ_ONCE(targ->evt); + cpu =3D READ_ONCE(targ->cpu); + + if (evt !=3D test_evt) { + sse_err("Received SSE event id %x instead of %x\n", test_evt, evt); + ret =3D -EINVAL; + } + + if (!sse_event_is_global(evt) && cpu !=3D smp_processor_id()) { + sse_err("Received SSE event %d on CPU %d instead of %d\n", evt, smp_proc= essor_id(), + cpu); + ret =3D -EINVAL; + } + +complete: + WRITE_ONCE(targ->args_ready, false); + /* Publish handler-side checks before waking the waiting CPU. */ + smp_store_release(&targ->completion, true); + + return ret; +} + +static int sse_run_fast_test_cpu(struct fast_test_arg *test_arg, + struct sse_event *event, u32 evt, int cpu) +{ + unsigned long timeout; + int ret; + + WRITE_ONCE(test_arg->completion, false); + WRITE_ONCE(test_arg->args_ready, false); + WRITE_ONCE(test_arg->evt, evt); + WRITE_ONCE(test_arg->cpu, cpu); + /* Publish all arguments before firmware can inject on another hart. */ + smp_store_release(&test_arg->args_ready, true); + + ret =3D sse_test_inject_event(event, evt, cpu); + if (ret) { + sse_err("event %s injection failed, err %d\n", + sse_evt_name(evt), ret); + return ret; + } + + timeout =3D jiffies + HZ / 100; + /* We can not use since they are not NMI safe */ + /* Pairs with the handler's completion publication. */ + while (!smp_load_acquire(&test_arg->completion) && + time_before(jiffies, timeout)) + cpu_relax(); + /* Acquire the handler's checks even if the loop observed a timeout. */ + if (!smp_load_acquire(&test_arg->completion)) { + sse_err("Failed to wait for event %s completion on CPU %d\n", + sse_evt_name(evt), cpu); + return -ETIMEDOUT; + } + + return sse_test_wait_not_running_on_cpu(evt, cpu); +} + +static void sse_run_fast_test(struct fast_test_arg *test_arg, + struct sse_event *event, u32 evt) +{ + int cpu; + + if (sse_event_is_global(evt)) { + /* Keep the selected target online through injection and completion. */ + cpu_hotplug_disable(); + for_each_online_cpu(cpu) { + if (sse_run_fast_test_cpu(test_arg, event, evt, cpu)) + break; + } + cpu_hotplug_enable(); + return; + } + + guard(cpus_read_lock)(); + for_each_online_cpu(cpu) { + if (sse_run_fast_test_cpu(test_arg, event, evt, cpu)) + return; + } +} + +static void sse_test_injection_fast(void) +{ + int i, ret =3D 0, j; + u32 evt; + struct sse_event *event; + + pr_info("Starting SSE test (fast)\n"); + + for (i =3D 0; i < ARRAY_SIZE(sse_event_descs); i++) { + evt =3D sse_event_descs[i].evt_id; + WRITE_ONCE(fast_test_arg.evt, evt); + WRITE_ONCE(fast_test_arg.cpu, -1); + WRITE_ONCE(fast_test_arg.args_ready, false); + WRITE_ONCE(fast_test_arg.completion, false); + + if (!sse_event_descs[i].can_inject) + continue; + + event =3D sse_test_event_register(evt, 0, sse_test_handler, + (void *)&fast_test_arg); + if (IS_ERR(event)) { + if (PTR_ERR(event) =3D=3D -EEXIST) { + pr_info("Event %s already registered, skipping\n", + sse_evt_name(evt)); + continue; + } + sse_err("Failed to register event %s, err %ld\n", sse_evt_name(evt), + PTR_ERR(event)); + continue; + } + + ret =3D sse_event_enable(event); + if (ret) { + sse_err("Failed to enable event %s, err %d\n", sse_evt_name(evt), ret); + goto err_disable; + } + + pr_info("Starting testing event %s\n", sse_evt_name(evt)); + + for (j =3D 0; j < RUN_LOOP_COUNT; j++) + sse_run_fast_test(&fast_test_arg, event, evt); + pr_info("Finished testing event %s\n", sse_evt_name(evt)); + +err_disable: + ret =3D sse_event_disable(event); + if (ret) + sse_err("Failed to disable event %s, err %d\n", + sse_evt_name(evt), ret); + ret =3D sse_test_event_unregister(event); + if (ret) { + sse_err("Failed to unregister event %s, err %d\n", + sse_evt_name(evt), ret); + return; + } + } + pr_info("Finished SSE test (fast)\n"); +} + +struct priority_test_arg { + unsigned long evt; + struct sse_event *event; + bool called; + bool enable_attempted; + u32 prio; + struct priority_test_arg *next_evt_arg; + void (*check_func)(struct priority_test_arg *arg); +}; + +/* A failed unregister may leave firmware holding these handler arguments.= */ +static struct priority_test_arg default_hi_prio_args[] =3D { + { .evt =3D SBI_SSE_EVENT_GLOBAL_SOFTWARE_INJECTED }, + { .evt =3D SBI_SSE_EVENT_LOCAL_SOFTWARE_INJECTED }, + { .evt =3D SBI_SSE_EVENT_GLOBAL_LOW_PRIO_RAS }, + { .evt =3D SBI_SSE_EVENT_LOCAL_LOW_PRIO_RAS }, + { .evt =3D SBI_SSE_EVENT_LOCAL_PMU_OVERFLOW }, + { .evt =3D SBI_SSE_EVENT_GLOBAL_HIGH_PRIO_RAS }, + { .evt =3D SBI_SSE_EVENT_LOCAL_DOUBLE_TRAP }, + { .evt =3D SBI_SSE_EVENT_LOCAL_HIGH_PRIO_RAS }, +}; + +static struct priority_test_arg default_low_prio_args[] =3D { + { .evt =3D SBI_SSE_EVENT_LOCAL_HIGH_PRIO_RAS }, + { .evt =3D SBI_SSE_EVENT_LOCAL_DOUBLE_TRAP }, + { .evt =3D SBI_SSE_EVENT_GLOBAL_HIGH_PRIO_RAS }, + { .evt =3D SBI_SSE_EVENT_LOCAL_PMU_OVERFLOW }, + { .evt =3D SBI_SSE_EVENT_LOCAL_LOW_PRIO_RAS }, + { .evt =3D SBI_SSE_EVENT_GLOBAL_LOW_PRIO_RAS }, + { .evt =3D SBI_SSE_EVENT_LOCAL_SOFTWARE_INJECTED }, + { .evt =3D SBI_SSE_EVENT_GLOBAL_SOFTWARE_INJECTED }, +}; + +static struct priority_test_arg set_prio_args[] =3D { + { .evt =3D SBI_SSE_EVENT_GLOBAL_SOFTWARE_INJECTED, .prio =3D 5 }, + { .evt =3D SBI_SSE_EVENT_LOCAL_SOFTWARE_INJECTED, .prio =3D 10 }, + { .evt =3D SBI_SSE_EVENT_GLOBAL_LOW_PRIO_RAS, .prio =3D 15 }, + { .evt =3D SBI_SSE_EVENT_LOCAL_LOW_PRIO_RAS, .prio =3D 20 }, + { .evt =3D SBI_SSE_EVENT_LOCAL_PMU_OVERFLOW, .prio =3D 25 }, + { .evt =3D SBI_SSE_EVENT_GLOBAL_HIGH_PRIO_RAS, .prio =3D 30 }, + { .evt =3D SBI_SSE_EVENT_LOCAL_DOUBLE_TRAP, .prio =3D 35 }, + { .evt =3D SBI_SSE_EVENT_LOCAL_HIGH_PRIO_RAS, .prio =3D 40 }, +}; + +static struct priority_test_arg same_prio_args[] =3D { + { .evt =3D SBI_SSE_EVENT_LOCAL_PMU_OVERFLOW, .prio =3D 0 }, + { .evt =3D SBI_SSE_EVENT_LOCAL_HIGH_PRIO_RAS, .prio =3D 10 }, + { .evt =3D SBI_SSE_EVENT_LOCAL_SOFTWARE_INJECTED, .prio =3D 10 }, + { .evt =3D SBI_SSE_EVENT_GLOBAL_SOFTWARE_INJECTED, .prio =3D 10 }, + { .evt =3D SBI_SSE_EVENT_GLOBAL_HIGH_PRIO_RAS, .prio =3D 20 }, +}; + +static int sse_hi_priority_test_handler(u32 evt, void *arg, + struct pt_regs *regs) +{ + struct priority_test_arg *targ =3D arg; + struct priority_test_arg *next =3D READ_ONCE(targ->next_evt_arg); + + atomic_inc(&sse_test_handler_count); + WRITE_ONCE(targ->called, 1); + + if (next) { + sse_test_signal(next->evt, smp_processor_id()); + if (!READ_ONCE(next->called)) { + sse_err("Higher priority event %s was not handled %s\n", + sse_evt_name(next->evt), sse_evt_name(evt)); + } + } + + return 0; +} + +static int sse_low_priority_test_handler(u32 evt, void *arg, struct pt_reg= s *regs) +{ + struct priority_test_arg *targ =3D arg; + struct priority_test_arg *next =3D READ_ONCE(targ->next_evt_arg); + + atomic_inc(&sse_test_handler_count); + WRITE_ONCE(targ->called, 1); + + if (next) { + sse_test_signal(next->evt, smp_processor_id()); + if (READ_ONCE(next->called)) { + sse_err("Lower priority event %s was handle before %s\n", + sse_evt_name(next->evt), sse_evt_name(evt)); + } + } + + return 0; +} + +static void sse_test_injection_priority_arg(struct priority_test_arg *args= , unsigned int args_size, + sse_event_handler_fn handler, const char *test_name) +{ + unsigned int i; + unsigned long timeout; + int ret; + int target_cpu; + struct sse_event *event; + struct priority_test_arg *arg, *first_arg =3D NULL, *prev_arg =3D NULL; + + pr_info("Starting SSE priority test (%s)\n", test_name); + /* Keep the complete priority chain on one CPU. */ + migrate_disable(); + target_cpu =3D smp_processor_id(); + + for (i =3D 0; i < args_size; i++) { + arg =3D &args[i]; + + if (!sse_test_can_inject_event(arg->evt)) + continue; + + WRITE_ONCE(arg->called, false); + WRITE_ONCE(arg->next_evt_arg, NULL); + WRITE_ONCE(arg->event, NULL); + WRITE_ONCE(arg->enable_attempted, false); + + event =3D sse_test_event_register(arg->evt, arg->prio, handler, + (void *)arg); + if (IS_ERR(event)) { + if (PTR_ERR(event) =3D=3D -EEXIST) { + pr_info("Event %s already registered, skipping\n", + sse_evt_name(arg->evt)); + continue; + } + sse_err("Failed to register event %s, err %ld\n", sse_evt_name(arg->evt= ), + PTR_ERR(event)); + goto release_events; + } + arg->event =3D event; + + if (sse_event_is_global(arg->evt)) { + /* Keep the chain on one stable CPU. */ + ret =3D sse_event_set_target_cpu(event, target_cpu); + if (ret) { + sse_err("Failed to set event %s target CPU, err %d\n", + sse_evt_name(arg->evt), ret); + goto release_events; + } + } + + WRITE_ONCE(arg->enable_attempted, true); + ret =3D sse_event_enable(event); + if (ret) { + sse_err("Failed to enable event %s, err %d\n", sse_evt_name(arg->evt), = ret); + goto release_events; + } + + if (prev_arg) + WRITE_ONCE(prev_arg->next_evt_arg, arg); + + prev_arg =3D arg; + + if (!first_arg) + first_arg =3D arg; + } + + if (!first_arg) { + pr_info("No injectable event available for %s priority test\n", + test_name); + goto out; + } + + /* Inject first event, handler should trigger the others in chain. */ + ret =3D sse_test_inject_event(first_arg->event, first_arg->evt, target_cp= u); + if (ret) { + sse_err("SSE event %s injection failed\n", sse_evt_name(first_arg->evt)); + goto release_events; + } + + /* Lower-priority events run after the handler that injected them complet= es. */ + arg =3D first_arg; + while (arg) { + timeout =3D jiffies + HZ; + while (!READ_ONCE(arg->called) && time_before(jiffies, timeout)) + cpu_relax(); + + if (!READ_ONCE(arg->called)) { + sse_err("Event %s handler was not called\n", + sse_evt_name(arg->evt)); + ret =3D -EINVAL; + } + + event =3D arg->event; + arg =3D READ_ONCE(arg->next_evt_arg); + } + +release_events: + + for (i =3D 0; i < args_size; i++) { + arg =3D &args[i]; + event =3D arg->event; + if (!event) + continue; + + ret =3D sse_test_wait_not_running_on_cpu(arg->evt, target_cpu); + if (ret) + sse_err("Event %s did not complete, err %d\n", + sse_evt_name(arg->evt), ret); + + if (arg->enable_attempted) { + ret =3D sse_event_disable(event); + if (ret) { + sse_err("Failed to disable event %s, err %d\n", + sse_evt_name(arg->evt), ret); + sse_test_queue_cleanup(event); + WRITE_ONCE(arg->event, NULL); + WRITE_ONCE(arg->enable_attempted, false); + continue; + } + } + + ret =3D sse_test_event_unregister(event); + if (ret) { + sse_err("Failed to unregister event %s, err %d\n", + sse_evt_name(arg->evt), ret); + continue; + } + + WRITE_ONCE(arg->event, NULL); + WRITE_ONCE(arg->enable_attempted, false); + } + + pr_info("Finished SSE priority test (%s)\n", test_name); +out: + migrate_enable(); +} + +static void sse_test_injection_priority(void) +{ + sse_test_injection_priority_arg(default_hi_prio_args, ARRAY_SIZE(default_= hi_prio_args), + sse_hi_priority_test_handler, "high"); + + sse_test_injection_priority_arg(default_low_prio_args, ARRAY_SIZE(default= _low_prio_args), + sse_low_priority_test_handler, "low"); + + sse_test_injection_priority_arg(set_prio_args, ARRAY_SIZE(set_prio_args), + sse_low_priority_test_handler, "set"); + + sse_test_injection_priority_arg(same_prio_args, ARRAY_SIZE(same_prio_args= ), + sse_low_priority_test_handler, "same_prio_args"); +} + +static int sse_get_inject_status(u32 evt, bool *can_inject) +{ + int ret; + unsigned long val; + + /* Check if injection is supported */ + ret =3D sse_event_attr_get(evt, SBI_SSE_ATTR_STATUS, &val); + if (ret =3D=3D sbi_err_map_linux_errno(SBI_ERR_NOT_SUPPORTED) || + ret =3D=3D sbi_err_map_linux_errno(SBI_ERR_INVALID_PARAM)) { + *can_inject =3D false; + return 0; + } + if (ret) + return ret; + + *can_inject =3D !!(val & BIT(SBI_SSE_ATTR_STATUS_INJECT_OFFSET)); + + return 0; +} + +static int sse_init_events(void) +{ + int i, injectable =3D 0, ret; + + for (i =3D 0; i < ARRAY_SIZE(sse_event_descs); i++) { + struct sse_event_desc *desc =3D &sse_event_descs[i]; + + ret =3D sse_get_inject_status(desc->evt_id, &desc->can_inject); + if (ret) { + sse_err("Failed to read injection status for %s, err %d\n", + desc->name, ret); + return ret; + } + + if (desc->can_inject) + injectable++; + else + pr_info("Can not inject event %s, tests using this event will be skippe= d\n", + desc->name); + + if (desc->evt_id =3D=3D SBI_SSE_EVENT_LOCAL_SOFTWARE_INJECTED) + sse_stress_event_can_inject =3D desc->can_inject; + } + + return injectable; +} + +struct stress_test_ctx { + struct sse_event *event; + struct hrtimer timer; + struct hrtimer stop_timer; + struct task_struct *monitor_task; + wait_queue_head_t wait_q; + atomic_t inject_count; + atomic_t handler_count; + atomic_t handler_errors; + u32 evt_id; + int layer; + bool running; + bool test_done; +}; + +static struct stress_test_ctx stress_ctx; +static DEFINE_PER_CPU(int, stress_reinject_cpu_depth); + +static int stress_handler_empty(u32 evt, void *arg, struct pt_regs *regs) +{ + struct stress_test_ctx *ctx =3D arg; + + atomic_inc(&sse_test_handler_count); + atomic_inc(&ctx->handler_count); + + return 0; +} + +static int stress_handler_ecall(u32 evt, void *arg, struct pt_regs *regs) +{ + struct stress_test_ctx *ctx =3D arg; + struct sbiret ret; + + ret =3D sbi_ecall(SBI_EXT_BASE, SBI_EXT_BASE_GET_SPEC_VERSION, + 0, 0, 0, 0, 0, 0); + if (ret.error) + atomic_inc(&ctx->handler_errors); + atomic_inc(&sse_test_handler_count); + atomic_inc(&ctx->handler_count); + + return 0; +} + +static int stress_handler_multi_ecall(u32 evt, void *arg, struct pt_regs *= regs) +{ + struct stress_test_ctx *ctx =3D arg; + struct sbiret ret; + + ret =3D sbi_ecall(SBI_EXT_BASE, SBI_EXT_BASE_GET_SPEC_VERSION, + 0, 0, 0, 0, 0, 0); + if (ret.error) + atomic_inc(&ctx->handler_errors); + ret =3D sbi_ecall(SBI_EXT_BASE, SBI_EXT_BASE_GET_IMP_ID, + 0, 0, 0, 0, 0, 0); + if (ret.error) + atomic_inc(&ctx->handler_errors); + ret =3D sbi_ecall(SBI_EXT_BASE, SBI_EXT_BASE_GET_IMP_VERSION, + 0, 0, 0, 0, 0, 0); + if (ret.error) + atomic_inc(&ctx->handler_errors); + atomic_inc(&sse_test_handler_count); + atomic_inc(&ctx->handler_count); + + return 0; +} + +static int stress_handler_reinject(u32 evt, void *arg, struct pt_regs *reg= s) +{ + struct stress_test_ctx *ctx =3D arg; + int *depth =3D this_cpu_ptr(&stress_reinject_cpu_depth); + + (*depth)++; + if (*depth < STRESS_REINJECT_DEPTH) + sse_test_signal(evt, smp_processor_id()); + else + *depth =3D 0; + + atomic_inc(&sse_test_handler_count); + atomic_inc(&ctx->handler_count); + + return 0; +} + +static sse_event_handler_fn *stress_handlers[] =3D { + stress_handler_empty, + stress_handler_ecall, + stress_handler_multi_ecall, + stress_handler_reinject, +}; + +static const char * const stress_layer_names[] =3D { + "empty handler", + "single SBI ecall in handler", + "multiple SBI ecalls in handler", + "self re-inject", +}; + +static enum hrtimer_restart stress_timer_callback(struct hrtimer *timer) +{ + struct stress_test_ctx *ctx =3D container_of(timer, struct stress_test_ct= x, timer); + + if (!READ_ONCE(ctx->running)) + return HRTIMER_NORESTART; + + if (!sse_test_signal(ctx->evt_id, smp_processor_id())) + atomic_inc(&ctx->inject_count); + hrtimer_forward_now(timer, ns_to_ktime(STRESS_INJECT_NS)); + + return HRTIMER_RESTART; +} + +static enum hrtimer_restart stress_stop_timer_callback(struct hrtimer *tim= er) +{ + struct stress_test_ctx *ctx; + + ctx =3D container_of(timer, struct stress_test_ctx, stop_timer); + WRITE_ONCE(ctx->test_done, true); + wake_up(&ctx->wait_q); + + return HRTIMER_NORESTART; +} + +static int stress_monitor_thread(void *data) +{ + struct stress_test_ctx *ctx =3D data; + unsigned long last_inject =3D 0, last_handler =3D 0; + + while (!kthread_should_stop()) { + unsigned long inject =3D atomic_read(&ctx->inject_count); + unsigned long handler =3D atomic_read(&ctx->handler_count); + + pr_info("stress layer %d: inject=3D%lu (+%lu), handler=3D%lu (+%lu)\n", + ctx->layer, inject, inject - last_inject, + handler, handler - last_handler); + + last_inject =3D inject; + last_handler =3D handler; + + schedule_timeout_interruptible(HZ); + } + + return 0; +} + +static int sse_stress_test_layer(int layer) +{ + struct sse_event *event; + int inject_count, handler_count; + int ret, target_cpu, unregister_ret; + + if (layer < 0 || layer >=3D ARRAY_SIZE(stress_handlers)) + return -EINVAL; + + pr_info("Starting SSE stress layer %d (%s)\n", + layer, stress_layer_names[layer]); + + memset(&stress_ctx, 0, sizeof(stress_ctx)); + stress_ctx.evt_id =3D SBI_SSE_EVENT_LOCAL_SOFTWARE_INJECTED; + stress_ctx.layer =3D layer; + WRITE_ONCE(stress_ctx.running, true); + atomic_set(&stress_ctx.inject_count, 0); + atomic_set(&stress_ctx.handler_count, 0); + atomic_set(&stress_ctx.handler_errors, 0); + init_waitqueue_head(&stress_ctx.wait_q); + + event =3D sse_test_event_register(stress_ctx.evt_id, 0, + stress_handlers[layer], &stress_ctx); + if (IS_ERR(event)) { + sse_err("Failed to register stress event, err %ld\n", + PTR_ERR(event)); + return PTR_ERR(event); + } + + stress_ctx.event =3D event; + + ret =3D sse_event_enable(event); + if (ret) { + sse_err("Failed to enable stress event, err %d\n", ret); + goto err_disable; + } + + stress_ctx.monitor_task =3D kthread_run(stress_monitor_thread, + &stress_ctx, "sse_stress_mon"); + if (IS_ERR(stress_ctx.monitor_task)) { + ret =3D PTR_ERR(stress_ctx.monitor_task); + sse_err("Failed to create stress monitor thread, err %d\n", ret); + goto err_disable; + } + + /* Keep the pinned timer and its local event on the selected CPU. */ + cpus_read_lock(); + migrate_disable(); + target_cpu =3D smp_processor_id(); + hrtimer_setup(&stress_ctx.timer, stress_timer_callback, + CLOCK_MONOTONIC, HRTIMER_MODE_PINNED); + hrtimer_start(&stress_ctx.timer, ns_to_ktime(STRESS_INJECT_NS), + HRTIMER_MODE_REL_PINNED); + migrate_enable(); + + hrtimer_setup(&stress_ctx.stop_timer, stress_stop_timer_callback, + CLOCK_MONOTONIC, HRTIMER_MODE_REL); + hrtimer_start(&stress_ctx.stop_timer, ms_to_ktime(STRESS_DURATION_MS), + HRTIMER_MODE_REL); + + wait_event(stress_ctx.wait_q, READ_ONCE(stress_ctx.test_done)); + + WRITE_ONCE(stress_ctx.running, false); + hrtimer_cancel(&stress_ctx.timer); + hrtimer_cancel(&stress_ctx.stop_timer); + kthread_stop(stress_ctx.monitor_task); + + pr_info("Finished SSE stress layer %d (%s): inject=3D%d, handler=3D%d\n", + layer, stress_layer_names[layer], + atomic_read(&stress_ctx.inject_count), + atomic_read(&stress_ctx.handler_count)); + + inject_count =3D atomic_read(&stress_ctx.inject_count); + handler_count =3D atomic_read(&stress_ctx.handler_count); + if (!inject_count || !handler_count) { + sse_err("Stress layer %d made no progress: inject=3D%d, handler=3D%d\n", + layer, inject_count, handler_count); + ret =3D -EIO; + } + if (atomic_read(&stress_ctx.handler_errors)) { + sse_err("Stress layer %d observed %d SBI call errors\n", layer, + atomic_read(&stress_ctx.handler_errors)); + ret =3D -EIO; + } + if (sse_test_wait_not_running_on_cpu(stress_ctx.evt_id, target_cpu)) { + sse_err("Stress event did not complete on CPU %d\n", target_cpu); + if (!ret) + ret =3D -ETIMEDOUT; + } + cpus_read_unlock(); + +err_disable: + if (sse_event_disable(event)) { + sse_err("Failed to disable stress event\n"); + if (!ret) + ret =3D -EIO; + } + unregister_ret =3D sse_test_event_unregister(event); + if (unregister_ret) { + sse_err("Failed to unregister stress event\n"); + if (!ret) + ret =3D unregister_ret; + } + stress_ctx.event =3D NULL; + + return ret; +} + +static void sse_stress_test_all_layers(void) +{ + int i, ret; + + pr_info("Starting SSE stress tests: duration=3D%d ms, interval=3D%d ns\n", + STRESS_DURATION_MS, STRESS_INJECT_NS); + + for (i =3D 0; i < ARRAY_SIZE(stress_handlers); i++) { + ret =3D sse_stress_test_layer(i); + if (ret) { + sse_err("Stress layer %d failed, err %d\n", i, ret); + break; + } + + msleep(100); + } + + pr_info("Finished SSE stress tests\n"); +} + +static int __init sse_test_init(void) +{ + int ret; + + pr_info("RUN %s BEGIN\n", run_id); + atomic_set(&sse_test_handler_count, 0); + + if (stress < SSE_STRESS_OFF || stress > SSE_STRESS_ONLY) { + sse_err("Invalid stress mode %d\n", stress); + pr_info("RUN %s END\n", run_id); + return -EINVAL; + } + + ret =3D sse_init_events(); + if (ret < 0) { + pr_info("RUN %s END\n", run_id); + return ret; + } + if (!ret) { + sse_skip("No injectable SSE event is available\n"); + pr_info("RUN %s END\n", run_id); + return 0; + } + if (stress =3D=3D SSE_STRESS_ONLY && !sse_stress_event_can_inject) { + sse_skip("Local software-injected event is unavailable for stress\n"); + pr_info("RUN %s END\n", run_id); + return 0; + } + + if (stress !=3D SSE_STRESS_ONLY) { + sse_test_injection_fast(); + sse_test_injection_priority(); + } + + if (stress =3D=3D SSE_STRESS_AFTER_SMOKE && !sse_stress_event_can_inject) + sse_skip("Local software-injected event is unavailable for stress\n"); + else if (stress !=3D SSE_STRESS_OFF) + sse_stress_test_all_layers(); + if (!atomic_read(&sse_test_handler_count)) + sse_skip("No SSE event was handled\n"); + + pr_info("RUN %s END\n", run_id); + + return 0; +} + +static void __exit sse_test_exit(void) +{ + cancel_delayed_work_sync(&sse_test_cleanup_work); +} + +module_init(sse_test_init); +module_exit(sse_test_exit); + +MODULE_LICENSE("GPL"); +MODULE_AUTHOR("Cl=C3=A9ment L=C3=A9ger "); +MODULE_DESCRIPTION("Test module for SSE"); diff --git a/tools/testing/selftests/riscv/sse/run_sse_test.sh b/tools/test= ing/selftests/riscv/sse/run_sse_test.sh new file mode 100644 index 000000000000..e70a2fd14b05 --- /dev/null +++ b/tools/testing/selftests/riscv/sse/run_sse_test.sh @@ -0,0 +1,59 @@ +#!/bin/bash +# SPDX-License-Identifier: GPL-2.0 +# +# Copyright (C) 2025 Rivos Inc. + +MODULE_NAME=3Driscv_sse_test +DRIVER=3D"./module/${MODULE_NAME}.ko" +ksft_skip=3D4 + +check_test_requirements() +{ + uid=3D$(id -u) + if [ $uid -ne 0 ]; then + echo "$0: Must be run as root" + exit $ksft_skip + fi + + if ! which insmod > /dev/null 2>&1; then + echo "$0: You need insmod installed" + exit $ksft_skip + fi + + if [ ! -f "$DRIVER" ]; then + echo "$0: SSE is disabled or ${MODULE_NAME} is not built" + exit $ksft_skip + fi +} + +check_test_requirements +run_id=3D"$$-$(date +%s)" + +if ! insmod "$DRIVER" run_id=3D"$run_id" "$@" > /dev/null 2>&1; then + echo "${MODULE_NAME}: failed to load, please check dmesg" + exit 1 +fi + +if ! rmmod "$MODULE_NAME"; then + echo "${MODULE_NAME}: failed to unload, please check dmesg" + exit 1 +fi + +run_log=3D$(dmesg | sed -n \ + "/${MODULE_NAME}: RUN ${run_id} BEGIN/,/${MODULE_NAME}: RUN ${run_id} END= /p") +if [ -z "$run_log" ]; then + echo "${MODULE_NAME}: unable to find log for run ${run_id}" + exit 1 +fi + +if echo "$run_log" | grep -q "${MODULE_NAME}: FAILED:"; then + echo "${MODULE_NAME} failed, please check dmesg" + exit 1 +fi + +if echo "$run_log" | grep -q "${MODULE_NAME}: SKIP:"; then + echo "${MODULE_NAME}: no injectable SSE event" + exit $ksft_skip +fi + +exit 0 --=20 2.50.1 (Apple Git-155) From nobody Fri Sep 25 21:02:40 2026 Received: from mail-pj2-f43.google.com (mail-pj2-f43.google.com [74.125.227.171]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A386D488225 for ; Mon, 21 Sep 2026 11:19:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.227.171 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789989552; cv=none; b=tcivlqERAVuhWuwSCcmLi2XiYPBi2FjgYkUCDPhMgIDy5WlDEOq5CbKUcraWtNycQ/4hoPAzaDsg/5z/oZ4+j5lja20C0C57ssKk9vOov8IgLoAeiGv18LoXEGALb3sZIZSq+LN5AjKcwD3HVnQYUdgzSXS0wJlrPnfZrZsWRdE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789989552; c=relaxed/simple; bh=C1sSDxP7g1P+V6PziinqLDO/+YfEMg+uT1/ui3lWsdQ=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=OqO1pLaI3JIdD5iLWf9MPlPB2z+cldUFnnqnESZsClxphNtdcDMuGl1P1ZyygrMbiY8AEbeeP4JhxoUyV1hsOHdLsqw19oXkjbOqzB3Oc5qsLPpIGZTgqV1swEEJ5Q57ckHbEMA5E4Nv2pEplAEk2Mmcfq2h6G07OdHuZ6MXkxI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com; spf=pass smtp.mailfrom=bytedance.com; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b=jOODWQ4V; arc=none smtp.client-ip=74.125.227.171 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=bytedance.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b="jOODWQ4V" Received: by mail-pj2-f43.google.com with SMTP id d9443c01a7336-2d747ec6188so17442025ad.3 for ; Mon, 21 Sep 2026 04:19:09 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=bytedance.com; s=google; t=1789989549; x=1790594349; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=uPqiENyHK/eI6mIkd6EtYTrMG0xLQpwgJVbdDlAHxyc=; b=jOODWQ4VD5SdDQnaNCG4ksIiJ736AADfB4YHnYRaGcLqfhK7IwXTTN+FLBgvdtEVD7 3fRfM2uzHp86ryAuUVOuQSb3WjKgkqHRRpPp7YSLaQEidURZN3YPAhlh8JQ1Gi5YSNJI T90sI5WyppuUzgdRl/JsLXoz5NA5R/CrLluWcjnLI7AvBcVTSkuCOH8gckWrY4nTEcRK zbV6rfA2pe9q/EyS6VgIpXWjmxVkCxZrcoycYkJ9VGxOs/p6OCt7PzFwqLEZinFqvLBM +OWOtVEIJgXeEq51t2bMNVmvTxdILKusHosv5Fa4ht+SQfiDdZQwp62r+VD8+OKd6zwO 5FNA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789989549; x=1790594349; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=uPqiENyHK/eI6mIkd6EtYTrMG0xLQpwgJVbdDlAHxyc=; b=Pn11kLrSudHNYjEGDOX8a+rAMDdn851svwXmDP9l/5HIqkpg/n/+EaylhkEA/Rs+wt /FtBsPqg8cCYSUoW4isgeRDUSCaUTmbwYw+m3Grq0uUh/jt/nyU/jUBoRxtSTODa6euV r9zhxOLSOnnJFvctHFt5mSINRB04uFbWUK7H0m2Qz1HyHSP5JRVk8wnwWlSMCqzN4gGm sPTq4OCyFWRs0KAvwuTuD4f4zM18kaiA2VQKYmzDhorqpXPedPOv1RuKFjaj8ngBSf9V 9hsnanGfvFVx8WMIDS1+SlQdkyOxpfTb4I0zvJsJUib7DdGQCGWF421mgFXOY9dfadCO c5jQ== X-Forwarded-Encrypted: i=1; AKwUvBwbSH9894j4fZl61kA9xty6vnG0Z0rBufpiw8WsBViNMGszfV5YcVDuYICkEfav5hf9Doe7k8I0dhIvgUI=@vger.kernel.org X-Gm-Message-State: AFuF++lvi37cuo9Gd9LHJi9C3w7VzYQbCDLYoqIxeu6A5LngXjCmYjxO QjBcp4ALhOnjysNF0h/s1KMmSypzbPS58SuT1yqViiSLFrxQRet9CBq4P72zb/Miu60= X-Gm-Gg: AYBFou0CpUiNswUbazKQw4L+2x8/qFbuYnlXHbh4ics5ZdM1pvlLqeUjU8ActQaHJV6 Fci6JuqGTPnnm+xC9kmP22w6IRiJ5gtny6zukgObWUugiJsNpqsJP68piRrgkhOA/9mW5HHPOYd 7LBekoyngEaaFkoB0u/aCwBKQ8pLUk9xs3HoSYaxfAR1/tSInO2jXcWhT9cx4oC91cUyx4LvOOJ ka8SUrmIiko3Mg1WuCmzHjYJh4Ey7GituZuJkQdpd6WUysfgkvd+1IKaV6MCcmxs36paM6EitYc ovGyiVYJqJCrGdmD3dAHTmwmXSzCaV3dNr9lT4YIIfl73tDQeNh8xSEB0HTzDgNEWM8Sv7UoNw5 iDW9mfE1UBhuAb+aK2FLke50Ai9NOkS3deSFbN2+Di3667RgoWsnBWI4yOkZ57gZ8yGCY8ZjfHZ GwjI/Li6JmEGpmfp6UTEOVZQJrlnXFR37gFs0MWYc9ypUYyDWrrW1N6Kev291t1ksLnr5os8SXk wCE6WmIzXKD1arAklycfV2G0Aiwr7mWdzLMV2O1K9XGpw== X-Received: by 2002:a17:903:3d8d:b0:2d8:df44:a5c2 with SMTP id d9443c01a7336-2ddb1b99fedmr107644825ad.20.1789989548896; Mon, 21 Sep 2026 04:19:08 -0700 (PDT) Received: from FJ7FR2JRQ3.bytedance.net ([240e:694:e20:401::8]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2ddc17ba651sm31831125ad.53.2026.09.21.04.18.45 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Mon, 21 Sep 2026 04:19:08 -0700 (PDT) From: Zhanpeng Zhang To: Paul Walmsley , Palmer Dabbelt , Albert Ou , Alexandre Ghiti , Himanshu Chauhan , Conor Dooley , Anup Patel Cc: =?UTF-8?q?Cl=C3=A9ment=20L=C3=A9ger?= , Yunhui Cui , Atish Patra , Peter Zijlstra , Ingo Molnar , Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Alexander Shishkin , Jiri Olsa , Ian Rogers , Adrian Hunter , James Clark , Will Deacon , Thomas Gleixner , Jonathan Corbet , Randy Dunlap , Shuah Khan , Shuah Khan , Yuanzhu , Yicong Yang , Susheng Yang , linux-riscv@lists.infradead.org, linux-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-arm-kernel@lists.infradead.org, Zhanpeng Zhang Subject: [PATCH v10 RESEND 9/9] selftests/riscv: add perf user-stack SSE copy regression test Date: Mon, 21 Sep 2026 19:15:06 +0800 Message-ID: <1aea2955f0769b00e99c530565d7acc67e47360e.1789974241.git.zhangzhanpeng.jasper@bytedance.com> X-Mailer: git-send-email 2.50.1 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" On RISC-V, PMU overflow interrupts can be delivered through the SBI Supervisor Software Events (SSE) mechanism. A perf event that samples the raw user stack (PERF_SAMPLE_STACK_USER, as perf record --call-graph dwarf does) then copies a chunk of the interrupted task's user stack from an NMI-like context. If that copy is allowed to take a nested page fault it can corrupt the interrupted task's kernel state and hang or crash the machine under load; this is what riscv_perf_out_copy_user() and the nofault page-fault change fix. The existing SSE selftest module exercises the framework (register, enable, inject, complete, priorities, stress) but never drives the perf user-stack copy that motivated the no-fault path. Add a userspace test that closes that gap: - Open a sampling hardware PMU event with PERF_SAMPLE_STACK_USER over a deep, partially non-resident user stack, drain the ring buffer, and verify every SAMPLE record is well formed and never reports more dumped bytes than were requested. This checks that a non-resident page truncates the dump cleanly instead of faulting or overrunning. - Drive a multi-CPU unix-socket + deep-recursion workload under high-frequency DWARF sampling; the pass criterion is simply that the machine survives, since the original bug took it down. The test reports SKIP when hardware PMU sampling is unavailable or perf_event_paranoid forbids it, so it is safe to run unprivileged or in constrained environments. It is placed under the RISC-V SSE selftests because SSE delivery is the RISC-V-specific condition it protects, and is wired into the sse subtarget Makefile alongside the module runner. Signed-off-by: Zhanpeng Zhang --- tools/testing/selftests/riscv/sse/Makefile | 5 + .../selftests/riscv/sse/sse_perf_ustack.c | 564 ++++++++++++++++++ 2 files changed, 569 insertions(+) create mode 100644 tools/testing/selftests/riscv/sse/sse_perf_ustack.c diff --git a/tools/testing/selftests/riscv/sse/Makefile b/tools/testing/sel= ftests/riscv/sse/Makefile index 7e2677fdce09..646b418de9ff 100644 --- a/tools/testing/selftests/riscv/sse/Makefile +++ b/tools/testing/selftests/riscv/sse/Makefile @@ -1,5 +1,10 @@ +CFLAGS +=3D -I$(top_srcdir)/tools/testing/selftests +LDLIBS +=3D -lpthread + TEST_GEN_MODS_DIR :=3D module =20 +TEST_GEN_PROGS :=3D sse_perf_ustack + TEST_PROGS :=3D run_sse_test.sh =20 include ../../lib.mk diff --git a/tools/testing/selftests/riscv/sse/sse_perf_ustack.c b/tools/te= sting/selftests/riscv/sse/sse_perf_ustack.c new file mode 100644 index 000000000000..9535b6d7ba4c --- /dev/null +++ b/tools/testing/selftests/riscv/sse/sse_perf_ustack.c @@ -0,0 +1,564 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Regression test for the RISC-V perf user-stack copy taken in SSE + * (NMI-like) context. + * + * On RISC-V, PMU overflow interrupts can be delivered through the SBI + * Supervisor Software Events (SSE) mechanism. A perf event that samples t= he + * raw user stack (PERF_SAMPLE_STACK_USER, as perf record --call-graph dwa= rf + * does) then copies a large chunk of the interrupted task's user stack fr= om + * that context. If that copy is allowed to take a nested page fault it can + * corrupt the interrupted task's kernel state and hang or crash the machi= ne + * under load. + * + * This test exercises that exact path: + * - It opens a sampling hardware PMU event with PERF_SAMPLE_STACK_USER. + * - It samples a child running on a controlled user stack followed by an + * inaccessible page, so the copy must truncate at that page boundary. + * - It checks that every user-stack sample record is well formed and th= at + * the dumped size never exceeds the requested size (i.e. the copy sto= ps + * cleanly rather than faulting on). + * - It then drives a multi-threaded unix-socket + deep-recursion worklo= ad + * under high-frequency per-CPU sampling and requires every active sam= pler + * to make progress without taking the machine down. + * + * The test is architecture independent in what it drives; it is placed un= der + * the RISC-V SSE selftests because SSE delivery is the RISC-V-specific + * condition it is meant to protect. + */ +#define _GNU_SOURCE + +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include +#include +#include +#include +#include +#include +#include + +#include "../../kselftest.h" + +#ifndef noinline +#define noinline __attribute__((noinline)) +#endif + +#define STACK_DUMP_SIZE 8192 /* 8 KiB, 8-byte aligned */ +#define RB_DATA_PAGES 64 /* power of two */ +#define SELF_SAMPLE_FREQ 4000 +#define STRESS_SAMPLE_FREQ 5000 +#define STRESS_SECONDS 5 +#define RECURSE_DEPTH 512 +#define TRUNCATION_RUN_MS 250 + +static long page_size; + +static int perf_event_open(struct perf_event_attr *attr, pid_t pid, int cp= u, + int group_fd, unsigned long flags) +{ + return syscall(__NR_perf_event_open, attr, pid, cpu, group_fd, flags); +} + +/* Prevent the compiler from optimizing away a stack buffer. */ +static void keep_alive(void *p) +{ + __asm__ __volatile__("" : : "r"(p) : "memory"); +} + +/* + * Consume a deep user stack and keep it live, so a raw user-stack sample = has + * many pages to copy. Returns a value derived from the stack so the compi= ler + * cannot elide the frames. + */ +static noinline unsigned long burn_stack(int depth, unsigned long *sink) +{ + unsigned long frame[32]; + unsigned int i; + + for (i =3D 0; i < ARRAY_SIZE(frame); i++) + frame[i] =3D (unsigned long)depth * i + *sink; + + if (depth > 0) + frame[depth & 31] +=3D burn_stack(depth - 1, sink); + + for (i =3D 0; i < ARRAY_SIZE(frame); i++) + *sink +=3D frame[i]; + + keep_alive(frame); + return *sink; +} + +static struct perf_event_attr sampling_attr(unsigned long freq) +{ + struct perf_event_attr attr =3D { + .type =3D PERF_TYPE_HARDWARE, + .size =3D sizeof(attr), + .config =3D PERF_COUNT_HW_INSTRUCTIONS, + .sample_type =3D PERF_SAMPLE_STACK_USER, + .sample_stack_user =3D STACK_DUMP_SIZE, + .freq =3D 1, + .sample_freq =3D freq, + .disabled =3D 1, + .exclude_kernel =3D 1, + .exclude_hv =3D 1, + }; + + return attr; +} + +static bool open_skip_reason(int err, const char **why) +{ + switch (err) { + case EACCES: + case EPERM: + *why =3D "insufficient privilege for PMU sampling (perf_event_paranoid)"; + return true; + case ENOENT: + case ENODEV: + case EOPNOTSUPP: + *why =3D "hardware PMU sampling event not available"; + return true; + default: + return false; + } +} + +static bool pmu_sse_route_testable(const char **why) +{ + char *line =3D NULL; + size_t line_size =3D 0; + FILE *interrupts; + bool testable =3D true; + + interrupts =3D fopen("/proc/interrupts", "re"); + if (!interrupts) { + *why =3D "cannot inspect the active PMU delivery route"; + return false; + } + + /* The SBI PMU driver registers this name only for ordinary IRQ delivery.= */ + while (getline(&line, &line_size, interrupts) >=3D 0) { + if (strstr(line, "riscv-pmu")) { + *why =3D "ordinary RISC-V PMU IRQ delivery is active"; + testable =3D false; + break; + } + } + + free(line); + fclose(interrupts); + return testable; +} + +static bool ring_copy_from(void *dst, const void *rb, size_t rb_bytes, + uint64_t pos, size_t size) +{ + size_t offset =3D pos % rb_bytes; + size_t first; + + if (size > rb_bytes) + return false; + + first =3D size < rb_bytes - offset ? size : rb_bytes - offset; + memcpy(dst, (const char *)rb + offset, first); + if (first !=3D size) + memcpy((char *)dst + first, rb, size - first); + + return true; +} + +static int truncation_child(void *arg) +{ + int ready_fd =3D (intptr_t)arg; + char ready =3D 1; + + if (write(ready_fd, &ready, sizeof(ready)) !=3D 1) + return 1; + + for (;;) + __asm__ __volatile__("" : : : "memory"); +} + +static pid_t start_truncation_child(void **stack_mapping) +{ + struct pollfd pfd =3D { .events =3D POLLIN }; + size_t mapping_size =3D 2 * page_size; + char ready; + void *stack; + pid_t pid; + int pipefd[2]; + int saved_errno; + + stack =3D mmap(NULL, mapping_size, PROT_NONE, + MAP_PRIVATE | MAP_ANONYMOUS, -1, 0); + if (stack =3D=3D MAP_FAILED) + return -1; + if (mprotect(stack, page_size, PROT_READ | PROT_WRITE)) + goto err_unmap; + if (pipe(pipefd)) + goto err_unmap; + + /* clone() starts the child below the inaccessible second page. */ + pid =3D clone(truncation_child, (char *)stack + page_size, SIGCHLD, + (void *)(intptr_t)pipefd[1]); + if (pid < 0) + goto err_pipe; + + close(pipefd[1]); + pfd.fd =3D pipefd[0]; + if (poll(&pfd, 1, 1000) !=3D 1 || + read(pipefd[0], &ready, sizeof(ready)) !=3D sizeof(ready)) { + saved_errno =3D ETIMEDOUT; + kill(pid, SIGKILL); + waitpid(pid, NULL, 0); + close(pipefd[0]); + errno =3D saved_errno; + goto err_unmap; + } + close(pipefd[0]); + + *stack_mapping =3D stack; + return pid; + +err_pipe: + saved_errno =3D errno; + close(pipefd[0]); + close(pipefd[1]); + errno =3D saved_errno; +err_unmap: + saved_errno =3D errno; + munmap(stack, mapping_size); + errno =3D saved_errno; + return -1; +} + +static void stop_truncation_child(pid_t pid, void *stack_mapping) +{ + kill(pid, SIGKILL); + while (waitpid(pid, NULL, 0) < 0 && errno =3D=3D EINTR) + ; + munmap(stack_mapping, 2 * page_size); +} + +/* + * Subtest 1: sample a child whose stack is followed by an inaccessible pa= ge. + * Every record must be well formed and at least one stack copy must trunc= ate + * at the controlled page boundary rather than fault or overrun. + */ +static void test_ustack_records_wellformed(void) +{ + struct perf_event_attr attr =3D sampling_attr(SELF_SAMPLE_FREQ); + size_t rb_bytes =3D (size_t)RB_DATA_PAGES * page_size; + struct perf_event_mmap_page *meta; + unsigned long samples =3D 0, truncated =3D 0; + void *child_stack; + const char *why; + void *rb; + pid_t child; + int fd; + + child =3D start_truncation_child(&child_stack); + if (child < 0) { + ksft_test_result_fail("ustack records: create guarded stack child: %s\n", + strerror(errno)); + return; + } + + fd =3D perf_event_open(&attr, child, -1, -1, PERF_FLAG_FD_CLOEXEC); + if (fd < 0) { + if (open_skip_reason(errno, &why)) + ksft_test_result_skip("ustack records: %s\n", why); + else + ksft_test_result_fail("ustack records: perf_event_open: %s\n", + strerror(errno)); + goto out_child; + } + + meta =3D mmap(NULL, page_size + rb_bytes, PROT_READ | PROT_WRITE, + MAP_SHARED, fd, 0); + if (meta =3D=3D MAP_FAILED) { + ksft_test_result_fail("ustack records: mmap ring buffer: %s\n", + strerror(errno)); + close(fd); + goto out_child; + } + rb =3D (char *)meta + page_size; + + ioctl(fd, PERF_EVENT_IOC_RESET, 0); + ioctl(fd, PERF_EVENT_IOC_ENABLE, 0); + usleep(TRUNCATION_RUN_MS * 1000); + ioctl(fd, PERF_EVENT_IOC_DISABLE, 0); + + /* Drain the ring buffer and validate every SAMPLE record. */ + { + uint64_t head =3D __atomic_load_n(&meta->data_head, __ATOMIC_ACQUIRE); + uint64_t tail =3D meta->data_tail; + bool ok =3D true; + + if (head < tail || head - tail > rb_bytes) + ok =3D false; + + while (ok && tail < head) { + struct perf_event_header hdr; + uint64_t available =3D head - tail; + + if (available < sizeof(hdr) || + !ring_copy_from(&hdr, rb, rb_bytes, tail, sizeof(hdr)) || + hdr.size < sizeof(hdr) || hdr.size > available || + hdr.size > rb_bytes) { + ok =3D false; + break; + } + + if (hdr.type =3D=3D PERF_RECORD_SAMPLE) { + uint64_t dump_size, dyn_size; + size_t cursor =3D sizeof(hdr); + + if (sizeof(dump_size) > hdr.size - cursor || + !ring_copy_from(&dump_size, rb, rb_bytes, + tail + cursor, sizeof(dump_size)) || + dump_size > STACK_DUMP_SIZE) { + ok =3D false; + break; + } + cursor +=3D sizeof(dump_size); + samples++; + if (dump_size) { + /* data blob then trailing dynamic size */ + if (dump_size > hdr.size - cursor) { + ok =3D false; + break; + } + cursor +=3D dump_size; + if (sizeof(dyn_size) > hdr.size - cursor || + !ring_copy_from(&dyn_size, rb, rb_bytes, + tail + cursor, + sizeof(dyn_size))) { + ok =3D false; + break; + } + if (dyn_size > dump_size) { + ok =3D false; + break; + } + if (dyn_size < dump_size) + truncated++; + } + } + tail +=3D hdr.size; + } + __atomic_store_n(&meta->data_tail, head, __ATOMIC_RELEASE); + + if (!ok) + ksft_test_result_fail("ustack records: malformed sample record\n"); + else if (samples =3D=3D 0) + ksft_test_result_skip("ustack records: no samples collected\n"); + else if (truncated =3D=3D 0) + ksft_test_result_fail("ustack records: no guarded-stack truncation\n"); + else + ksft_test_result_pass("ustack records: %lu samples, %lu truncated\n", + samples, truncated); + } + + munmap(meta, page_size + rb_bytes); + close(fd); +out_child: + stop_truncation_child(child, child_stack); +} + +/* ---- Subtest 2: multi-threaded per-CPU sampling stress ---- */ + +struct stress_thread { + pthread_t tid; + int cpu; + int *stop; + int fd; + void *rb; + size_t rb_bytes; +}; + +static void *stress_worker(void *arg) +{ + struct stress_thread *st =3D arg; + unsigned long sink =3D 1; + int sv[2]; + char buf[64]; + + if (socketpair(AF_UNIX, SOCK_STREAM, 0, sv) =3D=3D 0) { + while (!__atomic_load_n(st->stop, __ATOMIC_RELAXED)) { + /* unix-socket ping-pong: takes the socket locks the + * original bug corrupted, while sampling nests. + */ + if (write(sv[0], buf, sizeof(buf)) > 0) + (void)read(sv[1], buf, sizeof(buf)); + burn_stack(RECURSE_DEPTH, &sink); + /* Periodically consume the ring buffer so sampling + * keeps delivering rather than filling up and stopping. + */ + if (st->rb) { + struct perf_event_mmap_page *m =3D st->rb; + uint64_t h =3D __atomic_load_n(&m->data_head, + __ATOMIC_ACQUIRE); + __atomic_store_n(&m->data_tail, h, + __ATOMIC_RELEASE); + } + } + close(sv[0]); + close(sv[1]); + } + + return (void *)sink; +} + +static void test_sse_stress_no_crash(void) +{ + struct perf_event_attr attr =3D sampling_attr(STRESS_SAMPLE_FREQ); + size_t rb_bytes =3D (size_t)RB_DATA_PAGES * page_size; + struct stress_thread *threads; + cpu_set_t available; + long progressed =3D 0; + int stop =3D 0; + const char *why =3D NULL; + long started =3D 0; + long nproc; + long slot; + int cpu; + + if (sched_getaffinity(0, sizeof(available), &available)) { + ksft_test_result_fail("sse stress: sched_getaffinity: %s\n", + strerror(errno)); + return; + } + nproc =3D CPU_COUNT(&available); + if (nproc < 1) { + ksft_test_result_skip("sse stress: no available CPUs\n"); + return; + } + + threads =3D calloc(nproc, sizeof(*threads)); + if (!threads) { + ksft_test_result_fail("sse stress: out of memory\n"); + return; + } + + slot =3D 0; + for (cpu =3D 0; cpu < CPU_SETSIZE; cpu++) { + struct stress_thread *st; + pthread_attr_t thread_attr; + cpu_set_t set; + void *map; + int ret; + + if (!CPU_ISSET(cpu, &available)) + continue; + st =3D &threads[slot++]; + + st->fd =3D -1; + st->cpu =3D cpu; + st->stop =3D &stop; + st->rb_bytes =3D rb_bytes; + st->fd =3D perf_event_open(&attr, -1, cpu, -1, + PERF_FLAG_FD_CLOEXEC); + if (st->fd < 0) { + if (!started && open_skip_reason(errno, &why)) + break; + continue; + } + + map =3D mmap(NULL, page_size + rb_bytes, PROT_READ | PROT_WRITE, + MAP_SHARED, st->fd, 0); + if (map =3D=3D MAP_FAILED) { + close(st->fd); + st->fd =3D -1; + continue; + } + st->rb =3D map; + + CPU_ZERO(&set); + CPU_SET(cpu, &set); + pthread_attr_init(&thread_attr); + ret =3D pthread_attr_setaffinity_np(&thread_attr, sizeof(set), &set); + if (!ret) + ret =3D pthread_create(&st->tid, &thread_attr, + stress_worker, st); + pthread_attr_destroy(&thread_attr); + + if (ret) { + munmap(st->rb, page_size + rb_bytes); + close(st->fd); + st->rb =3D NULL; + st->fd =3D -1; + continue; + } + + ioctl(st->fd, PERF_EVENT_IOC_RESET, 0); + ioctl(st->fd, PERF_EVENT_IOC_ENABLE, 0); + started++; + } + + if (started =3D=3D 0) { + free(threads); + if (why) + ksft_test_result_skip("sse stress: %s\n", why); + else + ksft_test_result_skip("sse stress: could not start any sampler\n"); + return; + } + + sleep(STRESS_SECONDS); + __atomic_store_n(&stop, 1, __ATOMIC_RELAXED); + + for (slot =3D 0; slot < nproc; slot++) { + struct stress_thread *st =3D &threads[slot]; + struct perf_event_mmap_page *meta; + + if (st->fd < 0) + continue; + pthread_join(st->tid, NULL); + ioctl(st->fd, PERF_EVENT_IOC_DISABLE, 0); + meta =3D st->rb; + if (__atomic_load_n(&meta->data_head, __ATOMIC_ACQUIRE)) + progressed++; + munmap(st->rb, page_size + rb_bytes); + close(st->fd); + } + + free(threads); + if (progressed !=3D started) + ksft_test_result_fail("sse stress: %ld/%ld samplers made progress\n", + progressed, started); + else + ksft_test_result_pass("sse stress: %ld samplers x %ds made progress\n", + started, STRESS_SECONDS); +} + +int main(void) +{ + const char *why; + + page_size =3D sysconf(_SC_PAGESIZE); + + ksft_print_header(); + ksft_set_plan(2); + if (!pmu_sse_route_testable(&why)) { + ksft_test_result_skip("ustack records: %s\n", why); + ksft_test_result_skip("sse stress: %s\n", why); + ksft_finished(); + } + + test_ustack_records_wellformed(); + test_sse_stress_no_crash(); + + ksft_finished(); +} --=20 2.50.1 (Apple Git-155)