From nobody Fri Sep 25 01:20:33 2026 Received: from CH1PR05CU001.outbound.protection.outlook.com (mail-northcentralusazon11020133.outbound.protection.outlook.com [52.101.193.133]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 607DE49A3AA; Thu, 17 Sep 2026 22:59:37 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.193.133 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789685980; cv=fail; b=hDvaQHXXctqU3v+u9k4I+SWmcgdHnSSaEOZr4nC/hC21o5GxQQF+dKQMWgt1rpK3AUvFYMTvz8GAn+J+ba+aT/wBBlThLEsTNzCxDLq3Yi4Jkc+dfkZMbPT/aAKpLYZ69emShM4u7mPJTsKUR1F7fR+qB6GSwIKbfdcedwhn+Ps= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789685980; c=relaxed/simple; bh=3OFVAQgfwvvfpf4nlrXDFyGecnrLNdgtKUC2v1GeEGk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=As9Dya8sA6IVGLgDiWsMK6XLNYRZy4TwPBVIGg4G1QTe9GYHOZ3fuEr2APv8bK7OlUqQ0vcRLJZfPoIfQaAevTMs3bb9K+qmv6+VFjbfBEGIwfaWMfCj06Lx3x+GsnnXYYG3i15MFSrVCaRX3dcmtoNls3zJ2NwwUyVHSbemXzQ= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com; spf=fail smtp.mailfrom=rambus.com; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b=0w3k9DjE; arc=fail smtp.client-ip=52.101.193.133 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=rambus.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b="0w3k9DjE" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=v8BbifW1DdYYpHepq9LDbW+ckfReTFYmYarjta/BPBx7hCXmhlvxYuiaLS/EVlfAE/F47KgNFEPMUL47TqPED2h8GuoT/OvTSL3Q6zsf/X1+Oir+gdgXtrjwg97EiKqOpNk9lCNZt1za/bGRH6hTOdqwjoFqK1PxB4BQP6H8pKTOnRvJNPhqlQdt3jv3aFTu68N6iGFbFSQ6RjBCfI5hZOhJAviW4UHrY7+XRB4/MA1o1XwJgTytszSZVtC8/iOqUjWi8aUpIhRpw+gwwA7gftujkQO/YhT58V4dWPHCiWsNgP22CNAtMA14EVeUyRZt/0j/yM9ye/Sofz02PUl4Ag== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=oWic5r/uY4GNDLvn/cXWGegwvbClQwzTywLe77NBz08=; b=IIeAbQzMPtESIdoVNo3u9QZNr7kRfJNJHJhnfX/EzofQ+1u/oxF2D47pXMb52D+/0jZs3/V3P58a4mv6A01wGkkitoDtgkiis5I30bo7QRxy2mN6DhmJxIji/Ozzm34O8O4LehMBJaBYUHN0YDZcLbz+uQ4Zg5s/3eQCBI1DBa2lwfQdjLBqZouRAeb8o+bLAaxH8gbT80DRhvz1+T022aKOXg5pJWK0cZIs6Mr37E+0O0a04W20pT81EWa8zO1j83VdjYXo9kRCs9eWiQSbPhR6UvageTtjL/03cRlbJRuAJEQeIibtFMOnHxyXNv8BCkiXaTBb7M2eSmkh6sHrUQ== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 192.86.86.210) smtp.rcpttodomain=cryptography.com smtp.mailfrom=rambus.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=rambus.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=rambus.com; s=selector1; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=oWic5r/uY4GNDLvn/cXWGegwvbClQwzTywLe77NBz08=; b=0w3k9DjEg6juptFBkSlMZR08wgrjX6YqWnNGVC/cMJ1SuA0tFHsAUYBWjyL5BySOT64oTT0zKPTIYXpQz/u/coCDmZFunSlzTA213RefFJMFl1cNLW+xrcbaZolrdJ6SH9HxoqiPswOidYLiGhn1oPop029mkacldMfNGz/xj6U0hrMgwfBpHr51cSjIkLJcmzZj8wDvqZMV84ygbqRmozV+OJobnItfGjm8V8ITPRoRMLzvuhLQxXmhRXeLNtSAkjAmyoVJZdf3HorVqHRhPMRXBqQJjFOdlztpXEG9yqpT0i38r2MSyp9w6o+hSWmVGvBHUAFFRXW5aJxR1MxYbg== Received: from SJ0PR03CA0354.namprd03.prod.outlook.com (2603:10b6:a03:39c::29) by SA5PR04MB9915.namprd04.prod.outlook.com (2603:10b6:806:4d6::18) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.428.13; Thu, 17 Sep 2026 22:59:32 +0000 Received: from SJ1PEPF000026CA.namprd04.prod.outlook.com (2603:10b6:a03:39c:cafe::32) by SJ0PR03CA0354.outlook.office365.com (2603:10b6:a03:39c::29) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.428.12 via Frontend Transport; Thu, 17 Sep 2026 22:59:32 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 192.86.86.210) smtp.mailfrom=rambus.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=rambus.com; Received-SPF: Pass (protection.outlook.com: domain of rambus.com designates 192.86.86.210 as permitted sender) receiver=protection.outlook.com; client-ip=192.86.86.210; helo=hqxsv-psmtppxy02.rambus.com; pr=C Received: from hqxsv-psmtppxy02.rambus.com (192.86.86.210) by SJ1PEPF000026CA.mail.protection.outlook.com (10.167.244.107) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.451.8 via Frontend Transport; Thu, 17 Sep 2026 22:59:31 +0000 Received: from hqxsv-cmdev3-aousherovitch.rambus.com (hqn-lb-int-float.rambus.com [10.12.20.20]) by hqxsv-psmtppxy02.rambus.com (Postfix) with ESMTP id AC28C180174C; Thu, 17 Sep 2026 22:59:31 +0000 (UTC) From: Alex Ousherovitch To: Albert Ou , Alex Ousherovitch , Conor Dooley , "David S. Miller" , Herbert Xu , Jonathan Corbet , Krzysztof Kozlowski , Palmer Dabbelt , Paul Walmsley , Rob Herring , Saravanakrishnan Krishnamoorthy , Shuah Khan Cc: Alexandre Ghiti , devicetree@vger.kernel.org, Joel Wittenauer , linux-api@vger.kernel.org, linux-crypto@vger.kernel.org, linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-riscv@lists.infradead.org, Shuah Khan , Thi Nguyen Subject: [PATCH v5 01/19] dt-bindings: crypto: add Rambus CryptoManager Hub Date: Thu, 17 Sep 2026 15:59:10 -0700 Message-ID: <20260917225929.2494111-2-aousherovitch@rambus.com> X-Mailer: git-send-email 2.43.7 In-Reply-To: <20260917225929.2494111-1-aousherovitch@rambus.com> References: <20260917225929.2494111-1-aousherovitch@rambus.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: SJ1PEPF000026CA:EE_|SA5PR04MB9915:EE_ X-MS-Office365-Filtering-Correlation-Id: 4af533a3-d748-459e-2827-08df150f55d5 X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|23010399003|36860700016|1800799024|376014|7416014|82310400026|10067099003|3023799007|11063799006|6133799003|56012099006|921020|13003099007|18002099003|22082099003; X-Microsoft-Antispam-Message-Info: Jat0p+NpCRzMdkCjVNxLK+x85aVjDctb7n8AHpzI49KAD1c4egz3ICqxUyF3KbmYiMg8gZAeLUoh7BF0b3Jl9SUjJHl/yjc9TpdJL4UQ+zZeTJJ+mibInLBUp5j0rjRdlrAx/KEmKdm25ByOlV+6dUP1ErE+oWO9FuvPRADTdeecf/Tpd7edmXXmHoSt+I1cMy4lPhDfYPaAVoBPt3MfJRK+G4RxZKpZVdEvuZV5OifksAqKkvvm02TU0XFFJBt2m/nEXfM5aiZcmbg0VOLfwPAyKgHzLni0lqPsDzpL08eHVR+eZr7cGfFlqS+C2zw4OnpfJBIlCC2w48ipjavik637iUDnpSE21SQRC9Piy04u+d8ssZbjG1Sba0aeE65sy3+G1sTQf7mIdJ46tz4LoX7NRWEzBX5Ubns2NVOfd9AUiwyrXwrF0csPw+QeIK4eqzv4xPbX3/fE3nnLIhTEMI9Coauhtw6hf5srk9YoFKZQrAsiZ9jiXR5R9eNNuvpwzTA0cdiwLIpl43yJB5wJrVjatbwVtcSyoc7k/8Do/fesgkFH7AU+nP146EI9GWSM3IqFO96rXavGiHM0gc9Hn9tJtMWveDvvq0Ym2GTXUHGA50wRMSGcnfvW6mLu64Iy5pUShkdjRKLgxUKIcvtb1kxI3QtsvWL4RtbZ3zE1Y1Yq7lm3LcWGIZMSlPZatZ8U X-Forefront-Antispam-Report: CIP:192.86.86.210;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:hqxsv-psmtppxy02.rambus.com;PTR:InfoDomainNonexistent;CAT:NONE;SFS:(13230040)(23010399003)(36860700016)(1800799024)(376014)(7416014)(82310400026)(10067099003)(3023799007)(11063799006)(6133799003)(56012099006)(921020)(13003099007)(18002099003)(22082099003);DIR:OUT;SFP:1102; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: jk0xWmwaFE2Ek7/w4QrITOxybJ2q2rqihYlaggrGeNN8y9DVh2QwzOqbc1pj85TSHYLnzhj+psVp2Jnz7w/5zwUP03mMVsYQFTRpHBJUBKCunTzhRc5wda0GiaruW1OnAlUdsTayLMvpGf88GORb1tB7klSO2zMf5YAqgUCR6iXWe7jeAorwdXiiwzZ8jCTtLcidCj1Kfk3Hz9RoI+DitEy/NN1LOc5glpUIIWsJDz40mO+n2rhYOmLixYCCcEXt6jLhN8/wb5rwEUX2kkQoNL913XdD288zx9+VbQJMIpldRlDUQQ6bRI+Y2BpZY060aSH/XOiHQyofKhzfh6WSaMFgG3tMBI0UaDaNh1VaQxAxvFkvdaxuqeENjKyOAcpJj7In9ZFAVmgqvWdRNo6iJhk/YdSOlprfvcIuGpYl3dTP3PVqekwvALeu+bNSsc9A X-OriginatorOrg: rambus.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 17 Sep 2026 22:59:31.8171 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: 4af533a3-d748-459e-2827-08df150f55d5 X-MS-Exchange-CrossTenant-Id: bd0ba799-c2b9-413c-9c56-5d1731c4827c X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=bd0ba799-c2b9-413c-9c56-5d1731c4827c;Ip=[192.86.86.210];Helo=[hqxsv-psmtppxy02.rambus.com] X-MS-Exchange-CrossTenant-AuthSource: SJ1PEPF000026CA.namprd04.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: SA5PR04MB9915 Content-Type: text/plain; charset="utf-8" Add device tree binding schema for the Rambus CryptoManager Hub (CMH) hardware crypto accelerator. The binding describes the parent SoC-level node with its SIC register region and one queue@N child node per mailbox the host owns, each carrying a reg (mailbox instance index), an optional interrupt, VCQ ring geometry (rambus,num-slots / rambus,slot-stride-bytes) and a rambus,cores affinity list. Which crypto cores are present is discovered from the SIC CORE_ENABLE register at probe, not described in the device tree. Register the 'rambus' vendor prefix for Rambus Inc. Signed-off-by: Alex Ousherovitch Co-developed-by: Saravanakrishnan Krishnamoorthy Signed-off-by: Saravanakrishnan Krishnamoorthy --- .../bindings/crypto/rambus,cmh-v1030.yaml | 151 ++++++++++++++++++ .../devicetree/bindings/vendor-prefixes.yaml | 2 + 2 files changed, 153 insertions(+) create mode 100644 Documentation/devicetree/bindings/crypto/rambus,cmh-v10= 30.yaml diff --git a/Documentation/devicetree/bindings/crypto/rambus,cmh-v1030.yaml= b/Documentation/devicetree/bindings/crypto/rambus,cmh-v1030.yaml new file mode 100644 index 000000000000..a0df35bdc371 --- /dev/null +++ b/Documentation/devicetree/bindings/crypto/rambus,cmh-v1030.yaml @@ -0,0 +1,151 @@ +# SPDX-License-Identifier: (GPL-2.0-only OR BSD-2-Clause) +%YAML 1.2 +--- +$id: http://devicetree.org/schemas/crypto/rambus,cmh-v1030.yaml# +$schema: http://devicetree.org/meta-schemas/core.yaml# + +title: Rambus CryptoManager Hub (CMH) Hardware Crypto Accelerator + +maintainers: + - Alex Ousherovitch + - Saravanakrishnan Krishnamoorthy + - Joel Wittenauer + +description: | + The Rambus CryptoManager Hub (CMH) is a hardware cryptographic accelerat= or + accessed via a mailbox-based VCQ (Virtual Command Queue) interface. The + host writes VCQ command sequences into per-mailbox DMA queue buffers and + rings a doorbell; the CMH eSW processes them and signals completion via + interrupt. + + The management host statically partitions the hardware mailboxes across + the SoC's host interfaces at integration time; the set of mailboxes a + given host owns is therefore fixed and not runtime-discoverable (a + mailbox locked to a host reads as unavailable in the SIC availability + register). Each owned mailbox is described by a child node. Which + crypto cores are present is a fixed silicon-build property indicated by + the SIC CORE_ENABLE register, so cores are not described in the device + tree. Clock and reset are owned by the management host; a node + describing a non-management host has no clock or reset provider of its + own, so clocks and reset-gpios are optional. + + CMH gates access to a locked mailbox by a hardware HOST ID presented on + the bus with every access, permitting only the owning host's ID. An + integration must present a single, stable HOST ID for all accesses to a + given mailbox, independent of the issuing CPU (relevant on SMP hosts + whose interconnect encodes the issuing CPU in the HOST ID). + +properties: + compatible: + items: + - not: {} + description: SoC-specific compatible, e.g. vendor,soc-cmh + - const: rambus,cmh-v1030 + + reg: + maxItems: 1 + description: + SIC (System Interface Controller) MMIO region. The registers of + mailbox instance N are at offset N * 0x1000 within this region. + + clocks: + minItems: 1 + maxItems: 3 + description: + The "core" functional clock, and optionally the half-rate "core-div2" + clock (present only on configurations with side-channel-protected + cores) and/or the "rt" real-time tick clock. See clock-names for the + valid combinations. + + clock-names: + oneOf: + - items: + - const: core + - items: + - const: core + - const: core-div2 + - items: + - const: core + - const: rt + - items: + - const: core + - const: core-div2 + - const: rt + + reset-gpios: + maxItems: 1 + description: + Host-controlled reset for the CryptoManager Hub. The hub has two + external, active-low reset inputs -- a power-on reset and a hard + reset; where a board routes one of them to a host GPIO, that line is + described here. + + "#address-cells": + const: 1 + + "#size-cells": + const: 0 + +patternProperties: + "^queue@[0-9a-f]+$": + type: object + description: + One node per hardware mailbox (VCQ command queue) this host owns. + The set of owned mailboxes is fixed by the management host at + integration time and enumerated here. + properties: + reg: + maxItems: 1 + description: + 0-based mailbox instance index. The instance's registers are + at reg * 0x1000 within the SIC region. + + interrupts: + maxItems: 1 + description: Completion/error interrupt for this mailbox. + + rambus,num-slots: + $ref: /schemas/types.yaml#/definitions/uint32 + enum: [2, 4, 8, 16, 32, 64, 128, 256, 512, 1024, 2048, 4096, 8192, + 16384, 32768] + default: 64 + description: + Number of VCQ ring slots for this mailbox's command queue in + host DMA memory. This is a per-board, per-mailbox host-memory + ring geometry -- boards built around the same SoC (hence the + same compatible) may use different ring sizes, so it is + described per mailbox rather than derived from the compatible. + + rambus,slot-stride-bytes: + $ref: /schemas/types.yaml#/definitions/uint32 + enum: [128, 256, 512, 1024] + default: 512 + description: + Stride in bytes between consecutive VCQ ring slots for this + mailbox's command queue. Like rambus,num-slots this is a + per-board host-memory ring geometry, not derived from the + compatible. + + rambus,cores: + $ref: /schemas/types.yaml#/definitions/string-array + items: + enum: [hc, aes, sm4, sm3, hcq, qse, pke, ccp] + description: | + Core-affinity list: the crypto cores whose work is dispatched to + this mailbox. A core may appear on at most one mailbox. Cores + not listed on any mailbox are load-balanced across all mailboxes. + Optional (default: none -- the mailbox only serves the + load-balanced pool). + + required: + - reg + + additionalProperties: false + +required: + - compatible + - reg + - "#address-cells" + - "#size-cells" + +additionalProperties: false diff --git a/Documentation/devicetree/bindings/vendor-prefixes.yaml b/Docum= entation/devicetree/bindings/vendor-prefixes.yaml index ba2002969373..364e53c31045 100644 --- a/Documentation/devicetree/bindings/vendor-prefixes.yaml +++ b/Documentation/devicetree/bindings/vendor-prefixes.yaml @@ -1397,6 +1397,8 @@ patternProperties: description: RaidSonic Technology GmbH "^ralink,.*": description: Mediatek/Ralink Technology Corp. + "^rambus,.*": + description: Rambus Inc. "^ramtron,.*": description: Ramtron International "^raspberrypi,.*": --=20 2.43.7 From nobody Fri Sep 25 01:20:33 2026 Received: from CO1PR03CU002.outbound.protection.outlook.com (mail-westus2azon11020125.outbound.protection.outlook.com [52.101.46.125]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BE93F4B0E42; Thu, 17 Sep 2026 22:59:42 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.46.125 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789686012; cv=fail; b=FS4mezgftJWN57wAdAsx/jh6MNq33xJeFNWIHnl8flVXHQQ3oIsymtb3kaTQsDP834Pu5YVizhVodsXha8o3TmHf/ECpx2srAN6n9kO/vcP1H+yftDjSJ5NJitrQahgRTPFmZItiYrJ7bRnmEs1tflL/CPX7QJyAhfsKDMSZoGY= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789686012; c=relaxed/simple; bh=moc3vZWKNXoveDtHDrSsD6kTbhGbIY0te+MHmI6OTOI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=U3U8zSOFOChRTn1dxj28RypN6Fsy6j0acfjozsUcZMWuc7bISIEbDBa88AH/VSj6wWubiFEzXvNexRTgH4F6xdLPmoRBNo5Z2M7OCpsbmktxT7+ogjsJYLyzYIeewZQyKkdw+IHNWcYStJtLsfIdrUdRi1yt0Domcg7fUNYXu7Y= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com; spf=fail smtp.mailfrom=rambus.com; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b=YZ7zBd6Q; arc=fail smtp.client-ip=52.101.46.125 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=rambus.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b="YZ7zBd6Q" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=OCApODLmg16ICHFcDPBYlG6lk6VardmV/1zrUfzukOKHk79NLzYaldjgWtyjJmFhw7tuaaJjmVwSaDi7KDkVITaMD8MJBVHYbCoRroB1Wudz7jee0w4Ge1lHPr3mWfj8bTb4KBYA8CyTc6mDeMX8UCB4pIo8t0nYF47d83YqovQtvlXeQWpeMUaXOfayG8DRI/q3zOntS78f+AQnfHP+8CMhSSPdXDwEJm21X55SFYpXGWOhy/kmZt1cH5Hh71weIHjaEAzdWzbAsG2bbzFrbO0/luGZgS6ANju5QtLA0TzzxftQyezRH1xuEy2V+ysgvK2VL1fRusS01VXWYcsESw== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=OXGJn3eORrTkbLfMREWUacJT8vsYt87VmtfyAyz9bsk=; b=Ag/OZE7u74SBXRsyMpE4DRgIXMFYQ5eek53Bnp1cscydC8Pwkg5wnec27S4Pgh8eM8Qt4nJJHnL8eWU4RSm2v4BzKyWm2NimAFpOJhKnYi516zpkQDJjpjhDRLGPeqMJl9FTBN5Vw9b31vd5BroZ/pI0zy/8iAKNq8rnf97pkJYBJNpxNrnRMa827GA04M75r64WOGDkieSOneINkFsVlubZMR4cw7onnJiBkVkmKnvcHs8tN3RCecFn3X0ht6zSnrdGFrSSiLmgnWoG+/LDRoIDzhfyDroQN1YlkEHRHGTk8WBIJ3NZfzlXm5AH0kP21y26VUlWLjo7h9CdjqTa/Q== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 192.86.86.210) smtp.rcpttodomain=cryptography.com smtp.mailfrom=rambus.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=rambus.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=rambus.com; s=selector1; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=OXGJn3eORrTkbLfMREWUacJT8vsYt87VmtfyAyz9bsk=; b=YZ7zBd6QSwHTc6BJ2yPqpJ59u1NWx/CQcSI57k0TUyPQKamkMif/YsSgN9h40j6bXzBkN2/l3JpgKXZz6mM16nU4wkgXSHPKnMFLrsPznCsiroarU4s1zjQEfqcRp+sNAYBG9FfLHsPvfWPerXn2rBv0lTqxDzPBzydst5SYfjxaL0s2wIWsxEBEgtPbWVEVMEOt7dAeARG02oGT3Af0SQA6h+pf34l/FtflXFwR6C1jX58fA3L1vUmRFHs4GX06GfKfZMtJajp2NbRo2MREfSJddqGdKQi7mh4J9oMV65C6lHp2D5NA7IO3NhlveD0CajN/5pF/LCIBvQXVYDK+Tw== Received: from CH5PR04CA0008.namprd04.prod.outlook.com (2603:10b6:610:1f4::25) by SJ0PR04MB7887.namprd04.prod.outlook.com (2603:10b6:a03:304::13) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.428.13; Thu, 17 Sep 2026 22:59:33 +0000 Received: from BN2PEPF0000A802.namprd02.prod.outlook.com (2603:10b6:610:1f4:cafe::5e) by CH5PR04CA0008.outlook.office365.com (2603:10b6:610:1f4::25) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.428.12 via Frontend Transport; Thu, 17 Sep 2026 22:59:32 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 192.86.86.210) smtp.mailfrom=rambus.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=rambus.com; Received-SPF: Pass (protection.outlook.com: domain of rambus.com designates 192.86.86.210 as permitted sender) receiver=protection.outlook.com; client-ip=192.86.86.210; helo=hqxsv-psmtppxy02.rambus.com; pr=C Received: from hqxsv-psmtppxy02.rambus.com (192.86.86.210) by BN2PEPF0000A802.mail.protection.outlook.com (10.167.245.171) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.451.8 via Frontend Transport; Thu, 17 Sep 2026 22:59:32 +0000 Received: from hqxsv-cmdev3-aousherovitch.rambus.com (hqn-lb-int-float.rambus.com [10.12.20.20]) by hqxsv-psmtppxy02.rambus.com (Postfix) with ESMTP id B897A180175C; Thu, 17 Sep 2026 22:59:31 +0000 (UTC) From: Alex Ousherovitch To: Albert Ou , Alex Ousherovitch , Conor Dooley , "David S. Miller" , Herbert Xu , Jonathan Corbet , Krzysztof Kozlowski , Palmer Dabbelt , Paul Walmsley , Rob Herring , Saravanakrishnan Krishnamoorthy , Shuah Khan Cc: Alexandre Ghiti , devicetree@vger.kernel.org, Joel Wittenauer , linux-api@vger.kernel.org, linux-crypto@vger.kernel.org, linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-riscv@lists.infradead.org, Shuah Khan , Thi Nguyen Subject: [PATCH v5 02/19] crypto: cmh - add core platform driver Date: Thu, 17 Sep 2026 15:59:11 -0700 Message-ID: <20260917225929.2494111-3-aousherovitch@rambus.com> X-Mailer: git-send-email 2.43.7 In-Reply-To: <20260917225929.2494111-1-aousherovitch@rambus.com> References: <20260917225929.2494111-1-aousherovitch@rambus.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: BN2PEPF0000A802:EE_|SJ0PR04MB7887:EE_ X-MS-Office365-Filtering-Correlation-Id: fe30ac26-71c6-4c70-ab6c-08df150f564d X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|1800799024|23010399003|7416014|376014|82310400026|36860700016|921020|7136999003|18002099003|22082099003|56012099006|3023799007|11063799006|5023799004|10067099003|6133799003; X-Microsoft-Antispam-Message-Info: BuQ+AlbRyUa480q4BynQuPXatjZXJE4peUnNzhhhmV/cuGSSNFoNRuvFHqJYn1gMTGkhLJJMM7aaif80x9SK3VlG9nc2MZTv8CByRy+ZwiC4iCFMc7JxXx33YRqcyVtjHE/lvKtufXdxfDS+gAWZcRsTE1FoxnBZE7YuLduK2uPW8gZ8lQNVw/nycBoQ+yaEP3eiaMyWXqQhKErOVIEUWVuvqY4ddRq1hde1JbragCGhxS7O5wt/sKCYy4CuvmFhC+B/r8da94IS2NNKah6zNG7ZW5W73F7pzFDmdkk8oERzXgLJUM8GgGtZ8hyzTWeb/7bkhr6doWECJXzPhq875ze1pqNLdPgG8DOxlYRlvmZK0fpuA2FYxl7t364aZ0eIEfUZ9tyhyiTA/wsnPkPts3NMKvJPpJ7Vy8SnX9RhLIoWqeiQ9WdaGF0pjUqcPmKS0PciUtu4M8LohlN3/D8E7C+qT6TpSr4SSRJMcYNfhwfcv7y60HIKf3y3WTpKw5mdwVMyojm8jcUUcTgW79bh+KxLQDQ+Q00P6x4l951gcGabZQeiJ1NhZXfxgygacv5spp09iQWNPlNefugum6X55HSMLo1NM3kcqj5yrCRs9P5mawd3yeO7UJVNtHCg8wFzTh1QpiWnHXwqxNIqXSUJPAlp1ERk0vR+id/SYpcl4pdbVB1Bk2aTYxN739UsTt9n/jvd5uZmTzqpHjINjzr+aMK3MaEL7JsE2BxHeCKzamfdfIvTc1hoAefQov2Nqz1d X-Forefront-Antispam-Report: CIP:192.86.86.210;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:hqxsv-psmtppxy02.rambus.com;PTR:InfoDomainNonexistent;CAT:NONE;SFS:(13230040)(1800799024)(23010399003)(7416014)(376014)(82310400026)(36860700016)(921020)(7136999003)(18002099003)(22082099003)(56012099006)(3023799007)(11063799006)(5023799004)(10067099003)(6133799003);DIR:OUT;SFP:1102; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: eNKV/yHSdJBIQ9jY91QOLzz2OTOxvZkhDMUhRxuW/0YvPI5KJrpC6SRimb1hAIXt38beCl2gXD4d1n0lGPuANQuguesMWVtN/tj73PNCFHJP2vNSdoFqGrCiM+2uYZ+1pLn0+dAwaTLIqDnWM65laUE56rKyIJFEWFbMFYFl2kYpa7+EKyxRd0bxL8eWYl45bv5ZJWz5vx71lHZMydGmInKA3AUw5Q3LWLmbQdIwDP27oaiAbmTxoNarFPSnSFysip2m80gd9IT0FWwe04eEHlO0sQyXGeZndlwnIyBnUHnuPDr9v0f2/nfxabATMrcwl7mcy/iDGmxma2SKNSIN90MqW9Ze6OuZwaZjwn3r5TN+sJFKiEMs8UHzCZMoqGgLAoCFI+k8VQWTSaH1S3Oo0sx2d0OKR4Q03weIx5VlEnWxET1ICGpnqtOV00qQz+CI X-OriginatorOrg: rambus.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 17 Sep 2026 22:59:32.4160 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: fe30ac26-71c6-4c70-ab6c-08df150f564d X-MS-Exchange-CrossTenant-Id: bd0ba799-c2b9-413c-9c56-5d1731c4827c X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=bd0ba799-c2b9-413c-9c56-5d1731c4827c;Ip=[192.86.86.210];Helo=[hqxsv-psmtppxy02.rambus.com] X-MS-Exchange-CrossTenant-AuthSource: BN2PEPF0000A802.namprd02.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: SJ0PR04MB7887 Content-Type: text/plain; charset="utf-8" Add the Rambus CryptoManager Hub (CMH) hardware crypto accelerator core platform driver. This patch provides: - Platform driver registration and probe/remove lifecycle - Hardware configuration and core discovery - Mailbox Queue Interface (MQI) for VCQ command submission - Transaction manager with async completion and backlog support - Result handler with threaded-IRQ completion - DMA buffer management - Debugfs instrumentation (when CONFIG_CRYPTO_DEV_CMH_DEBUG=3Dy) - Sysfs attributes (fw_version, hw_version, boot_status, mbx_available, mbx_count) - Kconfig and Makefile integration No crypto algorithms are registered yet -- those follow in subsequent patches. The driver communicates with the hardware via a mailbox-based VCQ (Virtual Command Queue) interface. Each crypto operation is packed into VCQ command entries, submitted to a mailbox, and completed asynchronously via interrupt. MODULE_IMPORT_NS(CRYPTO_INTERNAL) imports two symbols: - crypto_cipher_setkey() [crypto/cipher.c, EXPORT_SYMBOL_NS_GPL] - crypto_cipher_encrypt_one() [crypto/cipher.c, EXPORT_SYMBOL_NS_GPL] These are the single-block cipher API used for software-fallback paths: CCM empty-input tag computation (2 ECB encryptions + XOR) and XCBC(SM4) empty-message workaround (3 ECB encryptions + XOR). No public wrapper exists; this is the same pattern used by in-tree crypto/ccm.c, crypto/cmac.c, and crypto/xcbc.c. Signed-off-by: Alex Ousherovitch Co-developed-by: Saravanakrishnan Krishnamoorthy Signed-off-by: Saravanakrishnan Krishnamoorthy --- Documentation/ABI/testing/debugfs-driver-cmh | 171 ++ Documentation/ABI/testing/sysfs-driver-cmh | 67 + Documentation/crypto/device_drivers/cmh.rst | 357 +++ Documentation/crypto/device_drivers/index.rst | 1 + drivers/crypto/Kconfig | 1 + drivers/crypto/Makefile | 1 + drivers/crypto/cmh/Kconfig | 65 + drivers/crypto/cmh/Makefile | 25 + drivers/crypto/cmh/cmh_config.c | 619 +++++ drivers/crypto/cmh/cmh_debugfs.c | 289 +++ drivers/crypto/cmh/cmh_dma.c | 379 +++ drivers/crypto/cmh/cmh_main.c | 343 +++ drivers/crypto/cmh/cmh_mqi.c | 344 +++ drivers/crypto/cmh/cmh_rh.c | 1178 +++++++++ drivers/crypto/cmh/cmh_sysfs.c | 108 + drivers/crypto/cmh/cmh_txn.c | 2173 +++++++++++++++++ drivers/crypto/cmh/include/cmh.h | 24 + drivers/crypto/cmh/include/cmh_aes_abi.h | 98 + drivers/crypto/cmh/include/cmh_ccp_abi.h | 108 + drivers/crypto/cmh/include/cmh_config.h | 105 + drivers/crypto/cmh/include/cmh_debugfs.h | 90 + drivers/crypto/cmh/include/cmh_dma.h | 219 ++ drivers/crypto/cmh/include/cmh_drbg_abi.h | 67 + drivers/crypto/cmh/include/cmh_eac_abi.h | 44 + drivers/crypto/cmh/include/cmh_hc_abi.h | 162 ++ drivers/crypto/cmh/include/cmh_hcq_abi.h | 221 ++ drivers/crypto/cmh/include/cmh_kic_abi.h | 77 + drivers/crypto/cmh/include/cmh_mqi.h | 35 + drivers/crypto/cmh/include/cmh_pke_abi.h | 272 +++ drivers/crypto/cmh/include/cmh_qse_abi.h | 181 ++ drivers/crypto/cmh/include/cmh_registers.h | 161 ++ drivers/crypto/cmh/include/cmh_rh.h | 93 + drivers/crypto/cmh/include/cmh_rng.h | 31 + drivers/crypto/cmh/include/cmh_sm3_abi.h | 79 + drivers/crypto/cmh/include/cmh_sm4_abi.h | 101 + drivers/crypto/cmh/include/cmh_sys_abi.h | 148 ++ drivers/crypto/cmh/include/cmh_sysfs.h | 14 + drivers/crypto/cmh/include/cmh_txn.h | 491 ++++ drivers/crypto/cmh/include/cmh_vcq.h | 288 +++ 39 files changed, 9230 insertions(+) create mode 100644 Documentation/ABI/testing/debugfs-driver-cmh create mode 100644 Documentation/ABI/testing/sysfs-driver-cmh create mode 100644 Documentation/crypto/device_drivers/cmh.rst create mode 100644 drivers/crypto/cmh/Kconfig create mode 100644 drivers/crypto/cmh/Makefile create mode 100644 drivers/crypto/cmh/cmh_config.c create mode 100644 drivers/crypto/cmh/cmh_debugfs.c create mode 100644 drivers/crypto/cmh/cmh_dma.c create mode 100644 drivers/crypto/cmh/cmh_main.c create mode 100644 drivers/crypto/cmh/cmh_mqi.c create mode 100644 drivers/crypto/cmh/cmh_rh.c create mode 100644 drivers/crypto/cmh/cmh_sysfs.c create mode 100644 drivers/crypto/cmh/cmh_txn.c create mode 100644 drivers/crypto/cmh/include/cmh.h create mode 100644 drivers/crypto/cmh/include/cmh_aes_abi.h create mode 100644 drivers/crypto/cmh/include/cmh_ccp_abi.h create mode 100644 drivers/crypto/cmh/include/cmh_config.h create mode 100644 drivers/crypto/cmh/include/cmh_debugfs.h create mode 100644 drivers/crypto/cmh/include/cmh_dma.h create mode 100644 drivers/crypto/cmh/include/cmh_drbg_abi.h create mode 100644 drivers/crypto/cmh/include/cmh_eac_abi.h create mode 100644 drivers/crypto/cmh/include/cmh_hc_abi.h create mode 100644 drivers/crypto/cmh/include/cmh_hcq_abi.h create mode 100644 drivers/crypto/cmh/include/cmh_kic_abi.h create mode 100644 drivers/crypto/cmh/include/cmh_mqi.h create mode 100644 drivers/crypto/cmh/include/cmh_pke_abi.h create mode 100644 drivers/crypto/cmh/include/cmh_qse_abi.h create mode 100644 drivers/crypto/cmh/include/cmh_registers.h create mode 100644 drivers/crypto/cmh/include/cmh_rh.h create mode 100644 drivers/crypto/cmh/include/cmh_rng.h create mode 100644 drivers/crypto/cmh/include/cmh_sm3_abi.h create mode 100644 drivers/crypto/cmh/include/cmh_sm4_abi.h create mode 100644 drivers/crypto/cmh/include/cmh_sys_abi.h create mode 100644 drivers/crypto/cmh/include/cmh_sysfs.h create mode 100644 drivers/crypto/cmh/include/cmh_txn.h create mode 100644 drivers/crypto/cmh/include/cmh_vcq.h diff --git a/Documentation/ABI/testing/debugfs-driver-cmh b/Documentation/A= BI/testing/debugfs-driver-cmh new file mode 100644 index 000000000000..21c2d0fbb171 --- /dev/null +++ b/Documentation/ABI/testing/debugfs-driver-cmh @@ -0,0 +1,171 @@ +What: /sys/kernel/debug/cmh/mbx/vcqs_submitted +Date: June 2026 +KernelVersion: 7.1 +Contact: linux-crypto@vger.kernel.org +Description: + (RO) Total number of VCQ command entries submitted to + mailbox N since the driver was loaded. + +What: /sys/kernel/debug/cmh/mbx/vcqs_completed +Date: June 2026 +KernelVersion: 7.1 +Contact: linux-crypto@vger.kernel.org +Description: + (RO) Total number of VCQ command completions received + from mailbox N. + +What: /sys/kernel/debug/cmh/mbx/vcqs_errors +Date: June 2026 +KernelVersion: 7.1 +Contact: linux-crypto@vger.kernel.org +Description: + (RO) Total number of error completions received from + mailbox N. + +What: /sys/kernel/debug/cmh/mbx/queue_full_count +Date: June 2026 +KernelVersion: 7.1 +Contact: linux-crypto@vger.kernel.org +Description: + (RO) Number of times the transaction manager skipped + mailbox N because its in-flight queue was full. + +What: /sys/kernel/debug/cmh/mbx/max_queue_depth +Date: June 2026 +KernelVersion: 7.1 +Contact: linux-crypto@vger.kernel.org +Description: + (RO) High-water mark of in-flight transactions on + mailbox N. + +What: /sys/kernel/debug/cmh/mbx/inject_abort +Date: June 2026 +KernelVersion: 7.1 +Contact: linux-crypto@vger.kernel.org +Description: + (WO) Write any value to inject an MBX_COMMAND_ABORT on + mailbox N. The abort triggers error-IRQ handling that + completes all in-flight transactions with -EIO and then + issues MBX_COMMAND_RESTART to resume the mailbox. + Only available when CONFIG_CRYPTO_DEV_CMH_DEBUG is enabled. + +What: /sys/kernel/debug/cmh/mbx/force_drain +Date: June 2026 +KernelVersion: 7.1 +Contact: linux-crypto@vger.kernel.org +Description: + (WO) Write any value to unconditionally FLUSH and drain + all pending transactions on mailbox N, completing each + with -ECANCELED, and reset all recovery bookkeeping + (including the wedged flag). The mailbox is re-enabled + for new work immediately; no hardware health verification + is performed. Use as a last-resort recovery when the eSW + is unresponsive and normal ABORT/RESTART escalation has + not recovered the mailbox. + Only available when CONFIG_CRYPTO_DEV_CMH_DEBUG is enabled. + +What: /sys/kernel/debug/cmh/tm/cmq_posts +Date: June 2026 +KernelVersion: 7.1 +Contact: linux-crypto@vger.kernel.org +Description: + (RO) Total number of cmh_tm_post_command() calls (one + per crypto request submitted to the transaction manager). + +What: /sys/kernel/debug/cmh/tm/cmq_depth_max +Date: June 2026 +KernelVersion: 7.1 +Contact: linux-crypto@vger.kernel.org +Description: + (RO) High-water mark of the command queue length. + +What: /sys/kernel/debug/cmh/tm/cmq_eagain_count +Date: June 2026 +KernelVersion: 7.1 +Contact: linux-crypto@vger.kernel.org +Description: + (RO) Number of times the command queue was full and + returned -EAGAIN to the caller. + +What: /sys/kernel/debug/cmh/tm/backoff_count +Date: June 2026 +KernelVersion: 7.1 +Contact: linux-crypto@vger.kernel.org +Description: + (RO) Number of times the transaction manager backed off + because all mailbox queues were full. + +What: /sys/kernel/debug/cmh/tm/async_timeout_count +Date: June 2026 +KernelVersion: 7.1 +Contact: linux-crypto@vger.kernel.org +Description: + (RO) Number of async crypto requests that timed out + waiting for hardware completion. + +What: /sys/kernel/debug/cmh/config/async_timeout_ms +Date: June 2026 +KernelVersion: 7.1 +Contact: linux-crypto@vger.kernel.org +Description: + (RW) Async request timeout in milliseconds. On timeout + the driver issues MBX_COMMAND_ABORT; if the eSW is + unresponsive, the watchdog escalates through RESTART, + FLUSH, and force-drain to bound D-state duration. + +What: /sys/kernel/debug/cmh/config/vcq_timeout_ms +Date: June 2026 +KernelVersion: 7.1 +Contact: linux-crypto@vger.kernel.org +Description: + (RW) VCQ command timeout in milliseconds. + +What: /sys/kernel/debug/cmh/config/slow_op_timeout_ms +Date: June 2026 +KernelVersion: 7.1 +Contact: linux-crypto@vger.kernel.org +Description: + (RW) Slow-operation timeout in milliseconds. Used for + operations known to take longer (e.g. RSA key generation, + PQC key generation). + +What: /sys/kernel/debug/cmh/config/drain_timeout_ms +Date: June 2026 +KernelVersion: 7.1 +Contact: linux-crypto@vger.kernel.org +Description: + (RW) Drain timeout in milliseconds. Maximum time to wait + for all in-flight transactions to complete during driver + removal or suspend. + +What: /sys/kernel/debug/cmh/config/watchdog_ms +Date: June 2026 +KernelVersion: 7.1 +Contact: linux-crypto@vger.kernel.org +Description: + (RW) Result-handler watchdog interval in milliseconds. + Detects missed IRQs, stuck mailboxes, and abort-stall + conditions. Clamped to a 10 ms minimum. + +What: /sys/kernel/debug/cmh/config/drbg_timeout_ms +Date: June 2026 +KernelVersion: 7.1 +Contact: linux-crypto@vger.kernel.org +Description: + (RW) DRBG self-seed timeout in milliseconds. + +What: /sys/kernel/debug/cmh/config/cmq_max_depth +Date: August 2026 +KernelVersion: 7.2 +Contact: linux-crypto@vger.kernel.org +Description: + (RW) Maximum number of pending commands in the central + Command Message Queue. + +What: /sys/kernel/debug/cmh/config/backlog_max_depth +Date: August 2026 +KernelVersion: 7.2 +Contact: linux-crypto@vger.kernel.org +Description: + (RW) Maximum depth of the backlog queue for + CRYPTO_TFM_REQ_MAY_BACKLOG requests (0 disables backlog). diff --git a/Documentation/ABI/testing/sysfs-driver-cmh b/Documentation/ABI= /testing/sysfs-driver-cmh new file mode 100644 index 000000000000..07216c4aed63 --- /dev/null +++ b/Documentation/ABI/testing/sysfs-driver-cmh @@ -0,0 +1,67 @@ +What: /sys/devices/platform//fw_version +Date: June 2026 +KernelVersion: 7.1 +Contact: linux-crypto@vger.kernel.org +Description: + Reports the CryptoManager Hub embedded software (eSW) firmware + version as a 32-bit hexadecimal value read from the SIC + SW_VERSION register. + + Example: "0x00010002" + + Read-only. + +What: /sys/devices/platform//hw_version +Date: June 2026 +KernelVersion: 7.1 +Contact: linux-crypto@vger.kernel.org +Description: + Reports the CryptoManager Hub hardware version as a 32-bit + hexadecimal value read from the SIC HW_VERSION0 register. + + Example: "0x00000000" + + Read-only. + +What: /sys/devices/platform//boot_status +Date: June 2026 +KernelVersion: 7.1 +Contact: linux-crypto@vger.kernel.org +Description: + Reports the CryptoManager Hub boot status register as a 32-bit + hexadecimal value. This reflects the firmware boot + progress and final state: + + 0x00000066 - firmware booted (post-self-test) + other - firmware boot in progress or failed + + Read-only. + +What: /sys/devices/platform//mbx_available +Date: June 2026 +KernelVersion: 7.1 +Contact: linux-crypto@vger.kernel.org +Description: + Reports the mailbox availability bitmap as a 32-bit + hexadecimal value read from the SIC MBX_AVAILABILITY + register. Each set bit indicates a hardware mailbox + instance that the firmware has made available. + + Example: "0x00000003" (mailboxes 0 and 1 available) + + Read-only. + +What: /sys/devices/platform//mbx_count +Date: June 2026 +KernelVersion: 7.1 +Contact: linux-crypto@vger.kernel.org +Description: + Reports the number of mailboxes the driver has configured, + as a decimal integer. This reflects the driver's active + configuration (from DT properties, or a debug-only + override), which may be fewer than illustrated by + mbx_available. + + Example: "2" + + Read-only. diff --git a/Documentation/crypto/device_drivers/cmh.rst b/Documentation/cr= ypto/device_drivers/cmh.rst new file mode 100644 index 000000000000..9accd42e27ab --- /dev/null +++ b/Documentation/crypto/device_drivers/cmh.rst @@ -0,0 +1,357 @@ +.. SPDX-License-Identifier: GPL-2.0 + +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D +Rambus CryptoManager Hub (CMH) Driver +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +Overview +=3D=3D=3D=3D=3D=3D=3D=3D + +The ``cmh`` driver supports the Rambus CryptoManager Hub hardware cryptogr= aphic +accelerator. The hardware is accessed through a mailbox-based VCQ +(Virtual Command Queue) interface: the driver writes command sequences +into per-mailbox DMA queue buffers and rings a doorbell register; the +CryptoManager Hub embedded software (eSW) processes the commands and signa= ls +completion via a per-mailbox interrupt. + +The driver registers algorithms with the Linux kernel crypto subsystem +and exposes a management character device (``/dev/cmh_mgmt``) for +operations that have no standard crypto API binding. + +Hardware Interface +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +The CryptoManager Hub is presented as a platform device matched via Device= Tree +(compatible ``"rambus,cmh-v1030"``). The driver maps a single MMIO region +(the SIC -- System Interface Controller) whose sub-regions contain +per-mailbox doorbell, status, and command queue registers. + +The driver manages a configurable number of mailboxes (default 2). +Each mailbox has a configurable number of slots (default 64) and a +configurable stride (default 512 bytes per slot). The driver allocates +DMA-coherent memory for each mailbox queue during probe. + +A mailbox is owned by a single host identity for the duration of a lock, +and the hardware permits only that identity to access it. The platform +must present one consistent HOST ID for all accesses to a given mailbox, +independent of the issuing CPU; see the device-tree binding for this +integration requirement. + +Interrupts are per-mailbox completion/error interrupts. The driver +registers a threaded IRQ handler for each configured mailbox. + +The eSW is loaded independently of this driver -- typically by the +boot firmware or a platform-specific loader -- so the driver does not +use ``request_firmware()``. Instead it waits for the eSW to reach +mission mode during probe, bounded by a fixed timeout. + +Supported Algorithms +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +The driver registers the following algorithm families: + +Hash (ahash) + SHA-224, SHA-256, SHA-384, SHA-512, SHA3-224, SHA3-256, SHA3-384, + SHA3-512, SHAKE-128, SHAKE-256, cSHAKE-128, cSHAKE-256, KMAC-128, + KMAC-256, SM3 (10 hash + 2 cSHAKE + 2 KMAC + 1 SM3 =3D 15 algorithms) + +HMAC (ahash) + HMAC-SHA-224, HMAC-SHA-256, HMAC-SHA-384, HMAC-SHA-512, + HMAC-SHA3-224, HMAC-SHA3-256, HMAC-SHA3-384, HMAC-SHA3-512 + (8 algorithms) + +Symmetric Ciphers (skcipher) + AES: ECB, CBC, CTR, CFB, XTS (5 algorithms) + SM4: ECB, CBC, CTR, CFB, XTS (5 algorithms) + ChaCha20 (1 algorithm) + +AEAD + AES-GCM, AES-CCM (2 algorithms) + SM4-GCM, SM4-CCM (2 algorithms) + ``rfc7539(chacha20,poly1305)``, ``rfc7539esp(chacha20,poly1305)`` + (2 algorithms) + +MAC (ahash) + CMAC(AES) (1 algorithm) + CMAC(SM4), XCBC(SM4) (2 algorithms) + Poly1305 (1 algorithm) + +Public-Key, Key Agreement, and PQC Signatures + RSA (akcipher, 1 algorithm) + ECDSA P-256, P-384, P-521 (sig, 3 algorithms) + SM2 (sig, verify-only, 1 algorithm) + ECDH P-256, P-384, X25519 (kpp, 3 algorithms) + ML-DSA-44, ML-DSA-65, ML-DSA-87 (sig, 3 algorithms) + SLH-DSA: all 12 parameter sets (sig, 12 algorithms) + LMS, LMS-HSS (sig, verify-only, 2 algorithms) + XMSS, XMSS-MT (sig, verify-only, 2 algorithms) + (ML-KEM keygen/encaps/decaps is available via ``/dev/cmh_mgmt`` + only -- see `Limitations`_.) + +Hardware RNG + DRBG-backed hwrng (``/dev/hwrng``, 1 algorithm) + +All algorithm driver names use the ``rambus-cmh-`` prefix (e.g. +``rambus-cmh-sha256``, ``rambus-cmh-ecb-aes``, ``rambus-cmh-gcm-aes``, +``rambus-cmh-mldsa44``). Names generally follow the kernel's hyphenated +template name; families that have no kernel template (e.g. ML-DSA) use +the concatenated upstream algorithm name (``mldsa44``). + +Most algorithms register at priority 300 (301 for AES-CCM). +The ML-DSA ``sig`` algorithms register at priority 5001 to +outrank the kernel's generic software ML-DSA (priority 5000, which is +verify-only); the CMH driver provides full hardware sign and verify. + +Request model +------------- + +All crypto API operations are asynchronous: the driver queues each +request to its transaction-manager kthread and returns +``-EINPROGRESS``, invoking the caller's completion callback when the +hardware finishes. Requests that set ``CRYPTO_TFM_REQ_MAY_BACKLOG`` +are queued on a bounded backlog when the command queue is full; +without that flag a full queue is reported as +``-EBUSY``. Hardware or eSW failures surface as ``-EIO``, malformed +requests as ``-EINVAL``, oversized requests as ``-EMSGSIZE`` or +``-EINVAL`` (see `Data-Size Limits`_), and unresponsive hardware as +``-ETIMEDOUT``. The ``/dev/cmh_mgmt`` ioctls are, by contrast, +synchronous -- each ioctl blocks until the hardware completes. + +Driver Architecture +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +The driver is structured as follows: + +Platform Driver + Matches DT compatible ``"rambus,cmh-v1030"``. Probe initializes all + subsystems in order; remove tears them down in reverse. + +Configuration + Parses and validates DT properties (mailbox counts, slot sizes, and + stride values). + +MQI (Mailbox Queue Interface) + Allocates DMA-coherent queue memory per mailbox. Manages slot + allocation, VCQ command writing, and doorbell ringing. + +Transaction Manager + A dedicated kthread dequeues crypto requests from a central command + queue, builds VCQ command sequences, and submits them to mailbox + slots. Completion is signaled via wait queues. + +Response Handler + Per-mailbox threaded IRQ handlers walk completed slots, parse + results, and fire request completions. A configurable watchdog + timer (the ``watchdog_ms`` debugfs knob, default 200 ms) detects + stuck requests and escalates through ABORT, RESTART, and FLUSH + recovery. + +Key Management (``/dev/cmh_mgmt``) + A misc character device providing ioctl-based access to datastore + key CRUD, key derivation (KIC), PKE operations (EdDSA, SM2), + PQC operations (ML-KEM, ML-DSA, SLH-DSA), + EAC error register readback, and DRBG runtime configuration. + See ``Documentation/ABI/testing/cmh-mgmt`` for the full ioctl list. + +Power Management + The driver implements ``DEFINE_SIMPLE_DEV_PM_OPS`` suspend/resume. + On suspend, the transaction-manager kthread is stopped and pending + transactions are drained, waiting up to ``drain_timeout_ms`` + (default 10000 ms); resume restarts the kthread. + +Module Parameters +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +The driver defines no production module parameters. All mailbox +topology, per-core affinity, slot counts, and strides are taken from +Device Tree properties; the eSW boot timeout and queue depths are +built-in constants, and the runtime-tunable knobs are exposed via +debugfs (see `debugfs Counters`_ below). + +A small set of debug-only parameters is compiled in only with +``CONFIG_CRYPTO_DEV_CMH_DEBUG``. They exist solely to force alternate +geometries at ``insmod`` time during bringup and validation (for +example, to drive the mailbox-contention and cross-mailbox dispatch +paths without rebuilding the Device Tree); they default to "use the DT +value" and have no effect in a production build: + +``mbx_count_override`` (uint, default 0) + Override the DT mailbox count (0 =3D use DT) to force fewer + mailboxes than the hardware provides. + +``mbx_slots_override`` (uint, default 0) + Override all per-mailbox slot counts (0 =3D use DT). + +``mbx_round_robin`` (bool, default false) + Ignore DT ``rambus,cores`` affinity and round-robin all cores + across the configured mailboxes (0 =3D use DT affinity). + +``skip_fw_check`` (bool, default false) + Skip the SIC boot status and eSW mission-mode checks at probe. + Allows the module to load before the eSW has booted. + +sysfs Attributes +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +The driver exposes five read-only attributes under the platform +device sysfs directory: ``fw_version``, ``hw_version``, +``boot_status``, ``mbx_available``, and ``mbx_count``. See +``Documentation/ABI/testing/sysfs-driver-cmh`` for the authoritative +per-attribute description. + +debugfs Counters +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +When built with ``CONFIG_CRYPTO_DEV_CMH_DEBUG``, the driver creates +``/sys/kernel/debug/cmh/`` with three groups: per-mailbox counters +(``mbxN/``), transaction-manager statistics (``tm/``), and +runtime-tunable knobs (``config/``, including ``drain_timeout_ms``, +``watchdog_ms``, ``cmq_max_depth``, and ``backlog_max_depth``). See +``Documentation/ABI/testing/debugfs-driver-cmh`` for the authoritative +per-file description. + +Device Tree Binding +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +See ``Documentation/devicetree/bindings/crypto/rambus,cmh-v1030.yaml`` for= the +full DT binding schema and complete, schema-validated examples. Each +mailbox owned by the host is a ``queue@N`` child node with its own +``reg`` (instance index), optional ``interrupts``, optional per-mailbox +VCQ ring geometry (``rambus,num-slots`` / ``rambus,slot-stride-bytes``, +defaulting to 64 and 512), and an optional ``rambus,cores`` list pinning +specific crypto cores to that mailbox. Which crypto cores are present is = read from the +SIC ``CORE_ENABLE`` register at probe, not described in the device tree. + +The parent node may also carry up to three ``clocks`` (the main core +clock, a half-rate clock present only on configurations with +side-channel-protected cores, and the real-time tick clock) and an +optional ``reset-gpios``. The driver enables every supplied clock and +acquires the reset line deasserted; it does not drive a reset sequence. +Both are absent when a separate management controller owns them, in +which case the driver drives neither. + +User-Space Interfaces +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +``/dev/cmh_mgmt`` + Management character device. Opening it requires ``CAP_SYS_ADMIN``. + See ``Documentation/ABI/testing/cmh-mgmt`` for ioctl documentation. + The UAPI header is ````. + +In-kernel crypto API + All algorithms register with the standard kernel crypto API and are + consumed by in-kernel users (dm-crypt, fscrypt, IPsec, kTLS, etc.). + + Keys provisioned inside the hardware via ``/dev/cmh_mgmt`` are + referenced by an opaque hardware key identifier and are operated on + through the ``/dev/cmh_mgmt`` ioctl interface, without ever exposing + plaintext key material to user space. See + ``Documentation/ABI/testing/cmh-mgmt`` for key provisioning. + +``/dev/hwrng`` + The DRBG-backed hardware RNG is available as a standard hwrng device. + The driver configures the DRBG at probe with built-in defaults where + the firmware permits it (stateless mode, or on the management host); + otherwise the management host configures it out of band or via the + ``/dev/cmh_mgmt`` ioctl, after which random data becomes available. + +Limitations +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +- LMS and XMSS support verify-only (no sign/keygen in hardware for + stateful hash-based signatures). +- SM2 sig registration is verify-only (sign via ``/dev/cmh_mgmt`` ioctl). +- EdDSA (Ed25519/Ed448) is available only through ``/dev/cmh_mgmt`` + ioctls; no kernel ``sig`` registration. +- ML-KEM operations (encapsulate/decapsulate/keygen) are available only + through ``/dev/cmh_mgmt`` ioctls; no standard kernel crypto API + binding exists for KEM. + +Data-Size Limits +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +The driver imposes data-size limits on several APIs. These are +driver-level safety caps for kernel memory allocation unless noted +otherwise. + +Symmetric / AEAD / MAC linearization caps: + +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D +Scope Limit Origin +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D +AES skcipher 32 MiB Driver-imposed DMA linearization = cap +SM4 skcipher 32 MiB Driver-imposed DMA linearization = cap +All AEAD + ChaCha20 skcipher 1 MiB Driver-imposed DMA linearization = cap +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D + +MAC and keyed-hash algorithms buffer all input in kernel memory because +the hardware exposes no keyed-MAC context save/restore. Rather than +enforce a hard limit, most of them fall back to a software implementation +once the buffered input -- or a clone request -- exceeds the window: + +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D =3D=3D=3D=3D= =3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D +Algorithm Window Behaviour past the window +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D =3D=3D=3D=3D= =3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D +``cmac(aes)`` 64 KiB switch to generic software cmac(aes) +``cmac(sm4)`` 64 KiB switch to generic software cmac(sm4) +``xcbc(sm4)`` 64 KiB switch to generic software xcbc(sm4) +``poly1305`` 64 KiB switch to the in-kernel Poly1305 library +``hmac(sha*)`` 64 KiB switch to generic software hmac(sha\*) +``hmac(sha3-*)`` 64 KiB switch to generic software hmac(sha3-\*) +``kmac128`` 64 KiB hard limit -- reject with -EINVAL +``kmac256`` 64 KiB hard limit -- reject with -EINVAL +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D =3D=3D=3D=3D= =3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +For hmac/cmac/xcbc the driver allocates and keys a matching generic +software MAC transform itself -- it registers with +``CRYPTO_ALG_NO_FALLBACK`` and does not rely on the crypto core's +automatic fallback; poly1305, which has no ``crypto_shash`` provider, +uses the in-kernel Poly1305 library directly. +After the switch, both arbitrarily long messages and transform ``clone`` +(``.export()``/``.import()``) work at any length. + +KMAC is the exception: there is neither a generic KMAC shash nor a KMAC +library, and the eSW rejects the save command while ``outlen !=3D 0`` +(always true for KMAC), so it can neither stream nor serialize its state. +It keeps a hard 64 KiB cap (``.update()`` returns ``-EINVAL`` past it) +and returns ``-EOPNOTSUPP`` from ``.export()``/``.import()``. For +HMAC-SHA3 the same software fallback also avoids exposing the invertible +Keccak sponge state, which would otherwise allow key recovery; the eSW +likewise does not expose HMAC-SHA2 save/restore. + +Pure hash algorithms (SHA-2, SHA-3, SHAKE, cSHAKE, SM3) have no data +limit because the hardware supports incremental save/restore. + +cSHAKE uses save/restore for ``.export()``/``.import()`` but accumulates +data in ``.update()`` by design (the Keccak sponge has no block-alignment +boundary to trigger per-update HW submission, and HC_CMD_GATHER amortizes +the cost into a single finalize-time submission). + +Asymmetric / PQC algorithm limits: + +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D +Scope Limit Origin +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D +RSA key size 4096 bit HW-imposed +ML-DSA message 10 KiB eSW-imposed (QSE ABI) +SLH-DSA message 128 B eSW-imposed (HCQ ABI) +SLH-DSA context 255 B Spec-imposed (FIPS 205) +LMS public key 60 B eSW-imposed (HCQ ABI) +LMS message 256 B eSW-imposed (HCQ ABI) +LMS signature 13,364 B eSW-imposed (HCQ ABI) +XMSS public key 136 B eSW-imposed (HCQ ABI) +XMSS message 64 B eSW-imposed (HCQ ABI) +XMSS signature 27,688 B eSW-imposed (HCQ ABI) +SM2 encrypt message 32 B eSW KDF (single SM3 block) +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D + +Miscellaneous limits: + +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D +Scope Limit Origin +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D +cSHAKE/KMAC customization 256 B VCQ slot layout constraint +KIC HKDF key 64 B Partially eSW-derived +KIC HKDF label 56 B VCQ slot layout constraint +Key/blob mgmt ioctls 256 KiB Driver-imposed sanity cap +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D diff --git a/Documentation/crypto/device_drivers/index.rst b/Documentation/= crypto/device_drivers/index.rst index c81d311ac61b..c0247fc97bf8 100644 --- a/Documentation/crypto/device_drivers/index.rst +++ b/Documentation/crypto/device_drivers/index.rst @@ -6,4 +6,5 @@ Hardware Device Driver Specific Documentation .. toctree:: :maxdepth: 1 =20 + cmh octeontx2 diff --git a/drivers/crypto/Kconfig b/drivers/crypto/Kconfig index 0189dfdcbbe1..9716b85b8e4b 100644 --- a/drivers/crypto/Kconfig +++ b/drivers/crypto/Kconfig @@ -831,4 +831,5 @@ source "drivers/crypto/starfive/Kconfig" source "drivers/crypto/inside-secure/eip93/Kconfig" source "drivers/crypto/ti/Kconfig" =20 +source "drivers/crypto/cmh/Kconfig" endif # CRYPTO_HW diff --git a/drivers/crypto/Makefile b/drivers/crypto/Makefile index 2c33b83f3cfa..dae0d5f0500d 100644 --- a/drivers/crypto/Makefile +++ b/drivers/crypto/Makefile @@ -46,3 +46,4 @@ obj-y +=3D intel/ obj-y +=3D starfive/ obj-y +=3D cavium/ obj-y +=3D ti/ +obj-$(CONFIG_CRYPTO_DEV_CMH) +=3D cmh/ diff --git a/drivers/crypto/cmh/Kconfig b/drivers/crypto/cmh/Kconfig new file mode 100644 index 000000000000..fca66d5e2f89 --- /dev/null +++ b/drivers/crypto/cmh/Kconfig @@ -0,0 +1,65 @@ +# SPDX-License-Identifier: GPL-2.0 +# +# Rambus CryptoManager Hub (CMH) hardware crypto accelerator +# +# The VCQ command and DMA descriptor ABI uses native (host-endian) +# integer types and is only valid on little-endian hosts: the CMH block +# and its eSW are co-located on the same little-endian SoC and there is +# no big-endian deployment. Gate the driver on !CPU_BIG_ENDIAN rather +# than byte-swapping an ABI that would never be exercised. +# + +config CRYPTO_DEV_CMH + tristate "Rambus CryptoManager Hub (CMH) hardware crypto accelerator" + depends on CRYPTO && OF && HAS_IOMEM && (64BIT || COMPILE_TEST) + depends on !CPU_BIG_ENDIAN + select CRYPTO_HASH + select CRYPTO_SKCIPHER + select CRYPTO_AEAD + select CRYPTO_AKCIPHER + select CRYPTO_SIG + select CRYPTO_KPP + select CRYPTO_ECC + select CRYPTO_RSA + select CRYPTO_AES + select CRYPTO_CCM + select CRYPTO_SM4_GENERIC + # Generic software MACs the keyed-MAC drivers allocate and key as + # arbitrary-length / transform-clone fallbacks past the 64 KB HW + # window: HMAC-SHA-2/SHA-3, CMAC (AES and SM4) and XCBC (SM4). + select CRYPTO_HMAC + select CRYPTO_SHA256 + select CRYPTO_SHA512 + select CRYPTO_SHA3 + select CRYPTO_CMAC + select CRYPTO_XCBC + # Poly1305 has no generic ahash/shash; its fallback drives the + # poly1305 library () directly. + select CRYPTO_LIB_POLY1305 + select HW_RANDOM + help + Driver for the Rambus CryptoManager Hub (CMH) hardware crypto accelerat= or. + Accesses the hardware via a mailbox-based VCQ (Virtual Command + Queue) interface and registers algorithms with the kernel + crypto subsystem. + + Supported algorithm families: AES (ECB/CBC/CTR/XTS/CFB), + SM4 (ECB/CBC/CTR/XTS/CFB), ChaCha20-Poly1305, AES-GCM, AES-CCM, + SHA-2, SHA-3, SHAKE, CSHAKE, KMAC, SM3, HMAC, AES-CMAC, + SM4-CMAC, SM4-XCBC, RSA, ECDSA, ECDH, SM2, and DRBG (hwrng). + Ioctl-only algorithms: EdDSA, ML-KEM. + + To compile this driver as a module, choose M here. + +config CRYPTO_DEV_CMH_DEBUG + bool "CMH debug instrumentation (debugfs counters)" + depends on CRYPTO_DEV_CMH && DEBUG_FS + help + Enable per-mailbox debugfs counters under + /sys/kernel/debug/cmh/ for the CMH driver. + Exposes VCQ submit/complete/error counts, queue depth + high-water marks, and transaction manager backoff statistics. + + Useful for bringup, validation, and performance analysis. + Not recommended for production. + diff --git a/drivers/crypto/cmh/Makefile b/drivers/crypto/cmh/Makefile new file mode 100644 index 000000000000..742e65e3917e --- /dev/null +++ b/drivers/crypto/cmh/Makefile @@ -0,0 +1,25 @@ +# SPDX-License-Identifier: GPL-2.0 +# +# Makefile for the Rambus CryptoManager Hub (CMH) hardware crypto accelera= tor driver. +# + +obj-$(CONFIG_CRYPTO_DEV_CMH) +=3D cmh.o + +cmh-y :=3D \ + cmh_main.o \ + cmh_config.o \ + cmh_mqi.o \ + cmh_txn.o \ + cmh_rh.o \ + cmh_dma.o \ + cmh_sysfs.o + +ccflags-y +=3D -I$(src)/include + +# Suppress -Woverride-init for the [0 ... N] =3D -1 range-initializer patt= ern +# (standard kernel idiom for sparse lookup tables with a default value). +CFLAGS_cmh_config.o +=3D -Wno-override-init + +# Debug instrumentation: per-mailbox debugfs counters. +# cmh_debugfs.o is linked into the composite cmh.o (same tristate). +cmh-$(CONFIG_CRYPTO_DEV_CMH_DEBUG) +=3D cmh_debugfs.o diff --git a/drivers/crypto/cmh/cmh_config.c b/drivers/crypto/cmh/cmh_confi= g.c new file mode 100644 index 000000000000..2ef588a3fcb1 --- /dev/null +++ b/drivers/crypto/cmh/cmh_config.c @@ -0,0 +1,619 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Configuration from Device Tree + * + * The CMH device tree node provides: + * - reg: SIC base + size (mandatory) + * + * Per-mailbox child nodes (queue@N): + * - reg: 0-based MBX instance index (mandatory) + * - interrupts: (optional) per-MBX completion IRQ; absent =3D> polling + * - rambus,num-slots / rambus,slot-stride-bytes: (optional) VCQ ring + * geometry (slot count / per-slot stride in bytes, both powers of two) + * - rambus,cores: (optional) crypto core names pinned to this mailbox + * + * Crypto cores are discovered from the SIC CORE_ENABLE register at probe + * (cmh_config_discover_cores), not described in the device tree. + */ + +#include +#include +#include +#include +#include +#include +#include + +#include "cmh_config.h" +#include "cmh_dma.h" + +/* + * Debug-only MBX overrides for stress testing. + * When non-zero, these override the corresponding DT values, enabling + * contention stress tests to force a minimal MBX config + * (e.g. mbx_count_override=3D1 mbx_slots_override=3D1 for 1 MBX, 2 slots). + */ +#ifdef CONFIG_CRYPTO_DEV_CMH_DEBUG +static unsigned int mbx_count_override; +module_param(mbx_count_override, uint, 0444); +MODULE_PARM_DESC(mbx_count_override, + "[debug] Override DT MBX count (0 =3D use DT, default: 0)"); + +static unsigned int mbx_slots_override; +module_param(mbx_slots_override, uint, 0444); +MODULE_PARM_DESC(mbx_slots_override, + "[debug] Override all MBX slots_log2 (0 =3D use DT, default: 0)"); + +static bool mbx_round_robin; +module_param(mbx_round_robin, bool, 0444); +MODULE_PARM_DESC(mbx_round_robin, + "[debug] Ignore DT rambus,cores affinity and round-robin all cores acro= ss MBXes (0 =3D use DT affinity, default: 0)"); +#endif + +/* -- Core ID -> core_type lookup --------------------------------------- = */ + +/* + * Map hardware core IDs (from DT child "reg") to enum cmh_core_type. + * + * Entries set to -1 are not dispatchable crypto cores: system cores + * (SYS, DMA, KIC, TIC, MPU, EMC, EAC) and the DRBG singleton + * (handled separately in cmh_rng.c). + */ +static const int core_id_to_type[CORE_ID_NUM] =3D { + [0 ... CORE_ID_NUM - 1] =3D -1, + [CORE_ID_HC] =3D CMH_CORE_HC, + [CORE_ID_AES] =3D CMH_CORE_AES, + [CORE_ID_SM4] =3D CMH_CORE_SM4, + [CORE_ID_SM3] =3D CMH_CORE_SM3, + [CORE_ID_CCP] =3D CMH_CORE_CCP, + [CORE_ID_PKE] =3D CMH_CORE_PKE, + [CORE_ID_QSE] =3D CMH_CORE_QSE, + [CORE_ID_HCQ] =3D CMH_CORE_HCQ, +}; + +/* Human-readable names for error messages */ +static const char * const core_type_names[CMH_NUM_CORE_TYPES] =3D { + [CMH_CORE_HC] =3D "hc", + [CMH_CORE_AES] =3D "aes", + [CMH_CORE_SM4] =3D "sm4", + [CMH_CORE_SM3] =3D "sm3", + [CMH_CORE_CCP] =3D "ccp", + [CMH_CORE_PKE] =3D "pke", + [CMH_CORE_QSE] =3D "qse", + [CMH_CORE_HCQ] =3D "hcq", +}; + +/* -- Hardware core discovery ------------------------------------------ */ + +/* + * Per-core-type discovery descriptor: the dual-rail CORE_ENABLE mask and + * the canonical hardware core ID for that type. DRBG is intentionally + * absent -- it is not a dispatchable crypto core (handled by cmh_rng.c). + */ +struct cmh_core_desc { + u32 enable_mask; + u32 core_id; +}; + +static const struct cmh_core_desc cmh_core_descs[CMH_NUM_CORE_TYPES] =3D { + [CMH_CORE_HC] =3D { SIC_CORE_ENABLE_HC, CORE_ID_HC }, + [CMH_CORE_AES] =3D { SIC_CORE_ENABLE_AES, CORE_ID_AES }, + [CMH_CORE_SM4] =3D { SIC_CORE_ENABLE_SM4, CORE_ID_SM4 }, + [CMH_CORE_SM3] =3D { SIC_CORE_ENABLE_SM3, CORE_ID_SM3 }, + [CMH_CORE_CCP] =3D { SIC_CORE_ENABLE_CCP, CORE_ID_CCP }, + [CMH_CORE_PKE] =3D { SIC_CORE_ENABLE_PKE, CORE_ID_PKE }, + [CMH_CORE_QSE] =3D { SIC_CORE_ENABLE_QSE, CORE_ID_QSE }, + [CMH_CORE_HCQ] =3D { SIC_CORE_ENABLE_HCQ, CORE_ID_HCQ }, +}; + +/* Dual-rail encoding: 0b01 (low bit set, high bit clear) means enabled. */ +static bool cmh_core_enabled(u32 core_enable, u32 mask) +{ + return (core_enable & (mask | (mask << 1))) =3D=3D mask; +} + +/* rambus,cores DT names -> hardware core IDs (see cmh_vcq.h). */ +static const struct cmh_dt_core_name { + const char *name; + u32 id; +} cmh_dt_core_names[] =3D { + { "hc", CORE_ID_HC }, + { "aes", CORE_ID_AES }, + { "sm4", CORE_ID_SM4 }, + { "sm3", CORE_ID_SM3 }, + { "hcq", CORE_ID_HCQ }, + { "qse", CORE_ID_QSE }, + { "pke", CORE_ID_PKE }, + { "ccp", CORE_ID_CCP }, +}; + +static int cmh_dt_core_id(const char *name) +{ + unsigned int i; + + for (i =3D 0; i < ARRAY_SIZE(cmh_dt_core_names); i++) + if (!strcmp(name, cmh_dt_core_names[i].name)) + return (int)cmh_dt_core_names[i].id; + return -1; +} + +/* + * Read a power-of-two DT geometry property (a slot count or a byte stride) + * and return its log2 -- the encoding the mailbox QUEUE registers use. An + * absent property yields @def_log2; a present value must be a power of two + * whose log2 lies in [@min_log2, @max_log2], otherwise -EINVAL. + */ +static int cmh_dt_read_geom_log2(const struct device_node *child, + const char *prop, u32 def_log2, + u32 min_log2, u32 max_log2, u32 *out_log2) +{ + u32 val; + + if (of_property_read_u32(child, prop, &val)) { + *out_log2 =3D def_log2; + return 0; + } + if (!is_power_of_2(val) || + ilog2(val) < min_log2 || ilog2(val) > max_log2) + return -EINVAL; + *out_log2 =3D ilog2(val); + return 0; +} + +/* + * Apply the per-mailbox rambus,cores affinity to the discovered cores. E= ach + * core ID listed on a mailbox pins that core's instance to the mailbox; a + * core listed on no mailbox keeps mbx =3D -1 (round-robin across mailboxe= s). + */ +static int cmh_config_apply_affinity(struct cmh_config *cfg) +{ + unsigned int mi, ci, inst; + + for (mi =3D 0; mi < cfg->mbx_count; mi++) { + struct cmh_mbx_config *m =3D &cfg->mailboxes[mi]; + + for (ci =3D 0; ci < m->num_cores; ci++) { + u32 core_id =3D m->cores[ci]; + struct cmh_core_type_cfg *ct; + int type; + + if (core_id >=3D CORE_ID_NUM || + core_id_to_type[core_id] < 0) { + dev_err(cmh_dev(), + "mbx[%u]: rambus,cores 0x%02x is not a dispatchable core\n", + mi, core_id); + return -EINVAL; + } + + type =3D core_id_to_type[core_id]; + ct =3D &cfg->core_types[type]; + + for (inst =3D 0; inst < ct->num_instances; inst++) { + if (ct->core_ids[inst] !=3D core_id) + continue; + if (ct->mbx[inst] >=3D 0) { + dev_err(cmh_dev(), + "core 0x%02x pinned to more than one mailbox\n", + core_id); + return -EINVAL; + } + ct->mbx[inst] =3D (s32)mi; + break; + } + } + } + +#ifdef CONFIG_CRYPTO_DEV_CMH_DEBUG + if (mbx_round_robin) { + unsigned int t, j; + + for (t =3D 0; t < CMH_NUM_CORE_TYPES; t++) + for (j =3D 0; j < cfg->core_types[t].num_instances; j++) + cfg->core_types[t].mbx[j] =3D -1; + dev_info(cmh_dev(), + "[debug] mbx_round_robin: dropped all rambus,cores affinity\n"); + } +#endif + + return 0; +} + +/* -- Validation -------------------------------------------------------- = */ + +static int cmh_config_validate_core_types(struct cmh_config *cfg) +{ + unsigned int i, j, k; + + for (i =3D 0; i < CMH_NUM_CORE_TYPES; i++) { + struct cmh_core_type_cfg *ct =3D &cfg->core_types[i]; + const char *name =3D core_type_names[i]; + + /* Zero instances is valid -- core absent from DT */ + if (ct->num_instances =3D=3D 0) + continue; + + if (ct->num_instances > CMH_MAX_CORE_INSTANCES) { + dev_err(cmh_dev(), "%s: num_instances %u > max %u\n", + name, ct->num_instances, + CMH_MAX_CORE_INSTANCES); + return -EINVAL; + } + + /* Validate MBX indices */ + for (j =3D 0; j < ct->num_instances; j++) { + if (ct->mbx[j] >=3D 0 && + (u32)ct->mbx[j] >=3D cfg->mbx_count) { +#ifdef CONFIG_CRYPTO_DEV_CMH_DEBUG + if (mbx_count_override > 0) { + dev_info(cmh_dev(), + "%s: mbx[%u]=3D%d >=3D overridden mbx_count %u, auto-assigning\n", + name, j, ct->mbx[j], + cfg->mbx_count); + ct->mbx[j] =3D -1; + continue; + } +#endif + dev_err(cmh_dev(), "%s: mbx[%u]=3D%d >=3D mbx_count %u\n", + name, j, ct->mbx[j], + cfg->mbx_count); + return -EINVAL; + } + } + + /* No duplicate core IDs within this type */ + for (j =3D 1; j < ct->num_instances; j++) { + for (k =3D 0; k < j; k++) { + if (ct->core_ids[j] =3D=3D ct->core_ids[k]) { + dev_err(cmh_dev(), + "%s: duplicate core_id 0x%02x at [%u] and [%u]\n", + name, ct->core_ids[j], + k, j); + return -EINVAL; + } + } + } + + /* No duplicate MBX within this type (if explicit) */ + for (j =3D 1; j < ct->num_instances; j++) { + if (ct->mbx[j] < 0) + continue; + for (k =3D 0; k < j; k++) { + if (ct->mbx[k] =3D=3D ct->mbx[j]) { + dev_err(cmh_dev(), + "%s: duplicate mbx %d at [%u] and [%u]\n", + name, ct->mbx[j], k, j); + return -EINVAL; + } + } + } + + /* All core IDs must fit in VCQ 8-bit field */ + for (j =3D 0; j < ct->num_instances; j++) { + if (ct->core_ids[j] > CORE_ID_MAX) { + dev_err(cmh_dev(), + "%s: core_ids[%u]=3D0x%02x > CORE_ID_MAX\n", + name, j, ct->core_ids[j]); + return -EINVAL; + } + } + } + + /* Cross-type: no core ID used by more than one type */ + for (i =3D 0; i < CMH_NUM_CORE_TYPES; i++) { + struct cmh_core_type_cfg *ct_i =3D &cfg->core_types[i]; + + for (j =3D i + 1; j < CMH_NUM_CORE_TYPES; j++) { + struct cmh_core_type_cfg *ct_j =3D &cfg->core_types[j]; + + for (k =3D 0; k < ct_i->num_instances; k++) { + unsigned int m; + + for (m =3D 0; m < ct_j->num_instances; m++) { + if (ct_i->core_ids[k] !=3D + ct_j->core_ids[m]) + continue; + dev_err(cmh_dev(), + "core_id 0x%02x conflict: %s[%u] and %s[%u]\n", + ct_i->core_ids[k], + core_type_names[i], k, + core_type_names[j], m); + return -EINVAL; + } + } + } + } + + return 0; +} + +static int cmh_config_validate(struct cmh_config *cfg) +{ + unsigned int i, j; + unsigned long max_instance_end; + + if (cfg->mbx_count =3D=3D 0 || cfg->mbx_count > CMH_MAX_CONFIGURED_MBX) { + dev_err(cmh_dev(), "mbx_count %u out of range (1..%u)\n", + cfg->mbx_count, CMH_MAX_CONFIGURED_MBX); + return -EINVAL; + } + + for (i =3D 0; i < cfg->mbx_count; i++) { + struct cmh_mbx_config *m =3D &cfg->mailboxes[i]; + + if (m->instance >=3D CMH_MAX_MBX_INSTANCES) { + dev_err(cmh_dev(), "mbx_instances[%u]=3D%u >=3D %u\n", + i, m->instance, CMH_MAX_MBX_INSTANCES); + return -EINVAL; + } + + if (m->slots_log2 < CMH_MBX_SLOTS_LOG2_MIN || + m->slots_log2 > CMH_MBX_SLOTS_LOG2_MAX) { + dev_err(cmh_dev(), "mbx_slots[%u]=3D%u out of range (%u..%u)\n", + i, m->slots_log2, + CMH_MBX_SLOTS_LOG2_MIN, CMH_MBX_SLOTS_LOG2_MAX); + return -EINVAL; + } + + if (m->stride_log2 < CMH_MBX_STRIDE_LOG2_MIN || + m->stride_log2 > CMH_MBX_STRIDE_LOG2_MAX) { + dev_err(cmh_dev(), "mbx_strides[%u]=3D%u out of range (%u..%u)\n", + i, m->stride_log2, + CMH_MBX_STRIDE_LOG2_MIN, CMH_MBX_STRIDE_LOG2_MAX); + return -EINVAL; + } + + /* Check for duplicate instance indices */ + for (j =3D 0; j < i; j++) { + if (cfg->mailboxes[j].instance =3D=3D m->instance) { + dev_err(cmh_dev(), "duplicate instance %u at indices %u and %u\n", + m->instance, j, i); + return -EINVAL; + } + } + } + + /* Ensure SIC region is large enough for all requested instances */ + max_instance_end =3D 0; + for (i =3D 0; i < cfg->mbx_count; i++) { + unsigned long end =3D ((unsigned long)cfg->mailboxes[i].instance + 1) + << CMH_MBX_INSTANCE_SHIFT; + if (end > max_instance_end) + max_instance_end =3D end; + } + + if (max_instance_end > cfg->sic_size) { + dev_err(cmh_dev(), "sic_size 0x%zx too small for instance requiring 0x%l= x\n", + cfg->sic_size, max_instance_end); + return -EINVAL; + } + + return 0; +} + +/* -- Public Interface -------------------------------------------------- = */ + +/** + * cmh_config_init() - Initialize device configuration from platform/DT da= ta + * @cfg: Configuration structure to populate + * @pdev: Platform device providing DT node and resources + * + * Parse the "rambus,cmh-v1030" device tree node for MMIO base address, + * interrupt specifiers, and per-mailbox properties (instance indices, slo= t counts, + * strides). When DT properties are absent, fall back to module parameter + * arrays. Populate per-core-type instance configuration from module + * parameters, then validate the complete configuration. + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_config_init(struct cmh_config *cfg, struct platform_device *pdev) +{ + struct device_node *np =3D pdev->dev.of_node; + struct resource *res; + struct device_node *child; + int ret; + + if (!np) { + dev_err(&pdev->dev, "device tree node required\n"); + return -ENODEV; + } + + /* SIC base + size from DT "reg" property (mandatory) */ + res =3D platform_get_resource(pdev, IORESOURCE_MEM, 0); + if (!res) { + dev_err(cmh_dev(), "missing DT reg resource\n"); + return -EINVAL; + } + cfg->sic_base =3D res->start; + cfg->sic_size =3D resource_size(res); + + /* + * Interrupts are per-mailbox (declared in each mailbox child node) + * and resolved per MBX by cmh_rh_resolve_irqs(). There is no + * top-level interrupt; if no mailbox has one the response handler + * falls back to watchdog-timer polling. + */ + cfg->sic_mapped =3D NULL; + cfg->fw_ready_timeout_ms =3D CMH_FW_READY_TIMEOUT_MS; + + /* -- Mailbox configuration from DT child nodes ----------------- */ + + cfg->mbx_count =3D 0; + for_each_child_of_node(np, child) { + struct cmh_mbx_config *m; + int nc, ci; + + if (!of_node_name_eq(child, "queue")) + continue; + + if (cfg->mbx_count >=3D CMH_MAX_CONFIGURED_MBX) { + dev_err(cmh_dev(), + "too many mailbox nodes in DT (max %u)\n", + CMH_MAX_CONFIGURED_MBX); + of_node_put(child); + return -EINVAL; + } + m =3D &cfg->mailboxes[cfg->mbx_count]; + + ret =3D of_property_read_u32(child, "reg", &m->instance); + if (ret) { + dev_err(cmh_dev(), "mailbox %pOFn: missing 'reg'\n", + child); + of_node_put(child); + return ret; + } + + ret =3D cmh_dt_read_geom_log2(child, "rambus,num-slots", + CMH_DEFAULT_SLOTS_LOG2, + CMH_MBX_SLOTS_LOG2_MIN, + CMH_MBX_SLOTS_LOG2_MAX, + &m->slots_log2); + if (ret) { + dev_err(cmh_dev(), + "mailbox %u: rambus,num-slots not power-of-2 in %u..%u\n", + m->instance, 1U << CMH_MBX_SLOTS_LOG2_MIN, + 1U << CMH_MBX_SLOTS_LOG2_MAX); + of_node_put(child); + return ret; + } + + ret =3D cmh_dt_read_geom_log2(child, "rambus,slot-stride-bytes", + CMH_DEFAULT_STRIDE_LOG2, + CMH_MBX_STRIDE_LOG2_MIN, + CMH_MBX_STRIDE_LOG2_MAX, + &m->stride_log2); + if (ret) { + dev_err(cmh_dev(), + "mailbox %u: rambus,slot-stride-bytes not power-of-2 in %u..%u\n", + m->instance, 1U << CMH_MBX_STRIDE_LOG2_MIN, + 1U << CMH_MBX_STRIDE_LOG2_MAX); + of_node_put(child); + return ret; + } + +#ifdef CONFIG_CRYPTO_DEV_CMH_DEBUG + if (mbx_slots_override > 0) + m->slots_log2 =3D mbx_slots_override; +#endif + + /* Optional per-mailbox interrupt (absent =3D> polling). */ + m->irq =3D of_irq_get(child, 0); + if (m->irq =3D=3D -EPROBE_DEFER) { + of_node_put(child); + return -EPROBE_DEFER; + } + if (m->irq < 0) + m->irq =3D -1; + + /* Optional rambus,cores affinity list (core-name strings). */ + m->num_cores =3D 0; + nc =3D of_property_count_strings(child, "rambus,cores"); + if (nc > 0) { + if (nc > CMH_NUM_CORE_TYPES) { + dev_err(cmh_dev(), + "mailbox %u: too many rambus,cores (%d > %u)\n", + m->instance, nc, CMH_NUM_CORE_TYPES); + of_node_put(child); + return -EINVAL; + } + for (ci =3D 0; ci < nc; ci++) { + const char *cname; + int cid; + + of_property_read_string_index(child, + "rambus,cores", + ci, &cname); + cid =3D cmh_dt_core_id(cname); + if (cid < 0) { + dev_err(cmh_dev(), + "mailbox %u: unknown rambus,cores core '%s'\n", + m->instance, cname); + of_node_put(child); + return -EINVAL; + } + m->cores[ci] =3D (u32)cid; + } + m->num_cores =3D nc; + } + + m->queue_size =3D (1UL << m->slots_log2) << m->stride_log2; + m->dma_handle =3D 0; + m->virt_addr =3D NULL; + m->reg_base =3D NULL; + cfg->mbx_count++; + } + + if (cfg->mbx_count =3D=3D 0) { + dev_err(cmh_dev(), "no mailbox child nodes in DT\n"); + return -EINVAL; + } + +#ifdef CONFIG_CRYPTO_DEV_CMH_DEBUG + if (mbx_count_override > 0) { + if (mbx_count_override > cfg->mbx_count) { + dev_err(cmh_dev(), + "mbx_count_override %u > DT count %u\n", + mbx_count_override, cfg->mbx_count); + return -EINVAL; + } + dev_info(cmh_dev(), "[debug] overriding mbx_count: %u -> %u\n", + cfg->mbx_count, mbx_count_override); + cfg->mbx_count =3D mbx_count_override; + } +#endif + + /* + * Cores are discovered later (cmh_config_discover_cores) once the + * SIC region is mapped and CORE_ENABLE is readable. Here we only + * validate the mailbox configuration. + */ + return cmh_config_validate(cfg); +} + +/** + * cmh_config_discover_cores() - Enumerate cores from the CORE_ENABLE regi= ster + * @cfg: Configuration structure (SIC must already be mapped) + * + * Reads the SIC CORE_ENABLE register to determine which crypto core types + * the silicon build provides, populates cfg->core_types[], applies the + * per-mailbox rambus,cores affinity, and validates the result. Called af= ter + * the SIC ioremap. + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_config_discover_cores(struct cmh_config *cfg) +{ + u32 core_enable; + unsigned int type; + int ret; + + if (!cfg->sic_mapped) { + dev_err(cmh_dev(), "discover_cores: SIC not mapped\n"); + return -EINVAL; + } + + core_enable =3D cmh_reg_read32(cfg->sic_mapped, R_SIC_CORE_ENABLE); + dev_dbg(cmh_dev(), "CORE_ENABLE=3D0x%08x\n", core_enable); + + for (type =3D 0; type < CMH_NUM_CORE_TYPES; type++) { + const struct cmh_core_desc *d =3D &cmh_core_descs[type]; + struct cmh_core_type_cfg *ct =3D &cfg->core_types[type]; + + if (!d->enable_mask) + continue; + if (!cmh_core_enabled(core_enable, d->enable_mask)) + continue; + + ct->core_ids[0] =3D d->core_id; + ct->mbx[0] =3D -1; + ct->num_instances =3D 1; + dev_dbg(cmh_dev(), "core %s (0x%02x) present\n", + core_type_names[type], d->core_id); + } + + ret =3D cmh_config_apply_affinity(cfg); + if (ret) + return ret; + + return cmh_config_validate_core_types(cfg); +} diff --git a/drivers/crypto/cmh/cmh_debugfs.c b/drivers/crypto/cmh/cmh_debu= gfs.c new file mode 100644 index 000000000000..19257d794843 --- /dev/null +++ b/drivers/crypto/cmh/cmh_debugfs.c @@ -0,0 +1,289 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- debugfs Per-MBX Counters and Fault Injection + * + * Creates the /sys/kernel/debug/cmh/ tree with: + * mbxN/vcqs_submitted (ro) Total VCQs sent to MBX N + * mbxN/vcqs_completed (ro) Total completions received + * mbxN/vcqs_errors (ro) Total error completions + * mbxN/queue_full_count (ro) Times select_mailbox() skipped this MBX + * mbxN/max_queue_depth (ro) High-water mark of in-flight transactions + * mbxN/inject_abort (wo) Write any value to inject MBX_COMMAND_ABORT + * mbxN/force_drain (wo) Write any value to force-drain all pending= txns + * tm/cmq_posts (ro) Total cmh_tm_post_command() calls + * tm/cmq_depth_max (ro) High-water mark of CMQ length + * tm/cmq_eagain_count (ro) Times CMQ was full (-EAGAIN) + * tm/backoff_count (ro) Times TM backed off (all MBX queues full) + * tm/async_timeout_count (ro) Async requests that timed out + * + * This file is only compiled when CONFIG_CRYPTO_DEV_CMH_DEBUG=3Dy (see Kb= uild). + * Requires CONFIG_DEBUG_FS=3Dy in the kernel (standard for dev builds). + */ + +#include +#include +#include +#include + +#include "cmh_debugfs.h" +#include "cmh_config.h" +#include "cmh_registers.h" +#include "cmh_dma.h" +#include "cmh_txn.h" +#include "cmh_rh.h" +#include "cmh_rng.h" + +/* -- Module State -------------------------------------------------------= --- */ + +static struct { + struct dentry *root; /* /sys/kernel/debug/cmh/ */ + struct cmh_mbx_stats *mbx; /* array[mbx_count] */ + struct cmh_tm_stats tm; + struct cmh_config *cfg; /* for inject_abort register access */ + u32 mbx_count; +} dbgfs; + +/* -- debugfs file ops for atomic64_t ------------------------------------= --- */ + +static int cmh_dbgfs_u64_get(void *data, u64 *val) +{ + *val =3D (u64)atomic64_read((atomic64_t *)data); + return 0; +} + +DEFINE_DEBUGFS_ATTRIBUTE(cmh_dbgfs_u64_ro_fops, + cmh_dbgfs_u64_get, NULL, "%llu\n"); + +/* -- Per-MBX directory --------------------------------------------------= --- */ + +/* + * inject_abort -- write-only debugfs file for fault injection. + * + * Writing any value triggers MBX_COMMAND_ABORT on this mailbox. + * The eSW calls mbx_abort() -> mbx_cmd_error(mbx, -EPIPE), fires the + * error IRQ, and the LKM RH completes in-flight transactions with -EIO + * then issues MBX_COMMAND_RESTART to resume the mailbox. + * + * Private data points to the MBX index (cast to void *). + */ +static ssize_t inject_abort_write(struct file *file, + const char __user *ubuf, + size_t count, loff_t *ppos) +{ + u32 idx =3D (u32)(unsigned long)file->private_data; + void __iomem *base; + + if (!dbgfs.cfg || idx >=3D dbgfs.cfg->mbx_count) + return -EINVAL; + + base =3D dbgfs.cfg->mailboxes[idx].reg_base; + dev_warn(cmh_dev(), "debugfs: injecting ABORT on mbx[%u]\n", idx); + cmh_reg_write32(MBX_COMMAND_ABORT, base, R_MBX_COMMAND); + + return count; +} + +static const struct file_operations inject_abort_fops =3D { + .owner =3D THIS_MODULE, + .open =3D simple_open, + .write =3D inject_abort_write, + .llseek =3D noop_llseek, +}; + +/* + * force_drain -- write-only debugfs file for administrative recovery. + * + * Writing any value issues MBX_COMMAND_FLUSH, drains all pending + * transactions on this mailbox (completing each with -ECANCELED), + * and resets all recovery bookkeeping (abort_stall_ticks, + * restart_pending, restart_retries, flush_count, wedged). + * + * Use this to recover D-state processes when the eSW is dead and + * normal ABORT/RESTART escalation has not recovered the mailbox. + */ +static ssize_t force_drain_write(struct file *file, + const char __user *ubuf, + size_t count, loff_t *ppos) +{ + u32 idx =3D (u32)(unsigned long)file->private_data; + + if (!dbgfs.cfg || idx >=3D dbgfs.cfg->mbx_count) + return -EINVAL; + + cmh_rh_force_drain_mbx(idx); + return count; +} + +static const struct file_operations force_drain_fops =3D { + .owner =3D THIS_MODULE, + .open =3D simple_open, + .write =3D force_drain_write, + .llseek =3D noop_llseek, +}; + +static void create_mbx_dir(u32 idx, struct dentry *parent) +{ + struct cmh_mbx_stats *s =3D &dbgfs.mbx[idx]; + struct dentry *d; + char name[16]; + + snprintf(name, sizeof(name), "mbx%u", idx); + d =3D debugfs_create_dir(name, parent); + + debugfs_create_file("vcqs_submitted", 0444, d, + &s->vcqs_submitted, &cmh_dbgfs_u64_ro_fops); + debugfs_create_file("vcqs_completed", 0444, d, + &s->vcqs_completed, &cmh_dbgfs_u64_ro_fops); + debugfs_create_file("vcqs_errors", 0444, d, + &s->vcqs_errors, &cmh_dbgfs_u64_ro_fops); + debugfs_create_file("queue_full_count", 0444, d, + &s->queue_full_count, &cmh_dbgfs_u64_ro_fops); + debugfs_create_file("max_queue_depth", 0444, d, + &s->max_queue_depth, &cmh_dbgfs_u64_ro_fops); + debugfs_create_file("inject_abort", 0200, d, + (void *)(uintptr_t)idx, &inject_abort_fops); + debugfs_create_file("force_drain", 0200, d, + (void *)(uintptr_t)idx, &force_drain_fops); +} + +/* -- TM directory -------------------------------------------------------= --- */ + +static void create_tm_dir(struct dentry *parent) +{ + struct cmh_tm_stats *s =3D &dbgfs.tm; + struct dentry *d; + + d =3D debugfs_create_dir("tm", parent); + + debugfs_create_file("cmq_posts", 0444, d, + &s->cmq_posts, &cmh_dbgfs_u64_ro_fops); + debugfs_create_file("cmq_depth_max", 0444, d, + &s->cmq_depth_max, &cmh_dbgfs_u64_ro_fops); + debugfs_create_file("cmq_eagain_count", 0444, d, + &s->cmq_eagain_count, &cmh_dbgfs_u64_ro_fops); + debugfs_create_file("backoff_count", 0444, d, + &s->backoff_count, &cmh_dbgfs_u64_ro_fops); + debugfs_create_file("async_timeout_count", 0444, d, + &s->async_timeout_count, &cmh_dbgfs_u64_ro_fops); +} + +/* -- Config directory: timeout tuning ---------------------------------- = */ + +static void create_config_dir(struct dentry *parent) +{ + struct dentry *d; + + d =3D debugfs_create_dir("config", parent); + + /* TM timeouts */ + debugfs_create_u32("async_timeout_ms", 0644, d, + cmh_tm_timeout_async_ptr()); + debugfs_create_u32("vcq_timeout_ms", 0644, d, + cmh_tm_timeout_vcq_ptr()); + debugfs_create_u32("slow_op_timeout_ms", 0644, d, + cmh_tm_timeout_slow_op_ptr()); + debugfs_create_u32("drain_timeout_ms", 0644, d, + cmh_tm_timeout_drain_ptr()); + + /* RH watchdog */ + debugfs_create_u32("watchdog_ms", 0644, d, + cmh_rh_timeout_watchdog_ptr()); + + /* DRBG timeout */ + debugfs_create_u32("drbg_timeout_ms", 0644, d, + cmh_rng_timeout_drbg_ptr()); + + /* TM queue depths */ + debugfs_create_u32("cmq_max_depth", 0644, d, + cmh_tm_cmq_max_depth_ptr()); + debugfs_create_u32("backlog_max_depth", 0644, d, + cmh_tm_backlog_max_depth_ptr()); +} + +/* -- Public Interface ---------------------------------------------------= --- */ + +/** + * cmh_debugfs_init() - Create debugfs directory hierarchy for CMH + * @cfg: Platform configuration containing mailbox count and register base= s. + * + * Allocates per-mailbox statistics and creates the /sys/kernel/debug/cmh/ + * tree with per-mailbox counters, fault-injection files, and transaction + * manager statistics. debugfs is optional; failure to create entries does + * not prevent module initialisation. + */ +void cmh_debugfs_init(struct cmh_config *cfg) +{ + u32 mbx_count =3D cfg->mbx_count; + u32 i; + + dbgfs.root =3D debugfs_create_dir("cmh", NULL); + if (IS_ERR_OR_NULL(dbgfs.root)) { + if (!IS_ERR(dbgfs.root)) + dev_warn(cmh_dev(), "debugfs: creation returned NULL -- counters disabl= ed\n"); + else + dev_warn(cmh_dev(), "debugfs: creation failed (rc=3D%ld) -- counters di= sabled\n", + PTR_ERR(dbgfs.root)); + dbgfs.root =3D NULL; + return; /* debugfs is optional -- never fail module init */ + } + + dbgfs.mbx_count =3D mbx_count; + dbgfs.cfg =3D cfg; + dbgfs.mbx =3D kcalloc(mbx_count, sizeof(*dbgfs.mbx), GFP_KERNEL); + if (!dbgfs.mbx) { + debugfs_remove_recursive(dbgfs.root); + dbgfs.root =3D NULL; + return; + } + + for (i =3D 0; i < mbx_count; i++) + create_mbx_dir(i, dbgfs.root); + + create_tm_dir(dbgfs.root); + + create_config_dir(dbgfs.root); + + dev_dbg(cmh_dev(), "debugfs: initialized (%u mailboxes)\n", mbx_count); +} + +/** + * cmh_debugfs_cleanup() - Remove all CMH debugfs entries + * + * Tears down the /sys/kernel/debug/cmh/ tree and frees per-mailbox + * statistics memory. Safe to call even if cmh_debugfs_init() was never + * called or failed. + */ +void cmh_debugfs_cleanup(void) +{ + debugfs_remove_recursive(dbgfs.root); + dbgfs.root =3D NULL; + kfree(dbgfs.mbx); + dbgfs.mbx =3D NULL; + dev_dbg(cmh_dev(), "debugfs: cleaned up\n"); +} + +/** + * cmh_debugfs_mbx_stats() - Return per-mailbox statistics pointer + * @mbx_idx: Zero-based mailbox index. + * + * Return: Pointer to the statistics structure for @mbx_idx, or NULL if + * debugfs is disabled or @mbx_idx is out of range. + */ +struct cmh_mbx_stats *cmh_debugfs_mbx_stats(u32 mbx_idx) +{ + if (!dbgfs.mbx || mbx_idx >=3D dbgfs.mbx_count) + return NULL; + return &dbgfs.mbx[mbx_idx]; +} + +/** + * cmh_debugfs_tm_stats() - Return transaction manager statistics pointer + * + * Return: Pointer to the singleton TM statistics structure. The pointer + * is always valid (points to static storage). + */ +struct cmh_tm_stats *cmh_debugfs_tm_stats(void) +{ + return &dbgfs.tm; +} diff --git a/drivers/crypto/cmh/cmh_dma.c b/drivers/crypto/cmh/cmh_dma.c new file mode 100644 index 000000000000..8707a39622e4 --- /dev/null +++ b/drivers/crypto/cmh/cmh_dma.c @@ -0,0 +1,379 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- DMA Operations + * + * Implements the cmh_dma.h interface using the kernel DMA API + * (dma_map_single, dma_alloc_coherent, etc.). + * + * Scatterlist linearization rationale + * ------------------------------------ + * The eSW firmware supports SCATTERGATHER commands for all core + * types (AES_CMD_SCATTERGATHER, SM4_CMD_SCATTERGATHER, + * CCP_CMD_SCATTERGATHER, HC_CMD_GATHER), using a proprietary + * linked-list-item (LLI) descriptor chain format. The hash driver + * already uses this via cmh_dma_build_sg() + HC_CMD_GATHER. + * + * For symmetric cipher and AEAD commands, the LKM currently + * linearizes scatterlist input into contiguous bounce buffers via + * scatterwalk_map_and_copy() rather than building LLI chains from + * kernel scatterlists. This is a deliberate first-submission + * simplification with a concrete technical justification: + * + * - The hash SG path is unidirectional (DMA_TO_DEVICE gather only). + * Skcipher and AEAD require bidirectional handling: separate src + * and dst scatterlists (which may alias for in-place operations), + * plus AAD and authentication tag regions with distinct DMA + * directions and alignment constraints. + * - The CMH LLI format requires 64-byte aligned descriptor chain + * pointers (the .lli field) with 32-bit length fields. This + * alignment is automatically satisfied by dma_alloc_coherent() + * for the descriptor array; data buffer addresses have no + * hardware alignment requirement. Kernel SG entries have no + * alignment guarantee for data, so direct SG-to-LLI translation + * requires per-segment validation, potential splitting at + * descriptor boundaries, and separate chains for src/dst/AAD -- + * substantially more complex than the unidirectional hash + * gather case. + * - Each skcipher/AEAD driver caps linearization at + * CMH_AES_MAX_CRYPTLEN / CMH_SM4_MAX_CRYPTLEN (32 MiB). + * Requests exceeding this cap are rejected with -EINVAL. + * In practice, crypto API callers (dm-crypt, IPsec, kernel TLS) + * send page-sized or smaller buffers, so the bounce allocation + * is typically <=3D PAGE_SIZE and succeeds even under GFP_ATOMIC. + * + * A shared SG-to-LLI adapter handling bidirectional mappings, + * alignment splitting, and in-place src=3D=3Ddst detection for the + * skcipher/AEAD/MAC paths is planned as a follow-up series once the + * core driver is accepted. + * + * This linearization pattern is consistent with other upstream HW + * crypto drivers that use bounce buffers in their initial + * submissions (e.g. ccree, sa2ul, omap-aes). + */ + +#include +#include +#include +#include +#include +#include + +#include "cmh_dma.h" + +/* Module-global device pointer, set in cmh_dma_init() */ +static struct device *cmh_device; + +/** + * cmh_dma_init() - Initialize the standard DMA backend + * @pdev: Platform device providing the struct device for DMA ops + * + * Stores the device pointer for use by all DMA wrapper functions. + * + * Return: 0 (always succeeds for the standard backend). + */ +int cmh_dma_init(struct platform_device *pdev) +{ + cmh_device =3D &pdev->dev; + return 0; +} + +/** + * cmh_dma_cleanup() - Tear down the standard DMA backend + * + * Clears the stored device pointer. + */ +void cmh_dma_cleanup(void) +{ + cmh_device =3D NULL; +} + +/** + * cmh_dev() - Return the platform device pointer + * + * Return: struct device pointer, or NULL outside probe/remove lifecycle. + */ +struct device *cmh_dev(void) +{ + return cmh_device; +} + +/* -- Streaming DMA ------------------------------------------------------= -- */ + +/** + * cmh_dma_map_single() - Map a kernel buffer for streaming DMA + * @buf: Kernel virtual address + * @size: Buffer length in bytes + * @dir: DMA direction + * + * Return: DMA address, or a DMA_MAPPING_ERROR value on failure. + */ +dma_addr_t cmh_dma_map_single(void *buf, size_t size, + enum dma_data_direction dir) +{ + return dma_map_single(cmh_device, buf, size, dir); +} + +/** + * cmh_dma_unmap_single() - Unmap a streaming DMA buffer + * @addr: DMA address returned by cmh_dma_map_single() + * @size: Buffer length in bytes + * @dir: DMA direction (must match the map call) + */ +void cmh_dma_unmap_single(dma_addr_t addr, size_t size, + enum dma_data_direction dir) +{ + dma_unmap_single(cmh_device, addr, size, dir); +} + +/** + * cmh_dma_sync_for_cpu() - Sync a DMA buffer for CPU access + * @addr: DMA address of the mapped buffer + * @size: Region length in bytes + * @dir: DMA direction + */ +void cmh_dma_sync_for_cpu(dma_addr_t addr, size_t size, + enum dma_data_direction dir) +{ + dma_sync_single_for_cpu(cmh_device, addr, size, dir); +} + +/** + * cmh_dma_sync_for_device() - Sync a DMA buffer for device access + * @addr: DMA address of the mapped buffer + * @size: Region length in bytes + * @dir: DMA direction + */ +void cmh_dma_sync_for_device(dma_addr_t addr, size_t size, + enum dma_data_direction dir) +{ + dma_sync_single_for_device(cmh_device, addr, size, dir); +} + +/** + * cmh_dma_map_error() - Check whether a DMA mapping failed + * @addr: DMA address to check + * + * Return: Non-zero if @addr indicates a mapping error. + */ +int cmh_dma_map_error(dma_addr_t addr) +{ + return dma_mapping_error(cmh_device, addr); +} + +/* -- Coherent DMA -------------------------------------------------------= -- */ + +/** + * cmh_dma_alloc() - Allocate coherent DMA memory + * @size: Allocation size in bytes + * @handle: Output DMA address + * @gfp: GFP allocation flags + * + * Return: Kernel virtual address, or NULL on failure. + */ +void *cmh_dma_alloc(size_t size, dma_addr_t *handle, gfp_t gfp) +{ + return dma_alloc_coherent(cmh_device, size, handle, gfp); +} + +/** + * cmh_dma_free() - Free coherent DMA memory + * @size: Allocation size (must match cmh_dma_alloc) + * @virt: Kernel virtual address + * @handle: DMA address + */ +void cmh_dma_free(size_t size, void *virt, dma_addr_t handle) +{ + dma_free_coherent(cmh_device, size, virt, handle); +} + +/* -- Buffer write helpers -----------------------------------------------= -- */ + +/** + * cmh_dma_write() - Copy data into a DMA buffer + * @dst: Destination (from cmh_dma_alloc) + * @src: Source kernel buffer + * @len: Byte count + */ +void cmh_dma_write(void *dst, const void *src, size_t len) +{ + memcpy(dst, src, len); +} + +/** + * cmh_dma_fence() - No-op on standard DMA API platforms (coherent) + * @ptr: Unused -- present for interface compatibility + */ +void cmh_dma_fence(void *ptr) +{ + /* Standard DMA API: coherent memory, no cross-slave fence needed */ +} + +/** + * cmh_dma_zero() - Zero a DMA buffer + * @dst: Destination (from cmh_dma_alloc) + * @len: Byte count + */ +void cmh_dma_zero(void *dst, size_t len) +{ + memset(dst, 0, len); +} + +/** + * cmh_dma_build_sg() - Build a scatter-gather DMA mapping + * @bufs: Array of buffer descriptors to map + * @count: Number of entries in @bufs + * @gfp: GFP flags for memory allocation + * + * Allocates a streaming-DMA descriptor array and maps each buffer in @bufs + * for DMA-to-device transfer, filling CMH eSW-format scatter-gather + * descriptors with linked-list pointers. + * + * The descriptor array uses streaming DMA (kmalloc + dma_map_single) rath= er + * than dma_alloc_coherent so that cmh_dma_free_sg() -- which calls + * dma_unmap_single + kfree -- is safe from any context including BH-disab= led + * completion callbacks. + * + * Return: Pointer to the allocated cmh_sg_map on success, NULL on failure. + */ +struct cmh_sg_map *cmh_dma_build_sg(const struct cmh_dma_buf *bufs, u32 co= unt, + gfp_t gfp) +{ + struct cmh_sg_map *sgm; + u32 i; + + if (!count) + return NULL; + + sgm =3D kzalloc(struct_size(sgm, bufs, count), gfp); + if (!sgm) + return NULL; + + sgm->count =3D count; + sgm->items_size =3D array_size(count, sizeof(*sgm->items)); + if (sgm->items_size =3D=3D SIZE_MAX) + goto err_free_sgm; + + /* + * Allocate descriptor array with kmalloc and map for streaming DMA. + * We map first to obtain items_dma (needed for .lli pointers), + * then sync-for-cpu, fill descriptors, and sync-for-device. + * + * The eSW SG walker (struct dma_scattergather_item: u64 lli/src/dst/ + * len) requires only natural u64/4-byte alignment of the descriptor + * and its LLI chain pointers -- there is no 64-byte-per-descriptor + * requirement -- so a plain kzalloc (>=3D ARCH_KMALLOC_MINALIGN) is + * sufficient and DMA-safe. + */ + sgm->items =3D kzalloc(sgm->items_size, gfp); + if (!sgm->items) + goto err_free_sgm; + + sgm->items_dma =3D cmh_dma_map_single(sgm->items, sgm->items_size, + DMA_TO_DEVICE); + if (cmh_dma_map_error(sgm->items_dma)) + goto err_free_items; + + /* Map each source buffer for device read */ + for (i =3D 0; i < count; i++) { + dma_addr_t dma; + + if (!bufs[i].len) + goto err_unmap; + sgm->bufs[i].len =3D bufs[i].len; + dma =3D cmh_dma_map_single(bufs[i].data, bufs[i].len, + DMA_TO_DEVICE); + if (cmh_dma_map_error(dma)) + goto err_unmap; + sgm->bufs[i].dma =3D dma; + } + + /* + * Reclaim CPU ownership of the descriptor buffer. After + * dma_map_single the device owns the mapping; we must call + * sync_for_cpu before writing regardless of direction. The + * direction matches the original mapping (DMA_TO_DEVICE) -- + * this tells the DMA layer which cache operations apply: + * invalidate so the CPU sees coherent data before we fill + * the SG descriptors and later sync_for_device. + */ + cmh_dma_sync_for_cpu(sgm->items_dma, sgm->items_size, + DMA_TO_DEVICE); + + /* Fill CMH eSW SG descriptors */ + for (i =3D 0; i < count; i++) { + u64 lli_val; + + if (i + 1 < count) + lli_val =3D (u64)(sgm->items_dma + + (i + 1) * sizeof(*sgm->items)); + else + lli_val =3D 0; + + sgm->items[i].lli =3D lli_val; + sgm->items[i].src =3D (u64)sgm->bufs[i].dma; + sgm->items[i].dst =3D 0; + sgm->items[i].len =3D (u64)bufs[i].len; + } + + /* Flush descriptor writes to device */ + cmh_dma_sync_for_device(sgm->items_dma, sgm->items_size, + DMA_TO_DEVICE); + + return sgm; + +err_unmap: + while (i--) + cmh_dma_unmap_single(sgm->bufs[i].dma, + sgm->bufs[i].len, DMA_TO_DEVICE); + cmh_dma_unmap_single(sgm->items_dma, sgm->items_size, + DMA_TO_DEVICE); +err_free_items: + kfree(sgm->items); +err_free_sgm: + kfree(sgm); + return NULL; +} + +/** + * cmh_dma_free_sg() - Unmap and free a scatter-gather mapping + * @sgm: Scatter-gather mapping created by cmh_dma_build_sg(), or NULL + * + * Unmaps all DMA-mapped buffers, unmaps and frees the descriptor array, + * and releases the cmh_sg_map structure. Safe to call from any context + * (including BH-disabled completion callbacks) because it uses only + * dma_unmap_single + kfree -- no vunmap/dma_free_coherent. + */ +void cmh_dma_free_sg(struct cmh_sg_map *sgm) +{ + u32 i; + + if (!sgm) + return; + + for (i =3D 0; i < sgm->count; i++) + cmh_dma_unmap_single(sgm->bufs[i].dma, + sgm->bufs[i].len, DMA_TO_DEVICE); + + cmh_dma_unmap_single(sgm->items_dma, sgm->items_size, + DMA_TO_DEVICE); + kfree(sgm->items); + kfree(sgm); +} + +/** + * cmh_dma_orphan_free() - Orphan cleanup callback for abandoned DMA buffe= rs + * @data: Pointer to a struct cmh_dma_orphan describing the orphaned mappi= ng + * + * Called by the transaction manager when a synchronous operation times out + * and the caller has already returned. Unmaps the DMA buffer and frees + * the backing memory and the orphan descriptor itself. + */ +void cmh_dma_orphan_free(void *data) +{ + struct cmh_dma_orphan *o =3D data; + + cmh_dma_unmap_single(o->addr, o->len, o->dir); + kfree_sensitive(o->buf); + kfree(o); +} diff --git a/drivers/crypto/cmh/cmh_main.c b/drivers/crypto/cmh/cmh_main.c new file mode 100644 index 000000000000..c037e8daa877 --- /dev/null +++ b/drivers/crypto/cmh/cmh_main.c @@ -0,0 +1,343 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Platform Driver Entry and Exit + * + * Responsibilities: + * - Match "rambus,cmh-v1030" DT node via platform_driver + * - Parse device-tree properties via cmh_config_init() + * - ioremap the SIC region + * - Verify CMH boot status (sanity check) + * - Compute per-instance register bases + * - Initialize MBX queues (MQI) + * - Start Transaction Manager kthread + * - Register Response Handler IRQ + * - Register Kernel Crypto API hash algorithms + * - Clean up in reverse order on exit or error + */ + +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "cmh.h" +#include "cmh_dma.h" +#include "cmh_mqi.h" +#include "cmh_txn.h" +#include "cmh_rh.h" +#include "cmh_registers.h" +#include "cmh_debugfs.h" +#include "cmh_sysfs.h" + +#include + +#ifdef CONFIG_CRYPTO_DEV_CMH_DEBUG +static bool skip_fw_check; +module_param(skip_fw_check, bool, 0444); +MODULE_PARM_DESC(skip_fw_check, + "[debug] Skip eSW boot status check at probe (default: false)"); +#else +#define skip_fw_check false +#endif + +/* SIC Sanity Check */ + +static int cmh_check_sic(struct cmh_config *cfg) +{ + const u32 ready =3D SIC_SW_BOOT_STATUS_MISSION | + SIC_SW_BOOT_STATUS_MISSION2; + u32 boot_status; + u32 sw_boot; + int ret; + + boot_status =3D cmh_reg_read32(cfg->sic_mapped, R_SIC_BOOT_STATUS); + + if ((boot_status & SIC_BOOT_STATUS_MASK) !=3D SIC_BOOT_STATUS_PASS) { + dev_err(cmh_dev(), "SIC boot status check failed (0x%02x !=3D 0x%02x)\n", + boot_status & SIC_BOOT_STATUS_MASK, SIC_BOOT_STATUS_PASS); + return -EIO; + } + + /* + * Wait for eSW readiness: MISSION signals the primary VCQ engine, + * MISSION2 the sidecar engine (set asynchronously). The driver + * uses both, so require both bits. + */ + ret =3D read_poll_timeout(ioread32, sw_boot, + (sw_boot & ready) =3D=3D ready, + 1000, + (unsigned long)cfg->fw_ready_timeout_ms * 1000UL, + false, + cfg->sic_mapped + R_SIC_SW_BOOT_STATUS); + if (ret) { + sw_boot =3D cmh_reg_read32(cfg->sic_mapped, R_SIC_SW_BOOT_STATUS); + dev_err(cmh_dev(), "CMH eSW not ready (sw_boot_status=3D0x%08x, timeout= =3D%ums)\n", + sw_boot, cfg->fw_ready_timeout_ms); + return -ETIMEDOUT; + } + + return 0; +} + +/* Module Init -- platform driver probe */ + +static int cmh_probe(struct platform_device *pdev) +{ + struct cmh_device *dev; + struct cmh_config *cfg; + struct clk_bulk_data *clks; + struct gpio_desc *reset; + unsigned int i; + int ret; + + dev =3D devm_kzalloc(&pdev->dev, sizeof(*dev), GFP_KERNEL); + if (!dev) + return -ENOMEM; + + dev->dev =3D &pdev->dev; + cfg =3D &dev->config; + + /* Declare DMA addressing capability */ + ret =3D dma_set_mask_and_coherent(&pdev->dev, DMA_BIT_MASK(64)); + if (ret) { + ret =3D dev_err_probe(&pdev->dev, ret, + "dma_set_mask_and_coherent failed\n"); + goto err_free_dev; + } + + /* Initialize DMA backend (standard API or FPGA pool) */ + ret =3D cmh_dma_init(pdev); + if (ret) { + ret =3D dev_err_probe(&pdev->dev, ret, "DMA init failed\n"); + goto err_free_dev; + } + + /* Step 1: Parse and validate configuration (DT + module params) */ + ret =3D cmh_config_init(cfg, pdev); + if (ret) + goto err_dma_init; + + /* + * Enable functional clocks and release reset. Both are optional -- + * integrations where a separate management/power controller owns the + * clock and reset lines describe neither, and these calls are + * no-ops. The hub gates its clocks internally, but + * clk_disable_unused() would otherwise gate an always-on input the + * driver never claimed, so the driver enables whatever clocks the + * device tree provides. Clocks come up before reset is released (a + * hard reset requires an active clock) and before any SIC register + * access. The reset line is acquired already deasserted; the driver + * does not drive a reset pulse -- the eSW boots independently and its + * mission-mode readiness is verified separately below, and no in-tree + * platform wires this line for a reset sequence to be exercised. + * devm unwinds both on remove or probe error. + */ + ret =3D devm_clk_bulk_get_all_enabled(&pdev->dev, &clks); + if (ret < 0) { + ret =3D dev_err_probe(&pdev->dev, ret, + "failed to enable clocks\n"); + goto err_dma_init; + } + + reset =3D devm_gpiod_get_optional(&pdev->dev, "reset", GPIOD_OUT_LOW); + if (IS_ERR(reset)) { + ret =3D dev_err_probe(&pdev->dev, PTR_ERR(reset), + "failed to acquire reset GPIO\n"); + goto err_dma_init; + } + + /* Step 2: ioremap the SIC region */ + cfg->sic_mapped =3D devm_platform_ioremap_resource(pdev, 0); + if (IS_ERR(cfg->sic_mapped)) { + ret =3D dev_err_probe(&pdev->dev, PTR_ERR(cfg->sic_mapped), + "ioremap failed for SIC region\n"); + cfg->sic_mapped =3D NULL; + goto err_dma_init; + } + + /* Step 3: Verify CMH is alive */ + if (!skip_fw_check) { + ret =3D cmh_check_sic(cfg); + if (ret) + goto err_dma_init; + } + + /* Step 3.5: Discover crypto cores from the SIC CORE_ENABLE register */ + ret =3D cmh_config_discover_cores(cfg); + if (ret) + goto err_dma_init; + + /* Step 4: Compute per-instance register bases */ + for (i =3D 0; i < cfg->mbx_count; i++) { + struct cmh_mbx_config *m =3D &cfg->mailboxes[i]; + + m->reg_base =3D cmh_mbx_instance_base(cfg->sic_mapped, + m->instance); + + dev_dbg(cmh_dev(), "mbx[%u] instance=3D%u reg_base=3D%p\n", + i, m->instance, m->reg_base); + } + + cmh_debugfs_init(cfg); + + /* Initialise mailbox queue interface */ + ret =3D cmh_mqi_init(cfg); + if (ret) + goto err_mqi_init; + + /* Initialise transaction manager */ + ret =3D cmh_tm_init(cfg); + if (ret) + goto err_tm_init; + + /* Initialise response handler */ + ret =3D cmh_rh_init(cfg); + if (ret) + goto err_rh_init; + + platform_set_drvdata(pdev, dev); + + return 0; + +err_rh_init: + cmh_tm_cleanup(); +err_tm_init: + cmh_mqi_cleanup(cfg); +err_mqi_init: + cmh_debugfs_cleanup(); +err_dma_init: + cmh_dma_cleanup(); +err_free_dev: + return ret; +} + +/* Module Exit -- platform driver remove */ + +static void cmh_remove(struct platform_device *pdev) +{ + struct cmh_device *dev =3D platform_get_drvdata(pdev); + struct cmh_config *cfg; + + if (!dev) + return; + + cfg =3D &dev->config; + + cmh_rh_cleanup(cfg); + cmh_tm_cleanup(); + cmh_mqi_cleanup(cfg); + cmh_debugfs_cleanup(); + cmh_dma_cleanup(); +} + +static const struct of_device_id cmh_of_match[] =3D { + { .compatible =3D "rambus,cmh-v1030" }, + { /* sentinel */ } +}; +MODULE_DEVICE_TABLE(of, cmh_of_match); + +/* + * PM suspend/resume. + * + * Suspend: drain the TM first (while the RH is still active and can + * deliver completions for in-flight transactions), then quiesce the + * RH (cancel watchdog, mask HW interrupts). This ordering ensures + * the drain_timeout_ms wait in cmh_tm_quiesce() can actually succeed + * -- if we suspended RH first, no completions would be delivered and + * the drain would always hit the force-cancel path. + * + * IRQ handlers remain registered (standard PM pattern: the kernel + * disables the IRQ lines during suspend, no need to free/re-request). + * + * Resume: re-check the SIC/SW boot status, re-synchronise the RH + * with hardware (head positions, interrupt masks, watchdog), then + * restart the TM kthread. + */ + +static int cmh_suspend(struct device *dev) +{ + struct cmh_device *cmh =3D dev_get_drvdata(dev); + + if (!cmh) + return 0; + + cmh_tm_quiesce(); + cmh_rh_suspend(&cmh->config); + return 0; +} + +static int cmh_resume(struct device *dev) +{ + struct cmh_device *cmh =3D dev_get_drvdata(dev); + int ret; + + if (!cmh) + return 0; + + ret =3D cmh_check_sic(&cmh->config); + if (ret) { + dev_err(dev, "resume: CMH eSW health check failed (%d)\n", + ret); + return ret; + } + + /* + * cmh_rh_resume() is void: it only re-syncs MMIO head pointers, + * clears stale interrupt status bits (W1C), re-enables interrupt + * masks, and re-arms the watchdog timer -- none of which can fail + * after the SIC health check above has confirmed HW accessibility. + */ + cmh_rh_resume(&cmh->config); + + ret =3D cmh_tm_resume(); + if (ret) { + dev_err(dev, "resume: TM restart failed (%d)\n", ret); + return ret; + } + return 0; +} + +static DEFINE_SIMPLE_DEV_PM_OPS(cmh_pm_ops, + cmh_suspend, + cmh_resume); + +/* + * Runtime PM is intentionally not implemented. The CMH hardware does + * not expose HLOS-accessible clock gates or power domains -- the eSW + * firmware manages HW power state independently. There is no mechanism + * for the kernel to idle, gate clocks, or power down the accelerator + * block from HLOS. If a future platform variant exposes power control + * to HLOS (e.g. via a SCMI power domain), runtime PM support can be + * added at that time using SET_RUNTIME_PM_OPS and pm_runtime_get/put + * around VCQ submission paths. + * + * System sleep (suspend/resume) is supported via DEFINE_SIMPLE_DEV_PM_OPS + * above: suspend quiesces the TM and masks IRQs; resume re-verifies + * eSW health (SIC status) and restarts the TM thread. + */ + +static struct platform_driver cmh_driver =3D { + .probe =3D cmh_probe, + .remove =3D cmh_remove, + .driver =3D { + .name =3D "cmh", + .of_match_table =3D cmh_of_match, + .dev_groups =3D cmh_sysfs_groups, + .pm =3D pm_sleep_ptr(&cmh_pm_ops), + }, +}; + +module_platform_driver(cmh_driver); + +MODULE_DESCRIPTION("Rambus CryptoManager Hub (CMH) hardware crypto acceler= ator"); +MODULE_AUTHOR("Alex Ousherovitch "); +MODULE_AUTHOR("Saravanakrishnan Krishnamoorthy "); +MODULE_AUTHOR("Joel Wittenauer "); +MODULE_IMPORT_NS("CRYPTO_INTERNAL"); +MODULE_LICENSE("GPL"); diff --git a/drivers/crypto/cmh/cmh_mqi.c b/drivers/crypto/cmh/cmh_mqi.c new file mode 100644 index 000000000000..cf1845820d92 --- /dev/null +++ b/drivers/crypto/cmh/cmh_mqi.c @@ -0,0 +1,344 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Mailbox Queue Initializer + * + * Responsibilities: + * - Allocate queue buffers for each configured mailbox + * - Execute the MBX lock/setup/enable register sequence + * - Readback-verify all critical register writes + * - Hold lock for MBX lifetime (CMH eSW requires it for host access) + * - Clean up (flush + unlock + free) on exit or error + * + * Register sequence per instance (per CMH MBX hardware specification): + * 1. Read R_MBX_LOCK -> non-zero =3D ownership token acquired + * 2. W1C stale R_MBX_INTERRUPT bits (avoids spurious error cascade) + * 3. Set R_MBX_INTERRUPT_MASK =3D MBX_IRQ_MASK + * 4. Write QUEUE_LO/HI, SLOTS, STRIDE (queue address + geometry) + * 5. Sync TAIL =3D HEAD (CMH eSW owns HEAD; avoids stale-queue parse) + * 6. Readback verify QUEUE_LO/HI/SLOTS/STRIDE + * 7. Write COMMAND =3D MBX_COMMAND_RUN + * 8. Lock stays held -- released only in teardown + */ + +#include +#include +#include +#include +#include +#include +#include + +#include "cmh_mqi.h" +#include "cmh_dma.h" +#include "cmh_registers.h" +#include "cmh_config.h" + +/* Flush polling: eSW clears R_MBX_COMMAND to 0 when flush completes */ +#define MBX_FLUSH_POLL_US 50 +#define MBX_FLUSH_TIMEOUT_US 1000000 /* 1 second */ + +/* MBX Lock / Unlock */ + +/* + * Attempt to acquire the MBX hardware lock. + * Returns the lock token (non-zero) on success, 0 on timeout. + */ +static u32 cmh_mbx_lock(void __iomem *reg_base, u32 instance) +{ + unsigned long deadline =3D jiffies + msecs_to_jiffies(MBX_LOCK_TIMEOUT_MS= ); + u32 lock; + + while (time_before(jiffies, deadline)) { + lock =3D cmh_reg_read32(reg_base, R_MBX_LOCK); + if (lock) { + dev_dbg(cmh_dev(), "mbx %u lock acquired (token=3D0x%08x)\n", + instance, lock); + return lock; + } + /* HW lock may be held by CMH eSW -- back off before retry */ + usleep_range(MBX_LOCK_POLL_MIN_US, MBX_LOCK_POLL_MAX_US); + } + + return 0; +} + +/* Release the MBX lock: clear interrupt mask, write token back */ +static void cmh_mbx_unlock(void __iomem *reg_base, u32 lock_val) +{ + cmh_reg_write32(0, reg_base, R_MBX_INTERRUPT_MASK); + cmh_reg_write32(lock_val, reg_base, R_MBX_LOCK); +} + +/* Register Readback Verification */ + +static int cmh_verify_reg(void __iomem *base, u32 offset, u32 expected, + const char *name, u32 instance) +{ + u32 actual =3D cmh_reg_read32(base, offset); + + if (actual !=3D expected) { + dev_err(cmh_dev(), "mbx %u %s readback mismatch: 0x%08x !=3D 0x%08x\n", + instance, name, actual, expected); + return -EIO; + } + return 0; +} + +/* Clear any stale interrupt bits left from a prior module lifecycle. */ +static void cmh_mbx_clear_stale_irqs(void __iomem *base, u32 instance) +{ + u32 stale =3D cmh_reg_read32(base, R_MBX_INTERRUPT); + + if (stale) { + cmh_reg_write32(stale, base, R_MBX_INTERRUPT); + dev_dbg(cmh_dev(), "mbx %u cleared stale irq bits=3D0x%x\n", + instance, stale); + } +} + +/* Read CMH eSW HEAD and set TAIL =3D HEAD so the queue appears empty. */ +static void cmh_mbx_sync_tail_to_head(void __iomem *base, u32 instance) +{ + u32 fw_head =3D cmh_reg_read32(base, R_MBX_QUEUE_HEAD); + + cmh_reg_write32(fw_head, base, R_MBX_QUEUE_TAIL); + if (fw_head) + dev_dbg(cmh_dev(), "mbx %u synced tail=3D%u to fw head\n", + instance, fw_head); +} + +/* Per-Mailbox Setup */ + +static int cmh_mbx_setup_one(struct cmh_mbx_config *mbx) +{ + void __iomem *base =3D mbx->reg_base; + u32 addr_lo =3D lower_32_bits(mbx->dma_handle); + u32 addr_hi =3D upper_32_bits(mbx->dma_handle); + u32 lock_val; + int ret; + + /* Step 1: Acquire exclusive access */ + lock_val =3D cmh_mbx_lock(base, mbx->instance); + if (!lock_val) { + dev_err(cmh_dev(), "mbx %u lock timeout after %u ms\n", + mbx->instance, MBX_LOCK_TIMEOUT_MS); + return -ETIMEDOUT; + } + + /* + * Step 1.5: Clear stale interrupt bits from a prior module lifecycle. + * + * After rmmod, the CMH eSW may leave ERROR_IRQ set in + * R_MBX_INTERRUPT even though STATUS is IDLE. If we enable + * the mask first, the stale bits immediately trigger the + * CMH eSW interrupt chain, which can cascade into ERROR + * status before the first hash operation. W1C-clear them first. + */ + cmh_mbx_clear_stale_irqs(base, mbx->instance); + + /* Step 2: Program interrupt mask (enable DONE/ERROR interrupts) */ + cmh_reg_write32(MBX_IRQ_MASK, base, R_MBX_INTERRUPT_MASK); + + /* Step 3: Configure queue address (64-bit split) */ + cmh_reg_write32(addr_lo, base, R_MBX_QUEUE_LO); + cmh_reg_write32(addr_hi, base, R_MBX_QUEUE_HI); + + /* Step 4: Configure queue geometry */ + cmh_reg_write32(mbx->slots_log2, base, R_MBX_QUEUE_SLOTS); + cmh_reg_write32(mbx->stride_log2, base, R_MBX_QUEUE_STRIDE); + + /* + * Step 5: Synchronise TAIL to CMH eSW's HEAD. + * + * R_MBX_QUEUE_HEAD is read-only from the host side -- only the + * CMH eSW can write it. On a fresh boot HEAD is 0; after an + * rmmod/insmod cycle it retains the value from the previous + * session (e.g. 44). Writing 0 from the host is silently + * dropped by the MBX HW. + * + * If we set TAIL=3D0 while HEAD=3D44 the CMH eSW sees a non-empty + * queue (head !=3D tail with wrap-around) and immediately tries + * to load a VCQ at the old head offset into our freshly-zeroed + * DMA buffer, causing an "Invalid VCQ" EFAULT -> ECHILD cascade. + * + * Fix: read HEAD and set TAIL =3D HEAD so the queue looks empty. + */ + cmh_mbx_sync_tail_to_head(base, mbx->instance); + + /* + * Step 6: Readback verify critical registers. + * HOST_INFO is deliberately deferred to after verification -- writing + * it tells the CMH eSW "MBX is ready" and the CMH eSW may inspect + * (and clear) the queue registers immediately. + */ + ret =3D cmh_verify_reg(base, R_MBX_QUEUE_LO, addr_lo, + "QUEUE_LO", mbx->instance); + if (ret) + goto err_unlock; + + ret =3D cmh_verify_reg(base, R_MBX_QUEUE_HI, addr_hi, + "QUEUE_HI", mbx->instance); + if (ret) + goto err_unlock; + + ret =3D cmh_verify_reg(base, R_MBX_QUEUE_SLOTS, mbx->slots_log2, + "QUEUE_SLOTS", mbx->instance); + if (ret) + goto err_unlock; + + ret =3D cmh_verify_reg(base, R_MBX_QUEUE_STRIDE, mbx->stride_log2, + "QUEUE_STRIDE", mbx->instance); + if (ret) + goto err_unlock; + + /* Step 7: Enable -- start the mailbox */ + cmh_reg_write32(MBX_COMMAND_RUN, base, R_MBX_COMMAND); + + /* Read status while we still hold the lock */ + dev_dbg(cmh_dev(), "mbx %u setup: dma=3D0x%08x%08x slots=3D%u stride=3D%u= status=3D0x%08x\n", + mbx->instance, addr_hi, addr_lo, + mbx->slots_log2, mbx->stride_log2, + cmh_reg_read32(base, R_MBX_STATUS)); + + /* + * Lock stays held for the lifetime of this MBX session. + * + * mbx->lock_val is the ownership token returned by R_MBX_LOCK at + * acquisition time. The CMH eSW validates this token on every + * register access and requires it to be written back to release. + * It is NOT a transient mutex -- it persists until teardown. + */ + mbx->lock_val =3D lock_val; + + return 0; + +err_unlock: + cmh_mbx_unlock(base, lock_val); + return ret; +} + +/* Per-Mailbox Teardown */ + +static void cmh_mbx_teardown_one(struct cmh_mbx_config *mbx) +{ + void __iomem *base =3D mbx->reg_base; + u32 status; + + if (!base || !mbx->lock_val) + return; + + if (MBX_STATUS_CODE(cmh_reg_read32(base, R_MBX_STATUS)) !=3D + MBX_STATUS_OFFLINE) { + cmh_reg_write32(MBX_COMMAND_FLUSH, base, R_MBX_COMMAND); + + /* + * Wait for the eSW to process the flush before releasing + * the DMA buffer. The eSW clears R_MBX_COMMAND to zero + * upon completion; if it doesn't within 1 s, log a + * warning and proceed (best-effort teardown). + * + * DMA safety: by this point the RH and TM are already + * shut down (remove order: algos -> RH -> TM -> MQI), + * so no new transactions can be submitted and no + * completions are in flight. The queue buffer is only + * read by the eSW during active command processing; + * after flush the eSW will not touch it again. + */ + if (read_poll_timeout(cmh_reg_read32, status, + status =3D=3D 0, + MBX_FLUSH_POLL_US, + MBX_FLUSH_TIMEOUT_US, + true, base, R_MBX_COMMAND)) + dev_warn(cmh_dev(), + "mbx %u flush timeout during teardown (status=3D0x%08x)\n", + mbx->instance, + cmh_reg_read32(base, R_MBX_STATUS)); + } + + cmh_mbx_unlock(base, mbx->lock_val); + mbx->lock_val =3D 0; +} + +/* Public Interface */ + +/** + * cmh_mqi_init() - Initialize all mailbox queues + * @cfg: CMH configuration describing the mailboxes to set up + * + * Allocates DMA queue buffers for each configured mailbox, then executes + * the MBX lock/setup/enable register sequence. On failure, all + * successfully initialized mailboxes are torn down and buffers freed. + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mqi_init(struct cmh_config *cfg) +{ + unsigned int i, j; + int ret; + + /* Allocate queue buffers */ + for (i =3D 0; i < cfg->mbx_count; i++) { + struct cmh_mbx_config *m =3D &cfg->mailboxes[i]; + + m->virt_addr =3D cmh_dma_alloc(m->queue_size, &m->dma_handle, + GFP_KERNEL); + if (!m->virt_addr) { + ret =3D -ENOMEM; + goto err_free_bufs; + } + + dev_dbg(cmh_dev(), "mqi[%u] alloc %zu bytes @ virt=3D%pK dma=3D%pad\n", + i, m->queue_size, m->virt_addr, &m->dma_handle); + } + + /* Lock/setup/enable each mailbox */ + for (i =3D 0; i < cfg->mbx_count; i++) { + ret =3D cmh_mbx_setup_one(&cfg->mailboxes[i]); + if (ret) { + dev_err(cmh_dev(), "mqi[%u] setup failed (rc=3D%d)\n", + i, ret); + goto err_teardown; + } + } + + return 0; + +err_teardown: + for (j =3D 0; j < i; j++) + cmh_mbx_teardown_one(&cfg->mailboxes[j]); +err_free_bufs: + for (j =3D 0; j < cfg->mbx_count; j++) { + if (cfg->mailboxes[j].virt_addr) + cmh_dma_free(cfg->mailboxes[j].queue_size, + cfg->mailboxes[j].virt_addr, + cfg->mailboxes[j].dma_handle); + cfg->mailboxes[j].virt_addr =3D NULL; + cfg->mailboxes[j].dma_handle =3D 0; + } + return ret; +} + +/** + * cmh_mqi_cleanup() - Clean up all mailbox queues + * @cfg: CMH configuration describing the mailboxes to tear down + * + * Tears down each mailbox (flush + unlock) and frees the DMA queue + * buffers allocated by cmh_mqi_init(). + */ +void cmh_mqi_cleanup(struct cmh_config *cfg) +{ + unsigned int i; + + for (i =3D 0; i < cfg->mbx_count; i++) { + struct cmh_mbx_config *m =3D &cfg->mailboxes[i]; + + cmh_mbx_teardown_one(m); + + if (m->virt_addr) + cmh_dma_free(m->queue_size, m->virt_addr, + m->dma_handle); + m->virt_addr =3D NULL; + m->dma_handle =3D 0; + } +} diff --git a/drivers/crypto/cmh/cmh_rh.c b/drivers/crypto/cmh/cmh_rh.c new file mode 100644 index 000000000000..62b6f71f741c --- /dev/null +++ b/drivers/crypto/cmh/cmh_rh.c @@ -0,0 +1,1178 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Response Handler + * + * IRQ-driven completion processing using request_threaded_irq(): + * + * Hardirq: For each MBX, read R_MBX_INTERRUPT. If any bit is set, + * W1C-clear it and mark the MBX for threaded processing. + * Return IRQ_WAKE_THREAD if any MBX had work. + * + * Thread: For each pending MBX, read R_MBX_QUEUE_HEAD. Walk the + * per-MBX transaction queue (oldest first): for every txn + * whose last_vcq_id < new_head, check status, fire the + * completion callback, and free the transaction object. + * + * The DT "rambus,cmh-v1030" node declares one PLIC interrupt per mailbox, + * matching the real CMH ch_sys_interrupt_mbx[N-1:0] topology. + * Each MBX gets its own Linux virq; the same hardirq/thread pair + * is registered for all of them. The handler still scans all + * mailboxes on every invocation -- this is intentional, as it + * provides robustness against coalesced or missed edges. + * + * IRQ source: resolved from the "rambus,cmh-v1030" DT node at init time. + * The module's irq=3D parameter can override with a single shared IRQ. + */ + +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "cmh_rh.h" +#include "cmh_txn.h" +#include "cmh_registers.h" +#include "cmh_config.h" +#include "cmh_debugfs.h" +#include "cmh_dma.h" + +/* Per-mailbox IRQ bookkeeping */ +struct cmh_rh_mbx { + u32 last_head; /* last-observed MBX head position */ + atomic_t irq_bits; /* interrupt bits saved by hardirq (atomic_or) */ + bool pending; /* threaded handler should process this MBX */ + bool restart_pending; /* RESTART issued, awaiting eSW ack */ + u32 restart_retries; /* watchdog ticks since RESTART issued */ + u32 flush_count; /* consecutive failed FLUSH escalations */ + bool wedged; /* recovery failed, MBX offline */ + u32 abort_stall_ticks; /* ticks since async timeout ABORT issued */ +}; + +/* Module-level RH state */ +static struct { + struct cmh_config *cfg; + int irqs[CMH_MAX_CONFIGURED_MBX]; /* per-MBX virqs */ + u32 nirqs; /* number of registered IRQs */ + struct cmh_rh_mbx *mbx; /* array[cfg->mbx_count] */ + atomic_t irq_count; /* hardirq invocation counter */ + bool active; +} rh; + +/* + * Serialise the read-last_head / process_mbx / update-last_head + * sequence between the threaded IRQ handler (process context) and + * the watchdog timer (softirq context). Without this, a timer + * softirq can preempt the kthread mid-sequence, causing both paths + * to process the same head advance and prematurely complete a + * subsequent transaction before the CMH eSW has written its DMA + * output -- leading to data corruption and SLAB freelist poisoning. + * + * The kthread acquires with spin_lock_bh (disables softirqs), the + * watchdog acquires with spin_lock (already in softirq context). + */ +static DEFINE_SPINLOCK(rh_process_lock); + +/* + * Watchdog timer -- missed-IRQ recovery. + * + * Fires every watchdog_ms while rh.active. Reads MBX head registers; + * if any head has advanced without an IRQ, processes completions and + * logs a notice. Standard kernel pattern, analogous to NIC watchdog + * timers. + * + * Safe from timer/softirq context: cmh_reg_read32() is an MMIO read, + * cmh_tm_pop_transaction() uses spin_lock_irqsave(), and TM completion + * callbacks (crypto_request_complete et al.) are documented safe from + * any context including softirq. rh_process_lock serialises the + * head-read / process / head-update sequence against the threaded + * IRQ handler to prevent double-processing of the same completion. + * + * Default 200 ms (5 fires/s) provides ~10 recovery attempts within + * the default vcq_timeout_ms (2 s). Tune via debugfs config/watchdog_ms + * for platforms where interrupt delivery is more reliable (e.g. MSI on + * FPGA/silicon -- 500 ms--1 s may suffice as a safety net). + */ +#define CMH_RH_WATCHDOG_MS_DEFAULT 200 + +/* + * Floor for watchdog_ms to prevent a zero/near-zero value from + * spinning the timer in a tight softirq loop. Enforced at the + * point of use so debugfs writes are never rejected. + */ +#define CMH_RH_WATCHDOG_MS_MIN 10 + +/* + * Maximum watchdog ticks to wait for the eSW to process RESTART + * before escalating to FLUSH. At the default 200 ms interval, + * 5 retries =3D 1 s -- generous for an operation that should take + * microseconds. If the eSW hasn't responded by then, issue + * MBX_COMMAND_FLUSH to hard-reset the mailbox state. + */ +#define CMH_RH_RESTART_MAX_RETRIES 5 + +/* + * Maximum consecutive FLUSH escalations before marking the MBX as + * wedged. Each FLUSH cycle takes RESTART_MAX_RETRIES watchdog ticks + * (~1 s at default interval). Two failed FLUSHes (~2 s total) + * strongly indicate the eSW is not processing MBX commands at all. + */ +#define CMH_RH_FLUSH_MAX_FAILURES 2 + +/* + * Time budget (ms) after an async timeout ABORT before escalating + * to FLUSH + force-drain. Converted to watchdog ticks at runtime + * via abort_stall_ms / watchdog_ms, so the actual wall-clock bound + * stays constant regardless of watchdog_ms tuning. + * + * The stall detector fires when: + * - The head-of-queue transaction is in TXN_TIMED_OUT state + * - HEAD hasn't advanced (eSW didn't process the ABORT) + * - abort_stall_ticks exceeds the derived threshold + * + * At that point we issue FLUSH + force-drain, completing all pending + * transactions with -ETIMEDOUT and waking any blocked waiters. + * + * Default 5000 ms bounds worst-case D-state to + * async_timeout (2 s) + abort_stall (5 s) =3D ~7 s. + */ +#define CMH_RH_ABORT_STALL_MS 5000 + +static unsigned int watchdog_ms =3D CMH_RH_WATCHDOG_MS_DEFAULT; + +/* + * Re-poke R_MBX_QUEUE_TAIL to generate a fresh interrupt to the eSW. + * Writing the current value back is a queue no-op but guarantees a + * SIC interrupt edge, ensuring the eSW wakes from WFI. + */ +static void cmh_rh_poke_tail(void __iomem *base) +{ + cmh_tm_poke_tail(base); +} + +/* + * Drain all remaining in-flight transactions for a mailbox, completing + * each with the given error code. Called after FLUSH (which discards + * all queued VCQs) or when marking a mailbox as wedged. Updates + * last_head to the current hardware HEAD so subsequent polls don't + * re-process the same (now-dead) VCQ IDs as successful completions. + * + * Caller must hold rh_process_lock. + */ +static void cmh_rh_drain_mbx(u32 mbx_idx, int error) +{ + struct transaction_obj *txn; + + while ((txn =3D cmh_tm_pop_transaction(mbx_idx)) !=3D NULL) { + dev_dbg(cmh_dev(), "rh: mbx[%u] drain vcq=3D%u..%u err=3D%d\n", + mbx_idx, txn->first_vcq_id, + txn->last_vcq_id, error); + cmh_txn_finish(txn, error); + cmh_tm_txq_completion_notify(); + } + + rh.mbx[mbx_idx].last_head =3D + cmh_reg_read32(rh.cfg->mailboxes[mbx_idx].reg_base, + R_MBX_QUEUE_HEAD); +} + +/** + * cmh_rh_force_drain_mbx() - FLUSH + drain a mailbox from external context + * @mbx_idx: Mailbox index to drain + * + * Issues MBX_COMMAND_FLUSH to the eSW, drains all pending transactions + * (completing each with -ECANCELED), and resets all recovery bookkeeping + * including the wedged flag. This is an administrative last-resort + * recovery path exposed via debugfs. + * + * Context: process context. Acquires rh_process_lock internally. + */ +void cmh_rh_force_drain_mbx(u32 mbx_idx) +{ + void __iomem *base; + + if (!rh.cfg || !rh.mbx || mbx_idx >=3D rh.cfg->mbx_count) + return; + + base =3D rh.cfg->mailboxes[mbx_idx].reg_base; + + dev_warn(cmh_dev(), "rh: force-drain mbx[%u] (debugfs)\n", mbx_idx); + spin_lock_bh(&rh_process_lock); + cmh_reg_write32(MBX_IRQ_MASK, base, R_MBX_INTERRUPT); + cmh_reg_write32(MBX_COMMAND_FLUSH, base, R_MBX_COMMAND); + cmh_rh_poke_tail(base); + cmh_rh_drain_mbx(mbx_idx, -ECANCELED); + rh.mbx[mbx_idx].abort_stall_ticks =3D 0; + WRITE_ONCE(rh.mbx[mbx_idx].restart_pending, false); + rh.mbx[mbx_idx].restart_retries =3D 0; + rh.mbx[mbx_idx].flush_count =3D 0; + WRITE_ONCE(rh.mbx[mbx_idx].wedged, false); + spin_unlock_bh(&rh_process_lock); +} + +/** + * cmh_rh_mbx_is_wedged() - Check if a mailbox is permanently wedged + * @mbx_idx: Mailbox index to check + * + * Return: true if the mailbox has failed recovery and is offline. + */ +bool cmh_rh_mbx_is_wedged(u32 mbx_idx) +{ + if (!rh.mbx || !rh.cfg || mbx_idx >=3D rh.cfg->mbx_count) + return false; + + return READ_ONCE(rh.mbx[mbx_idx].wedged); +} + +/** + * cmh_rh_abort_mbx() - Issue MBX_COMMAND_ABORT under rh_process_lock + * @mbx_idx: Mailbox index to abort + * + * Serialises the ABORT write with RESTART/FLUSH commands issued by the + * watchdog, preventing command-register clobber races. Safe to call + * from any context (uses spin_lock_bh). + */ +void cmh_rh_abort_mbx(u32 mbx_idx) +{ + void __iomem *base; + + if (!rh.cfg || !rh.mbx || mbx_idx >=3D rh.cfg->mbx_count) + return; + + base =3D rh.cfg->mailboxes[mbx_idx].reg_base; + + spin_lock_bh(&rh_process_lock); + cmh_reg_write32(MBX_COMMAND_ABORT, base, R_MBX_COMMAND); + spin_unlock_bh(&rh_process_lock); +} + +static struct timer_list rh_watchdog; + +/* + * Hardirq handler -- runs with interrupts disabled. + * + * Read and W1C-clear R_MBX_INTERRUPT for each mailbox. + * If any MBX had a pending interrupt, return IRQ_WAKE_THREAD. + * Shared-IRQ safe: returns IRQ_NONE if we didn't handle anything. + */ +static irqreturn_t cmh_rh_hardirq(int irq, void *data) +{ + struct cmh_config *cfg =3D data; + bool handled =3D false; + u32 i; + + for (i =3D 0; i < cfg->mbx_count; i++) { + void __iomem *base =3D cfg->mailboxes[i].reg_base; + u32 bits; + + bits =3D cmh_reg_read32(base, R_MBX_INTERRUPT); + if (!bits) + continue; + + /* W1C: write back the set bits to clear them */ + cmh_reg_write32(bits, base, R_MBX_INTERRUPT); + + /* + * Accumulate bits atomically so a second hardirq + * firing while the threaded handler runs does not + * overwrite the first set of bits. + */ + atomic_or((int)bits, &rh.mbx[i].irq_bits); + WRITE_ONCE(rh.mbx[i].pending, true); + handled =3D true; + } + + /* + * Ordering: the kernel IRQ threading infrastructure + * performs a full barrier between hardirq return and + * the threaded handler invocation. + */ + if (handled) + atomic_inc(&rh.irq_count); + + return handled ? IRQ_WAKE_THREAD : IRQ_NONE; +} + +static void cmh_rh_stat_inc_errors(u32 mbx_idx) +{ + struct cmh_mbx_stats *s =3D cmh_debugfs_mbx_stats(mbx_idx); + + if (s) + atomic64_inc(&s->vcqs_errors); +} + +static void cmh_rh_stat_add_completed(u32 mbx_idx, u32 first_vcq, u32 last= _vcq) +{ + struct cmh_mbx_stats *s =3D cmh_debugfs_mbx_stats(mbx_idx); + + if (s) + atomic64_add(last_vcq - first_vcq + 1, &s->vcqs_completed); +} + +/* + * Process completions for a single mailbox. + * + * Walk the per-MBX transaction queue FIFO. For each transaction + * whose last_vcq_id is strictly less than the new head, fire the + * completion callback and free the object. + * + * "Strictly less than" using signed (s32) arithmetic handles wrap-around: + * the CMH eSW uses monotonically increasing 32-bit VCQ IDs. + */ +static void cmh_rh_process_mbx(u32 mbx_idx, u32 new_head, u32 irq_bits) +{ + struct transaction_obj *txn; + int error =3D 0; + + /* Determine error state from saved IRQ bits */ + if (irq_bits & MBX_ERROR_IRQ) { + void __iomem *base =3D rh.cfg->mailboxes[mbx_idx].reg_base; + u32 status =3D cmh_reg_read32(base, R_MBX_STATUS); + + error =3D -EIO; + dev_dbg(cmh_dev(), "rh: mbx[%u] error status=3D0x%08x (code=3D%u cmd_idx= =3D%u)\n", + mbx_idx, status, + MBX_STATUS_ERROR_CODE(status), + MBX_STATUS_CMD_INDEX(status)); + + /* + * ECHILD (10) in the parent status means a child VCQ + * failed internally. Read R_MBX_CHILD for the actual + * root cause (real errno, child core ID, child cmd idx). + */ + if (MBX_STATUS_ERROR_CODE(status) =3D=3D ECHILD) { + u32 child =3D cmh_reg_read32(base, R_MBX_CHILD); + + dev_dbg(cmh_dev(), + "rh: mbx[%u] child error=3D0x%08x (core=3D%u code=3D%u cmd_idx=3D%u)\n= ", + mbx_idx, child, + MBX_STATUS_CORE_ID(child), + MBX_STATUS_ERROR_CODE(child), + MBX_STATUS_CMD_INDEX(child)); + } + + /* + * CMH eSW does not advance head on error -- the MBX is + * stuck in ERROR state until the host issues a recovery + * command. However, HEAD may have advanced past one or + * more already-completed transactions before the error + * occurred (their completion IRQ may not have been + * processed yet). Retire those normally first, then + * force-complete the NEXT transaction (the one that + * actually failed) with -EIO. + * + * MBX command semantics after ERROR: + * CONTINUE -- re-run the same VCQ at HEAD (retry) + * RESTART -- advance HEAD+1, skip failed, resume + * FLUSH -- HEAD=3DTAIL, flush all HWCs, discard queue + */ + + /* First: retire transactions completed before the error */ + while ((txn =3D cmh_tm_peek_transaction(mbx_idx)) !=3D NULL) { + if ((s32)(new_head - txn->last_vcq_id) <=3D 0) + break; + txn =3D cmh_tm_pop_transaction(mbx_idx); + if (!txn) + break; + dev_dbg(cmh_dev(), + "rh: mbx[%u] pre-error complete vcq=3D%u..%u\n", + mbx_idx, txn->first_vcq_id, + txn->last_vcq_id); + cmh_txn_finish(txn, 0); + cmh_tm_txq_completion_notify(); + } + + /* + * The transaction that errored is now at the head. If it is + * multi-VCQ (first_vcq_id !=3D last_vcq_id) it needs special + * care: the eSW has no notion of a VCQ group and RESTART only + * advances HEAD by one, so it would resume the surviving slots + * of THIS transaction -- whose DMA cmh_txn_finish() is about + * to free -- a use-after-free. FLUSH discards the whole queue + * (HEAD=3DTAIL) so no surviving slot runs; drain then completes + * this and any following transactions with the error. + */ + txn =3D cmh_tm_peek_transaction(mbx_idx); + if (txn && txn->first_vcq_id !=3D txn->last_vcq_id) { + dev_warn_ratelimited(cmh_dev(), + "rh: mbx[%u] multi-VCQ error vcq=3D%u..%u -- FLUSH+drain\n", + mbx_idx, txn->first_vcq_id, + txn->last_vcq_id); + cmh_rh_stat_inc_errors(mbx_idx); + cmh_reg_write32(MBX_IRQ_MASK, base, R_MBX_INTERRUPT); + cmh_reg_write32(MBX_COMMAND_FLUSH, base, R_MBX_COMMAND); + cmh_rh_poke_tail(base); + cmh_rh_drain_mbx(mbx_idx, error); + return; + } + + /* Now pop and fail the transaction that actually errored */ + txn =3D cmh_tm_pop_transaction(mbx_idx); + if (txn) { + dev_dbg(cmh_dev(), "rh: mbx[%u] error-complete vcq=3D%u..%u\n", + mbx_idx, txn->first_vcq_id, + txn->last_vcq_id); + cmh_txn_finish(txn, error); + cmh_tm_txq_completion_notify(); + } else { + u32 head_reg, tail_reg; + + head_reg =3D cmh_reg_read32(base, R_MBX_QUEUE_HEAD); + tail_reg =3D cmh_reg_read32(base, R_MBX_QUEUE_TAIL); + dev_warn_ratelimited(cmh_dev(), + "rh: mbx[%u] ERROR with empty txn queue (orphaned) status=3D0x%0= 8x head=3D%u tail=3D%u core=3D%u ecode=3D%u cmd_idx=3D%u\n", + mbx_idx, status, + head_reg, tail_reg, + MBX_STATUS_CORE_ID(status), + MBX_STATUS_ERROR_CODE(status), + MBX_STATUS_CMD_INDEX(status)); + } + cmh_rh_stat_inc_errors(mbx_idx); + + /* + * W1C-clear R_MBX_INTERRUPT before issuing RESTART. + * + * The eSW sets MBX_ERROR_IRQ in R_MBX_INTERRUPT when + * it writes ERROR status. On platforms where the + * hardirq handler runs (IRQ wired to GIC), this bit + * is cleared there. On polling-only platforms (no + * IRQ line), it must be cleared explicitly before + * issuing a recovery command to de-assert the + * MBX-to-SIC interrupt line. + */ + cmh_reg_write32(MBX_IRQ_MASK, base, R_MBX_INTERRUPT); + cmh_reg_write32(MBX_COMMAND_RESTART, base, R_MBX_COMMAND); + + /* + * Poke R_MBX_QUEUE_TAIL to guarantee the eSW receives + * an interrupt. + * + * Writing R_MBX_COMMAND alone may not produce a new + * SIC interrupt edge if the MBX-to-SIC line is still + * asserted from prior error processing. The eSW RUN + * handler re-writes ERROR_IRQ to R_MBX_INTERRUPT on + * every spurious wakeup while in ERROR state, which + * can keep the SIC line high on level-triggered HW. + * + * R_MBX_QUEUE_TAIL writes always generate a fresh + * interrupt to the eSW (this is the normal VCQ + * submission path). Writing the current TAIL value + * back is a no-op from the queue perspective but + * ensures the eSW wakes from WFI and processes the + * RESTART command. + */ + cmh_rh_poke_tail(base); + WRITE_ONCE(rh.mbx[mbx_idx].restart_pending, true); + rh.mbx[mbx_idx].restart_retries =3D 0; + return; + } + + /* + * Pop completed transactions. A transaction is complete when + * the CMH eSW has advanced head past its last VCQ ID: + * (s32)(new_head - txn->last_vcq_id) > 0 + * Using signed comparison for correct wrap-around handling. + * + * Multi-VCQ note: a transaction can span multiple parent VCQs + * (e.g. SLH-DSA). The CMH eSW advances HEAD one VCQ at a time + * and raises a completion interrupt per VCQ, so for a multi-VCQ + * transaction HEAD can legitimately be observed partway through + * the group while the remaining VCQs are still running. The + * transaction is treated atomically here: it is completed only + * once HEAD has advanced past its last_vcq_id. An intermediate + * HEAD position is expected and simply causes us to wait. + */ + while ((txn =3D cmh_tm_peek_transaction(mbx_idx)) !=3D NULL) { + if ((s32)(new_head - txn->last_vcq_id) <=3D 0) { + /* + * Not yet complete. An intermediate HEAD within a + * multi-VCQ group is normal (the eSW advances HEAD + * per VCQ), so log at debug level and wait for the + * group to finish. + */ + if (txn->first_vcq_id !=3D txn->last_vcq_id && + (s32)(new_head - txn->first_vcq_id) > 0) + dev_dbg_ratelimited(cmh_dev(), + "rh: mbx[%u] head %u mid-group %u..%u\n", + mbx_idx, new_head, + txn->first_vcq_id, + txn->last_vcq_id); + break; + } + + txn =3D cmh_tm_pop_transaction(mbx_idx); + if (!txn) + break; + + dev_dbg(cmh_dev(), "rh: mbx[%u] complete vcq=3D%u..%u err=3D%d\n", + mbx_idx, txn->first_vcq_id, txn->last_vcq_id, + error); + + cmh_rh_stat_add_completed(mbx_idx, txn->first_vcq_id, + txn->last_vcq_id); + + cmh_txn_finish(txn, error); + cmh_tm_txq_completion_notify(); + } +} + +/* + * Threaded IRQ handler -- runs in process context. + * + * Walk all MBXes that had pending interrupts. After processing the + * pending set, do a final hardware poll of all MBX head registers to + * catch completions whose PLIC interrupt was consumed during an + * earlier register access (e.g. an inline interrupt notification + * during MMIO can cause the PLIC edge to be claimed before the + * hardirq sees it). + */ +static irqreturn_t cmh_rh_thread(int irq, void *data) +{ + struct cmh_config *cfg =3D data; + u32 i; + bool recheck; + + do { + recheck =3D false; + + for (i =3D 0; i < cfg->mbx_count; i++) { + u32 new_head, irq_bits; + + if (!READ_ONCE(rh.mbx[i].pending)) + continue; + + irq_bits =3D (u32)atomic_xchg(&rh.mbx[i].irq_bits, 0); + WRITE_ONCE(rh.mbx[i].pending, false); + + spin_lock_bh(&rh_process_lock); + new_head =3D cmh_reg_read32(cfg->mailboxes[i].reg_base, + R_MBX_QUEUE_HEAD); + + if (new_head =3D=3D rh.mbx[i].last_head && !irq_bits) { + spin_unlock_bh(&rh_process_lock); + continue; + } + + cmh_rh_process_mbx(i, new_head, irq_bits); + rh.mbx[i].last_head =3D new_head; + spin_unlock_bh(&rh_process_lock); + } + + /* + * Re-check: if the hardirq fired again while we were + * processing, pending flags will be set again. + */ + for (i =3D 0; i < cfg->mbx_count; i++) { + if (READ_ONCE(rh.mbx[i].pending)) { + recheck =3D true; + break; + } + } + } while (recheck); + + /* + * Final hardware poll: read every MBX head register and status + * to catch completions or errors whose interrupt was missed. + */ + for (i =3D 0; i < cfg->mbx_count; i++) { + u32 new_head; + u32 status; + u32 poll_irq_bits =3D 0; + + spin_lock_bh(&rh_process_lock); + new_head =3D cmh_reg_read32(cfg->mailboxes[i].reg_base, + R_MBX_QUEUE_HEAD); + status =3D cmh_reg_read32(cfg->mailboxes[i].reg_base, + R_MBX_STATUS); + + if (MBX_STATUS_CODE(status) =3D=3D MBX_STATUS_ERROR) { + if (READ_ONCE(rh.mbx[i].wedged)) { + spin_unlock_bh(&rh_process_lock); + continue; + } + if (READ_ONCE(rh.mbx[i].restart_pending)) { + /* + * HEAD advanced while restart_pending means + * RESTART worked but next VCQ also failed. + * Clear restart state and process new error. + */ + if (new_head !=3D rh.mbx[i].last_head) { + WRITE_ONCE(rh.mbx[i].restart_pending, + false); + rh.mbx[i].restart_retries =3D 0; + } else { + spin_unlock_bh(&rh_process_lock); + continue; + } + } + poll_irq_bits =3D MBX_ERROR_IRQ; + } else { + WRITE_ONCE(rh.mbx[i].restart_pending, false); + rh.mbx[i].restart_retries =3D 0; + rh.mbx[i].flush_count =3D 0; + } + + if (new_head !=3D rh.mbx[i].last_head || poll_irq_bits) { + cmh_rh_process_mbx(i, new_head, poll_irq_bits); + rh.mbx[i].last_head =3D new_head; + } + spin_unlock_bh(&rh_process_lock); + } + + return IRQ_HANDLED; +} + +/* + * Watchdog timer callback -- missed-IRQ recovery. + * + * Reads all MBX head registers. If any head advanced without a + * corresponding IRQ, process the completions here. Re-arms itself + * while rh.active is true. + */ +static void cmh_rh_watchdog_fn(struct timer_list *t) +{ + u32 i; + + if (!rh.active || !rh.cfg || !rh.mbx) + return; + + for (i =3D 0; i < rh.cfg->mbx_count; i++) { + u32 new_head; + u32 status; + u32 irq_bits =3D 0; + + spin_lock(&rh_process_lock); + new_head =3D cmh_reg_read32(rh.cfg->mailboxes[i].reg_base, + R_MBX_QUEUE_HEAD); + status =3D cmh_reg_read32(rh.cfg->mailboxes[i].reg_base, + R_MBX_STATUS); + + if (MBX_STATUS_CODE(status) =3D=3D MBX_STATUS_ERROR) { + if (READ_ONCE(rh.mbx[i].wedged)) { + spin_unlock(&rh_process_lock); + continue; + } + /* + * Back-to-back failure scenario: the crypto API + * (e.g. testmgr) may submit requests continuously. + * If RESTART succeeds but the next VCQ also fails, + * the entire RESTART->IDLE->RUN->ERROR cycle can + * complete within a single 200ms watchdog period. + * Without the HEAD-advance check below, the watchdog + * would mistake the new error for a failed RESTART, + * increment restart_retries, and eventually escalate + * to FLUSH -- wedging the mailbox unnecessarily. + */ + if (READ_ONCE(rh.mbx[i].restart_pending)) { + void __iomem *base =3D + rh.cfg->mailboxes[i].reg_base; + + /* + * HEAD advanced since RESTART was issued: + * RESTART succeeded, this is a fresh error. + * Clear recovery state and process normally. + */ + if (new_head !=3D rh.mbx[i].last_head) { + dev_dbg(cmh_dev(), + "rh: watchdog: mbx[%u] head advanced %u->%u during restart -- new er= ror\n", + i, rh.mbx[i].last_head, + new_head); + WRITE_ONCE(rh.mbx[i].restart_pending, + false); + rh.mbx[i].restart_retries =3D 0; + goto new_error; + } + + rh.mbx[i].restart_retries++; + if (rh.mbx[i].restart_retries > + CMH_RH_RESTART_MAX_RETRIES) { + rh.mbx[i].flush_count++; + if (rh.mbx[i].flush_count >=3D + CMH_RH_FLUSH_MAX_FAILURES) { + u32 hb, ei, cmd; + + cmd =3D cmh_reg_read32(base, R_MBX_COMMAND); + hb =3D cmh_reg_read32(rh.cfg->sic_mapped, + R_SIC_SW_HEARTBEAT); + ei =3D cmh_reg_read32(rh.cfg->sic_mapped, + R_SIC_SW_ERROR_INFO); + dev_crit(cmh_dev(), + "rh: mbx[%u] wedged after %u FLUSHes (cmd=3D0x%x status=3D0x%x hb= =3D0x%x err=3D0x%x)\n", + i, + rh.mbx[i].flush_count, + cmd, status, + hb, ei); + WRITE_ONCE(rh.mbx[i].wedged, + true); + cmh_rh_drain_mbx(i, -EIO); + spin_unlock(&rh_process_lock); + continue; + } + /* + * Backstop: eSW did not respond + * to RESTART within the retry + * budget. Escalate to FLUSH + * which is a harder reset of + * the eSW mailbox state. + */ + dev_err(cmh_dev(), + "rh: watchdog: mbx[%u] RESTART unresponsive after %u ticks, escalati= ng to FLUSH (attempt %u/%u)\n", + i, rh.mbx[i].restart_retries, + rh.mbx[i].flush_count, + CMH_RH_FLUSH_MAX_FAILURES); + cmh_reg_write32(MBX_IRQ_MASK, + base, + R_MBX_INTERRUPT); + cmh_reg_write32(MBX_COMMAND_FLUSH, + base, + R_MBX_COMMAND); + cmh_rh_poke_tail(base); + cmh_rh_drain_mbx(i, -EIO); + WRITE_ONCE(rh.mbx[i].restart_pending, + false); + rh.mbx[i].restart_retries =3D 0; + spin_unlock(&rh_process_lock); + continue; + } + /* + * RESTART was already issued on a prior + * tick but the eSW hasn't cleared the + * ERROR status yet. Do NOT pop another + * transaction -- that would cascade-kill + * unrelated in-flight work. Re-poke TAIL + * in case the eSW missed the interrupt. + */ + cmh_rh_poke_tail(base); + dev_dbg_ratelimited(cmh_dev(), + "rh: watchdog: mbx[%u] restart pending (%u/%u) status=3D0x%08x, = re-poke\n", + i, + rh.mbx[i].restart_retries, + CMH_RH_RESTART_MAX_RETRIES, + status); + spin_unlock(&rh_process_lock); + continue; + } +new_error: + dev_dbg_ratelimited(cmh_dev(), + "rh: watchdog: mbx[%u] error status=3D0x%08x (missed error IRQ) h= ead=3D%u tail=3D%u core=3D%u ecode=3D%u cmd_idx=3D%u\n", + i, status, new_head, + cmh_reg_read32(rh.cfg->mailboxes[i].reg_base, + R_MBX_QUEUE_TAIL), + MBX_STATUS_CORE_ID(status), + MBX_STATUS_ERROR_CODE(status), + MBX_STATUS_CMD_INDEX(status)); + irq_bits =3D MBX_ERROR_IRQ; + } else { + /* eSW cleared ERROR -- recovery succeeded */ + WRITE_ONCE(rh.mbx[i].restart_pending, false); + rh.mbx[i].restart_retries =3D 0; + rh.mbx[i].flush_count =3D 0; + } + + if (new_head !=3D rh.mbx[i].last_head || irq_bits) { + if (new_head !=3D rh.mbx[i].last_head) + dev_dbg_ratelimited(cmh_dev(), + "rh: watchdog: mbx[%u] head %u->%u (missed IRQ recovery)\n", + i, rh.mbx[i].last_head, + new_head); + cmh_rh_process_mbx(i, new_head, irq_bits); + rh.mbx[i].last_head =3D new_head; + rh.mbx[i].abort_stall_ticks =3D 0; + } + + /* + * Abort-stall detector: if the head-of-queue transaction + * timed out (state =3D=3D TXN_TIMED_OUT) but the eSW hasn't + * responded (HEAD didn't advance, no ERROR status): + * + * tick 1: issue MBX_COMMAND_ABORT (serialised + * under rh_process_lock -- safe against + * concurrent RESTART/FLUSH) + * ticks 2..N-1: wait for eSW to respond with ERROR + * tick N: escalate to FLUSH + force-drain + * + * If the eSW responds with ERROR between ticks, the ERROR + * status branch above handles RESTART recovery and resets + * abort_stall_ticks via the restart_pending guard. + */ + if (!READ_ONCE(rh.mbx[i].wedged) && + !READ_ONCE(rh.mbx[i].restart_pending)) { + struct transaction_obj *head_txn; + + head_txn =3D cmh_tm_peek_transaction(i); + if (head_txn && + atomic_read(&head_txn->state) =3D=3D TXN_TIMED_OUT) { + unsigned int stall_max; + void __iomem *base =3D + rh.cfg->mailboxes[i].reg_base; + + rh.mbx[i].abort_stall_ticks++; + + if (rh.mbx[i].abort_stall_ticks =3D=3D 1) { + dev_warn(cmh_dev(), + "rh: watchdog: mbx[%u] head txn timed out, issuing ABORT\n", + i); + cmh_reg_write32(MBX_COMMAND_ABORT, + base, + R_MBX_COMMAND); + } + + stall_max =3D DIV_ROUND_UP(CMH_RH_ABORT_STALL_MS, + max(watchdog_ms, + CMH_RH_WATCHDOG_MS_MIN)); + if (rh.mbx[i].abort_stall_ticks >=3D + stall_max) { + dev_err(cmh_dev(), + "rh: watchdog: mbx[%u] abort stall (%u ticks) -- FLUSH + drain\n", + i, rh.mbx[i].abort_stall_ticks); + cmh_reg_write32(MBX_COMMAND_FLUSH, + base, R_MBX_COMMAND); + cmh_rh_drain_mbx(i, -ETIMEDOUT); + rh.mbx[i].abort_stall_ticks =3D 0; + } + } else { + rh.mbx[i].abort_stall_ticks =3D 0; + } + } + spin_unlock(&rh_process_lock); + } + + if (rh.active) { + unsigned int wdog =3D max(watchdog_ms, CMH_RH_WATCHDOG_MS_MIN); + + mod_timer(&rh_watchdog, + jiffies + msecs_to_jiffies(wdog)); + } +} + +/* + * Resolve per-MBX Linux virqs for the CMH interrupt lines. + * + * Each mailbox declares its own completion interrupt in its device-tree + * child node; cmh_config_init() resolves these to Linux virqs and stores + * them in cfg->mailboxes[i].irq (-1 when the mailbox has no interrupt). + * IRQ mode requires every configured mailbox to have an interrupt; if + * none do (or only some), the response handler uses watchdog polling. + * + * Populates rh.irqs[] and rh.nirqs. Returns 0 on success, or a + * negative errno if no IRQs could be resolved (polling-only mode). + */ +static int cmh_rh_resolve_irqs(struct cmh_config *cfg) +{ + u32 i, nwith =3D 0; + + rh.nirqs =3D 0; + + for (i =3D 0; i < cfg->mbx_count; i++) + if (cfg->mailboxes[i].irq >=3D 0) + nwith++; + + if (nwith =3D=3D 0) + return -ENODEV; + + if (nwith !=3D cfg->mbx_count) { + dev_warn(cmh_dev(), + "rh: only %u/%u mailboxes have IRQs -- falling back to polling\n", + nwith, cfg->mbx_count); + return -ENODEV; + } + + /* + * Collect distinct virqs. Multiple mailboxes may be wired to the + * same line; the threaded handler scans all mailboxes, so each + * distinct line needs only one registration. Requesting (and later + * freeing) a shared line once per mailbox would fail -EBUSY / double + * free. + */ + rh.nirqs =3D 0; + for (i =3D 0; i < cfg->mbx_count; i++) { + int irq =3D cfg->mailboxes[i].irq; + bool seen =3D false; + u32 j; + + for (j =3D 0; j < rh.nirqs; j++) + if (rh.irqs[j] =3D=3D irq) { + seen =3D true; + break; + } + if (seen) + continue; + + rh.irqs[rh.nirqs++] =3D irq; + dev_dbg(cmh_dev(), "rh: MBX%u -> IRQ %d\n", i, irq); + } + + return 0; +} + +/** + * cmh_rh_init() - Initialize the response handler + * @cfg: Device configuration (mailbox count, MMIO bases, IRQ info) + * + * Resolve per-mailbox IRQs from the device tree (or module parameter + * override), register threaded IRQ handlers (hardirq + kthread), and + * arm the missed-IRQ software watchdog timer. If no IRQs can be + * resolved, falls back to watchdog-only polling mode. + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_rh_init(struct cmh_config *cfg) +{ + unsigned long irqflags; + int ret; + u32 i; + + rh.cfg =3D cfg; + rh.nirqs =3D 0; + rh.active =3D false; + atomic_set(&rh.irq_count, 0); + + /* Allocate per-MBX tracking */ + rh.mbx =3D kcalloc(cfg->mbx_count, sizeof(*rh.mbx), GFP_KERNEL); + if (!rh.mbx) + return -ENOMEM; + + /* Resolve per-MBX IRQs */ + if (cmh_rh_resolve_irqs(cfg) < 0) { + /* + * No IRQs available. The watchdog timer provides + * a polling fallback: it reads MBX head registers + * periodically and processes completions. This is + * slower than IRQ-driven completion but functional. + * + * Completion latency in polling-only mode is bounded + * by the watchdog interval (default 200 ms, tunable + * via debugfs config/watchdog_ms). + */ + dev_warn(cmh_dev(), + "rh: no IRQs -- using watchdog polling (interval %u ms)\n", + watchdog_ms); + + /* Seed last_head from HW before first watchdog tick */ + for (i =3D 0; i < cfg->mbx_count; i++) + rh.mbx[i].last_head =3D + cmh_reg_read32(cfg->mailboxes[i].reg_base, + R_MBX_QUEUE_HEAD); + + rh.active =3D true; + timer_setup(&rh_watchdog, cmh_rh_watchdog_fn, 0); + mod_timer(&rh_watchdog, jiffies + + msecs_to_jiffies(max(watchdog_ms, + CMH_RH_WATCHDOG_MS_MIN))); + return 0; + } + + /* Initialize per-MBX state: read current head positions */ + for (i =3D 0; i < cfg->mbx_count; i++) + rh.mbx[i].last_head =3D cmh_reg_read32(rh.cfg->mailboxes[i].reg_base, + R_MBX_QUEUE_HEAD); + + /* + * Register threaded IRQ handlers, one per DISTINCT line + * (cmh_rh_resolve_irqs() de-duplicated shared virqs). The handler + * scans all mailboxes unconditionally, so one registration per line + * covers every mailbox wired to it. + * + * When de-dup collapsed lines (nirqs < mbx_count) at least one line + * is shared across mailboxes -- register IRQF_SHARED. + */ + irqflags =3D (rh.nirqs < cfg->mbx_count) ? IRQF_SHARED : 0; + + for (i =3D 0; i < rh.nirqs; i++) { + ret =3D request_threaded_irq(rh.irqs[i], + cmh_rh_hardirq, + cmh_rh_thread, + irqflags, + "cmh", cfg); + if (ret) { + dev_err(cmh_dev(), "rh: request_threaded_irq(%d) for MBX%u failed (rc= =3D%d)\n", + rh.irqs[i], i, ret); + /* Unwind previously registered IRQs */ + while (i--) + free_irq(rh.irqs[i], cfg); + rh.nirqs =3D 0; + kfree(rh.mbx); + rh.mbx =3D NULL; + return ret; + } + } + + rh.active =3D true; + + /* Enable MBX completion interrupts (DONE + ERROR) */ + for (i =3D 0; i < cfg->mbx_count; i++) { + u32 stale; + + /* + * W1C any interrupt bits that accumulated between + * MQI setup and now (e.g. CMH eSW processing stale + * commands) before enabling the mask. + */ + stale =3D cmh_reg_read32(cfg->mailboxes[i].reg_base, + R_MBX_INTERRUPT); + if (stale) + cmh_reg_write32(stale, cfg->mailboxes[i].reg_base, + R_MBX_INTERRUPT); + + cmh_reg_write32(MBX_IRQ_MASK, + cfg->mailboxes[i].reg_base, + R_MBX_INTERRUPT_MASK); + } + + /* Arm missed-IRQ watchdog timer */ + timer_setup(&rh_watchdog, cmh_rh_watchdog_fn, 0); + mod_timer(&rh_watchdog, jiffies + + msecs_to_jiffies(max(watchdog_ms, + CMH_RH_WATCHDOG_MS_MIN))); + + return 0; +} + +/** + * cmh_rh_suspend() - Suspend the response handler + * @cfg: Device configuration + * + * Stop the watchdog timer and mask mailbox interrupts at the hardware + * level. The IRQ handlers remain registered so that resume can + * re-enable them without re-requesting. + */ +void cmh_rh_suspend(struct cmh_config *cfg) +{ + u32 i; + + if (!rh.active) + return; + + /* + * Clear rh.active before deleting the timer: the watchdog re-arms + * itself only while rh.active is set, so an in-flight callback will + * not re-queue the timer once this is cleared. Keep + * timer_delete_sync (not _shutdown) so cmh_rh_resume() can re-arm. + */ + rh.active =3D false; + + /* Stop the watchdog before masking HW interrupts */ + timer_delete_sync(&rh_watchdog); + + /* Mask MBX interrupts at the hardware level */ + for (i =3D 0; i < cfg->mbx_count; i++) + cmh_reg_write32(0, cfg->mailboxes[i].reg_base, + R_MBX_INTERRUPT_MASK); + + /* + * Ensure no threaded IRQ handler is still in-flight. + * After masking, a handler may already have been scheduled. + * synchronize_irq() waits for it to complete before we + * proceed with suspend (which tears down TM state). + */ + for (i =3D 0; i < rh.nirqs; i++) + synchronize_irq(rh.irqs[i]); +} + +/** + * cmh_rh_resume() - Resume the response handler after suspend + * @cfg: Device configuration + * + * Re-synchronize per-mailbox head tracking with hardware, clear stale + * interrupt bits accumulated during the power transition, re-enable + * mailbox completion interrupts, and re-arm the watchdog timer. + */ +void cmh_rh_resume(struct cmh_config *cfg) +{ + u32 i; + + if (!rh.mbx || !cfg) + return; + + /* Re-sync per-MBX head tracking with hardware */ + for (i =3D 0; i < cfg->mbx_count; i++) { + u32 stale; + + rh.mbx[i].last_head =3D + cmh_reg_read32(cfg->mailboxes[i].reg_base, + R_MBX_QUEUE_HEAD); + + /* W1C any stale interrupt bits from the power transition */ + stale =3D cmh_reg_read32(cfg->mailboxes[i].reg_base, + R_MBX_INTERRUPT); + if (stale) + cmh_reg_write32(stale, cfg->mailboxes[i].reg_base, + R_MBX_INTERRUPT); + + /* Re-enable MBX completion interrupts */ + cmh_reg_write32(MBX_IRQ_MASK, cfg->mailboxes[i].reg_base, + R_MBX_INTERRUPT_MASK); + } + + rh.active =3D true; + + /* Re-arm the watchdog */ + mod_timer(&rh_watchdog, jiffies + + msecs_to_jiffies(max(watchdog_ms, + CMH_RH_WATCHDOG_MS_MIN))); +} + +/** + * cmh_rh_cleanup() - Clean up the response handler + * @cfg: Device configuration + * + * Stop the watchdog timer, mask mailbox interrupts at the hardware + * level, release all registered IRQ handlers, and free per-mailbox + * tracking state. Safe to call even if init was never completed. + */ +void cmh_rh_cleanup(struct cmh_config *cfg) +{ + if (rh.active) { + u32 i; + + /* + * Clear rh.active first (the watchdog re-arm is gated on + * it), then shut the timer down: after timer_shutdown_sync() + * a stray mod_timer() is a no-op, so the callback cannot + * resurrect the timer before rh.mbx is freed. A later probe + * re-inits it via timer_setup() in cmh_rh_init(). + */ + rh.active =3D false; + + /* Cancel watchdog before disabling interrupts */ + timer_shutdown_sync(&rh_watchdog); + + /* Disable MBX interrupts before releasing handlers */ + for (i =3D 0; i < cfg->mbx_count; i++) + cmh_reg_write32(0, + cfg->mailboxes[i].reg_base, + R_MBX_INTERRUPT_MASK); + + /* Release all per-MBX IRQs */ + for (i =3D 0; i < rh.nirqs; i++) + free_irq(rh.irqs[i], cfg); + dev_dbg(cmh_dev(), "rh: %u IRQs released\n", rh.nirqs); + rh.nirqs =3D 0; + } + + dev_dbg(cmh_dev(), "rh: %u IRQs handled\n", + atomic_read(&rh.irq_count)); + + kfree(rh.mbx); + rh.mbx =3D NULL; +} + +/* -- debugfs timeout accessor ------------------------------------------ = */ + +#ifdef CONFIG_CRYPTO_DEV_CMH_DEBUG +/** + * cmh_rh_timeout_watchdog_ptr() - Return pointer to watchdog_ms for debug= fs + * + * Exposes the Response Handler watchdog timeout for runtime tuning + * via debugfs config/ directory. + * + * Return: pointer to the static watchdog_ms variable. + */ +unsigned int *cmh_rh_timeout_watchdog_ptr(void) { return &watchdog_ms; } +#endif diff --git a/drivers/crypto/cmh/cmh_sysfs.c b/drivers/crypto/cmh/cmh_sysfs.c new file mode 100644 index 000000000000..ab482a222167 --- /dev/null +++ b/drivers/crypto/cmh/cmh_sysfs.c @@ -0,0 +1,108 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- sysfs Device Attributes + * + * Exposes hardware identity and status as read-only sysfs attributes + * under /sys/devices/platform/cmh/. Wired via .dev_groups in the + * platform_driver struct -- the driver core creates and removes these + * automatically around .probe() / .remove(). + * + * Because .dev_groups is used (not manual sysfs_create_group), the + * driver core guarantees that attributes are created after .probe() + * sets drvdata and removed before .remove() clears it. Therefore + * platform_get_drvdata() cannot return NULL in any show callback and + * no NULL check is needed. Same pattern as caam/ctrl.c and + * ccree/cc_sysfs.c. + */ + +#include +#include +#include + +#include "cmh.h" +#include "cmh_registers.h" +#include "cmh_sysfs.h" + +static ssize_t fw_version_show(struct device *dev, + struct device_attribute *attr, char *buf) +{ + struct cmh_device *cmh =3D platform_get_drvdata(to_platform_device(dev)); + struct cmh_config *cfg =3D &cmh->config; + + if (!cfg->sic_mapped) + return -ENODEV; + + return sysfs_emit(buf, "0x%08x\n", + cmh_reg_read32(cfg->sic_mapped, R_SIC_SW_VERSION)); +} +static DEVICE_ATTR_RO(fw_version); + +static ssize_t hw_version_show(struct device *dev, + struct device_attribute *attr, char *buf) +{ + struct cmh_device *cmh =3D platform_get_drvdata(to_platform_device(dev)); + struct cmh_config *cfg =3D &cmh->config; + + if (!cfg->sic_mapped) + return -ENODEV; + + return sysfs_emit(buf, "0x%08x\n", + cmh_reg_read32(cfg->sic_mapped, R_SIC_HW_VERSION0)); +} +static DEVICE_ATTR_RO(hw_version); + +static ssize_t boot_status_show(struct device *dev, + struct device_attribute *attr, char *buf) +{ + struct cmh_device *cmh =3D platform_get_drvdata(to_platform_device(dev)); + struct cmh_config *cfg =3D &cmh->config; + + if (!cfg->sic_mapped) + return -ENODEV; + + return sysfs_emit(buf, "0x%08x\n", + cmh_reg_read32(cfg->sic_mapped, R_SIC_BOOT_STATUS)); +} +static DEVICE_ATTR_RO(boot_status); + +static ssize_t mbx_available_show(struct device *dev, + struct device_attribute *attr, char *buf) +{ + struct cmh_device *cmh =3D platform_get_drvdata(to_platform_device(dev)); + struct cmh_config *cfg =3D &cmh->config; + + if (!cfg->sic_mapped) + return -ENODEV; + + return sysfs_emit(buf, "0x%08x\n", + cmh_reg_read32(cfg->sic_mapped, R_SIC_MBX_AVAILABILITY)); +} +static DEVICE_ATTR_RO(mbx_available); + +static ssize_t mbx_count_show(struct device *dev, + struct device_attribute *attr, char *buf) +{ + struct cmh_device *cmh =3D platform_get_drvdata(to_platform_device(dev)); + + return sysfs_emit(buf, "%u\n", cmh->config.mbx_count); +} +static DEVICE_ATTR_RO(mbx_count); + +static struct attribute *cmh_sysfs_attrs[] =3D { + &dev_attr_fw_version.attr, + &dev_attr_hw_version.attr, + &dev_attr_boot_status.attr, + &dev_attr_mbx_available.attr, + &dev_attr_mbx_count.attr, + NULL, +}; + +static const struct attribute_group cmh_sysfs_group =3D { + .attrs =3D cmh_sysfs_attrs, +}; + +const struct attribute_group *cmh_sysfs_groups[] =3D { + &cmh_sysfs_group, + NULL, +}; diff --git a/drivers/crypto/cmh/cmh_txn.c b/drivers/crypto/cmh/cmh_txn.c new file mode 100644 index 000000000000..0900f903a58a --- /dev/null +++ b/drivers/crypto/cmh/cmh_txn.c @@ -0,0 +1,2173 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Transaction Manager + * + * Dedicated kthread that dequeues command messages, builds VCQs in + * DMA queue slots, and rings the MBX doorbell. + * + * Command flow: + * 1. Caller posts command_msg via cmh_tm_post_command() + * 2. TM thread wakes, dequeues msg from CMQ + * 3. Selects mailbox (core-to-MBX affinity, or caller-pinned) + * 4. Copies pre-built VCQ entries into DMA slot at tail + * 5. Creates transaction_obj, appends to per-MBX txn queue + * 6. Writes tail+1 -> R_MBX_QUEUE_TAIL (doorbell) + * + * The Response Handler (cmh_rh.c) walks per-MBX txn queues + * when an IRQ fires and the head advances, firing completion callbacks. + * + * Transaction state machine + * ------------------------- + * Each async transaction moves through the following states. DMA + * buffers remain mapped and owned by the HW until the COMPLETE state + * is reached -- only then are they safe to unmap/free. + * + * QUEUED --[TM posts to HW]--> INFLIGHT + * (cmq) | | \ + * | | \--[timer fires]--> + * | | TIMED_OUT + * | | | + * | [HW completes / [HW completes / + * | RH pops txn] RH pops txn] + * | | | + * | v v + * | COMPLETE COMPLETE + * | (err=3DHW rc) (err=3D-ETIMEDOUT) + * | + * +--[pre-submit fail]--> freed (callback never fires) + * + * Note: QUEUED is the command_msg phase (sitting in the CMQ list, + * not yet a transaction_obj). The transaction_obj states tracked + * by atomic_cmpxchg are INFLIGHT, TIMED_OUT, and COMPLETE only. + * + * Completion callback context guarantee: + * The crypto_request_complete() callback is invoked from one of: + * - The RH threaded IRQ handler (process context, BH disabled) + * - The watchdog timer (softirq / timer context) + * - The TM kthread during queue drain/cleanup (process context) + * + * It is NEVER invoked from hardirq context. + * + * The watchdog path runs from timer softirq because it must recover + * missed IRQs without sleeping. This is crypto-API-compliant: + * crypto_request_complete() is documented safe from any context + * (including softirq). Callers must NOT assume process context in + * their completion callbacks -- all operations therein must be + * softirq-safe (no mutex, no GFP_KERNEL, no sleeping locks). + * + * For backlog promotion the -EINPROGRESS notification is issued with + * the CMQ spinlock dropped, so a consumer may safely resubmit (or take + * sleeping locks) from that callback without re-entering cmq_lock. + */ + +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "cmh_txn.h" +#include "cmh_rh.h" +#include "cmh_registers.h" +#include "cmh_config.h" +#include "cmh_vcq.h" +#include "cmh_debugfs.h" +#include "cmh_dma.h" + +/* Module State */ + +static struct { + struct cmh_config *cfg; + struct task_struct *thread; + /* + * Gates submission acceptance in cmh_tm_post_command(); the TM + * thread's lifetime is governed by kthread_should_stop(), not this. + */ + bool running; + + /* Command Message Queue (CMQ) */ + struct list_head cmq; + spinlock_t cmq_lock; /* protects cmq + backlog lists */ + wait_queue_head_t cmq_waitq; + + /* Backlog queue for CRYPTO_TFM_REQ_MAY_BACKLOG requests */ + struct list_head backlog; + u32 backlog_depth; + + /* Per-mailbox transaction queues */ + struct cmh_mbx_txq *txqs; /* array[cfg->mbx_count] */ + + /* Round-robin mailbox selector */ + u32 next_mbx; +} tm; + +/* + * Command Message Queue and backlog depths. Built-in defaults; a debug + * build can tune them at runtime via debugfs config/{cmq_max_depth, + * backlog_max_depth}. Change the #defines to alter the compiled-in + * values (backlog 0 disables the backlog). + */ +#define CMH_TM_CMQ_MAX_DEPTH 256 +#define CMH_TM_BACKLOG_MAX_DEPTH 1024 + +static unsigned int cmq_max_depth =3D CMH_TM_CMQ_MAX_DEPTH; +static unsigned int backlog_max_depth =3D CMH_TM_BACKLOG_MAX_DEPTH; + +static unsigned int async_timeout_ms =3D 2000; + +#define CMH_TM_BACKOFF_MIN_US 100 /* queue-full backoff range (us) */ +#define CMH_TM_BACKOFF_MAX_US 500 +static unsigned int cmq_depth; /* current CMQ depth, protected by tm= .cmq_lock */ + +/* + * Monotonically increasing counter bumped by cmh_tm_txq_completion_notify= (). + * Used as a generation check in the queue-full backoff predicate so that + * wait_event_interruptible_timeout() returns immediately when a TXQ + * completion frees a slot, rather than sleeping for the full timeout. + */ +static atomic_t txq_completion_gen; + +/* -- Debugfs stat helpers (avoid anonymous compound blocks) -------------= */ + +static void cmh_stat_inc_mbx_queue_full(u32 mbx_idx) +{ + struct cmh_mbx_stats *s =3D cmh_debugfs_mbx_stats(mbx_idx); + + if (s) + atomic64_inc(&s->queue_full_count); +} + +static void cmh_stat_record_vcq_submit(u32 mbx_idx, u32 num_vcqs, u32 dept= h) +{ + struct cmh_mbx_stats *s =3D cmh_debugfs_mbx_stats(mbx_idx); + + if (s) { + atomic64_add(num_vcqs, &s->vcqs_submitted); + cmh_stat_update_max(&s->max_queue_depth, (s64)depth); + } +} + +static void cmh_stat_inc_tm_backoff(void) +{ + struct cmh_tm_stats *s =3D cmh_debugfs_tm_stats(); + + if (s) + atomic64_inc(&s->backoff_count); +} + +static void cmh_stat_inc_cmq_eagain(void) +{ + struct cmh_tm_stats *s =3D cmh_debugfs_tm_stats(); + + if (s) + atomic64_inc(&s->cmq_eagain_count); +} + +static void cmh_stat_record_cmq_post(u32 depth) +{ + struct cmh_tm_stats *s =3D cmh_debugfs_tm_stats(); + + if (s) { + atomic64_inc(&s->cmq_posts); + cmh_stat_update_max(&s->cmq_depth_max, (s64)depth); + } +} + +static void cmh_stat_inc_async_timeout(void) +{ + struct cmh_tm_stats *s =3D cmh_debugfs_tm_stats(); + + if (s) + atomic64_inc(&s->async_timeout_count); +} + +/* + * Drop one reference on a command_msg; free when the last ref is dropped. + * Used by cmh_tm_submit_sync() to share msg ownership between the + * waiter (caller) and the TM subsystem (thread or cleanup drain). + */ +static void command_msg_put(struct command_msg *msg) +{ + if (refcount_dec_and_test(&msg->refs)) { + kfree(msg->vcq_data); + kfree(msg); + } +} + +/* + * Drop one reference on a transaction_obj; free when the last ref drops. + * Two references are held when the per-request timeout timer is armed: + * one for the TXQ owner (RH/cleanup), one for the timer callback. + * When no timer is armed, only the owner ref exists. + */ +static void txn_put(struct transaction_obj *txn) +{ + if (refcount_dec_and_test(&txn->refs)) + kfree(txn); +} + +/* + * Per-request async timeout callback (runs in softirq / timer context). + * + * This function ONLY marks the transaction state as TIMED_OUT via + * atomic cmpxchg and drops the timer reference. It does NOT fire + * the completion callback, does NOT touch DMA buffers, and does NOT + * write any MBX registers. + * + * Rationale: the HW may still be writing to DMA buffers at this + * point. Unmapping or freeing them here would be a use-after-free. + * The actual -ETIMEDOUT completion fires later, from process + * context, when the RH threaded IRQ pops the transaction after the + * HW finishes (or after MBX abort/drain on rmmod/suspend). + * + * MBX_COMMAND_ABORT is NOT issued here. It is issued by the RH + * watchdog abort-stall detector under rh_process_lock, which + * serialises it against RESTART/FLUSH recovery commands. Writing + * ABORT from timer softirq without the lock caused a race where + * concurrent timeouts clobbered an in-progress RESTART, wedging + * the mailbox. + * + * Context: softirq (timer). Must not sleep. + */ +static void txn_timeout_fn(struct timer_list *t) +{ + struct transaction_obj *txn =3D timer_container_of(txn, t, timeout_timer); + int old; + + old =3D atomic_cmpxchg(&txn->state, TXN_INFLIGHT, TXN_TIMED_OUT); + if (old =3D=3D TXN_INFLIGHT) { + dev_err_ratelimited(cmh_dev(), + "tm: async timeout vcq=3D%u..%u mbx=3D%u cmd_id=3D0x%08x\n", + txn->first_vcq_id, + txn->last_vcq_id, txn->mailbox_idx, + txn->command_id); + cmh_stat_inc_async_timeout(); + } + + txn_put(txn); /* drop timer ref */ +} + +/** + * cmh_txn_finish() - Complete a popped transaction with FSM + timer clean= up + * @txn: Transaction object to complete + * @error: Error code from HW (0 on success) + * + * Three cases: + * 1. Normal: state INFLIGHT -> COMPLETE. Fire callback with HW error. + * 2. Timed out: state already TXN_TIMED_OUT (timer marked it). + * Fire callback with -ETIMEDOUT. DMA is now safe because the + * HW has finished and HEAD has advanced past this VCQ. + * 3. Force-cancel (drain/quiesce): handled by caller, not here. + */ +void cmh_txn_finish(struct transaction_obj *txn, int error) +{ + int old; + + old =3D atomic_cmpxchg(&txn->state, TXN_INFLIGHT, TXN_COMPLETE); + + /* Dequeue the timer if still pending; drop timer ref if we did */ + if (timer_delete(&txn->timeout_timer)) + txn_put(txn); + + if (old =3D=3D TXN_INFLIGHT) { + /* HW completion (may carry error) */ + if (txn->complete) + txn->complete(txn->completion_data, error); + } else if (old =3D=3D TXN_TIMED_OUT) { + /* Timer won earlier; now HW is done -- deliver -ETIMEDOUT */ + if (txn->complete) + txn->complete(txn->completion_data, -ETIMEDOUT); + } + + txn_put(txn); /* drop owner ref */ +} + +/* Mailbox Slot Addressing */ + +/* + * Return a kernel-virtual pointer to the VCQ slot for the given vcqid. + * Mirrors CMH eSW's mbx_queue_addr() but uses the kernel virt_addr. + */ +static void *mbx_slot_ptr(struct cmh_mbx_config *mbx, u32 vcqid) +{ + u32 slot_mask =3D (1U << mbx->slots_log2) - 1U; + u32 slot_offset =3D (vcqid & slot_mask) << mbx->stride_log2; + + return (u8 *)mbx->virt_addr + slot_offset; +} + +/* + * Return the number of free slots in a mailbox queue. + */ +static u32 mbx_free_slots(struct cmh_mbx_config *mbx) +{ + u32 head =3D cmh_reg_read32(mbx->reg_base, R_MBX_QUEUE_HEAD); + u32 tail =3D cmh_reg_read32(mbx->reg_base, R_MBX_QUEUE_TAIL); + u32 size =3D 1U << mbx->slots_log2; + + return size - (u32)(tail - head); +} + +/** + * cmh_tm_max_cmds_per_vcq() - Return max commands per VCQ slot + * + * Scans all mailbox configurations and returns the minimum number of + * VCQ command entries that fit in a single slot, clamped to the + * MIN_VCQ_CMDS..MAX_VCQ_CMDS range. + * + * Return: Maximum usable VCQ command count per slot. + */ +u32 cmh_tm_max_cmds_per_vcq(void) +{ + u32 i, min_cmds =3D MAX_VCQ_CMDS; + + for (i =3D 0; i < tm.cfg->mbx_count; i++) { + u32 stride =3D 1U << tm.cfg->mailboxes[i].stride_log2; + u32 cmds =3D stride / (u32)sizeof(struct vcq_cmd); + + if (cmds < min_cmds) + min_cmds =3D cmds; + } + + if (min_cmds < MIN_VCQ_CMDS) + min_cmds =3D MIN_VCQ_CMDS; + + return min_cmds; +} + +/** + * cmh_tm_mbx_count() - Return the number of configured mailboxes + * + * Return: Number of mailboxes in the current configuration. + */ +u32 cmh_tm_mbx_count(void) +{ + return tm.cfg->mbx_count; +} + +/* Core-to-MBX Affinity -- Config-Driven Multi-Instance Support */ + +/* + * Per-core-type configuration table. Each entry holds one or more + * (core_id, mbx_idx) instances. Defaults: single instance per core + * type with the standard CORE_ID_* and MBX auto-assigned on first use + * (mbx_idx =3D -1). Module params can override for explicit assignment + * and multi-instance support. + * + * Round-robin across instances for each new crypto operation. + */ + +struct core_instance_info { + u32 core_id; /* VCQ dispatch core_id */ + /* + * Assigned MBX index, or -1 (sentinel) for auto-assign on first + * use. Uses atomic_t for a lockless once-only latch: the first + * caller does atomic_cmpxchg(&mbx_idx, -1, new_mbx); all later + * callers see the winning value via atomic_read(). + */ + atomic_t mbx_idx; +}; + +struct core_type_info { + u32 num_instances; + struct core_instance_info instances[CMH_MAX_CORE_INSTANCES]; + atomic_t next_instance; /* round-robin counter */ +}; + +static struct core_type_info core_types[CMH_NUM_CORE_TYPES] =3D { + [CMH_CORE_HC] =3D { .num_instances =3D 1, + .instances =3D { { .core_id =3D CORE_ID_HC, + .mbx_idx =3D ATOMIC_INIT(-1) } } }, + [CMH_CORE_AES] =3D { .num_instances =3D 1, + .instances =3D { { .core_id =3D CORE_ID_AES, + .mbx_idx =3D ATOMIC_INIT(-1) } } }, + [CMH_CORE_SM4] =3D { .num_instances =3D 1, + .instances =3D { { .core_id =3D CORE_ID_SM4, + .mbx_idx =3D ATOMIC_INIT(-1) } } }, + [CMH_CORE_SM3] =3D { .num_instances =3D 1, + .instances =3D { { .core_id =3D CORE_ID_SM3, + .mbx_idx =3D ATOMIC_INIT(-1) } } }, + [CMH_CORE_CCP] =3D { .num_instances =3D 1, + .instances =3D { { .core_id =3D CORE_ID_CCP, + .mbx_idx =3D ATOMIC_INIT(-1) } } }, + [CMH_CORE_PKE] =3D { .num_instances =3D 1, + .instances =3D { { .core_id =3D CORE_ID_PKE, + .mbx_idx =3D ATOMIC_INIT(-1) } } }, + [CMH_CORE_QSE] =3D { .num_instances =3D 1, + .instances =3D { { .core_id =3D CORE_ID_QSE, + .mbx_idx =3D ATOMIC_INIT(-1) } } }, + [CMH_CORE_HCQ] =3D { .num_instances =3D 1, + .instances =3D { { .core_id =3D CORE_ID_HCQ, + .mbx_idx =3D ATOMIC_INIT(-1) } } }, +}; + +/* Round-robin counter for auto-assigning MBXes to core instances */ +static atomic_t affinity_next_mbx =3D ATOMIC_INIT(0); + +/** + * cmh_tm_affinity_reset() - Reset core-to-MBX affinity state + * + * Clears all auto-assigned MBX bindings and resets round-robin + * counters for both the global MBX allocator and per-core-type + * instance selectors. + */ +void cmh_tm_affinity_reset(void) +{ + u32 i, j; + + atomic_set(&affinity_next_mbx, 0); + + /* Reset multi-instance table */ + for (i =3D 0; i < CMH_NUM_CORE_TYPES; i++) { + struct core_type_info *ct =3D &core_types[i]; + + atomic_set(&ct->next_instance, 0); + for (j =3D 0; j < ct->num_instances; j++) + atomic_set(&ct->instances[j].mbx_idx, -1); + } +} + +/** + * cmh_core_default_id() - Return default core_id for a core type + * @type: Core type selector + * + * Returns the first-instance core_id for @type without advancing the + * round-robin counter. Used by callers pinned to a fixed MBX (e.g. + * mgmt ioctls on MGMT_MBX) that only need the VCQ core_id field. + * + * Return: VCQ core_id value for the default instance of @type. + */ +u32 cmh_core_default_id(enum cmh_core_type type) +{ + if (WARN_ON_ONCE(type >=3D CMH_NUM_CORE_TYPES)) + return CORE_ID_NUM; + + return core_types[type].instances[0].core_id; +} + +/** + * cmh_core_select_instance() - Select a core instance via round-robin + * @type: Core type selector + * + * Round-robin across configured instances, each permanently pinned to + * its MBX (auto-assigned on first use if mbx_idx was -1). + * + * Uses atomic_inc_return (pre-increment), so the very first call for a + * given type returns instance[1 % N]. Over the lifetime of the module + * the distribution is perfectly balanced; the off-by-one only affects + * the first cycle. + * + * The (u32) cast before the modulo ensures correct behaviour across + * the INT_MAX -> INT_MIN wraparound of atomic_t: (u32)INT_MIN =3D + * 0x80000000, and 0x80000000 % N still yields a valid index. + * + * Return: A core_dispatch with (core_id, mbx_idx) for the selected + * instance. + */ +struct core_dispatch cmh_core_select_instance(enum cmh_core_type type) +{ + struct core_type_info *ct; + struct core_instance_info *inst; + struct core_dispatch d; + u32 idx, count; + s32 mbx, new_mbx, old; + + if (WARN_ON_ONCE(type >=3D CMH_NUM_CORE_TYPES)) + return (struct core_dispatch){ .core_id =3D CORE_ID_NUM, .mbx_idx =3D -1= }; + + ct =3D &core_types[type]; + + /* + * A core type absent from CORE_ENABLE has num_instances =3D=3D 0 (algs + * for it should never have been reached). Return CORE_ID_NUM, the + * out-of-range core sentinel: if a caller ignores this and submits, + * the eSW rejects the unknown core rather than the VCQ landing on + * CORE_ID_SYS (0). Also avoids the divide-by-zero below. + */ + if (WARN_ONCE(!ct->num_instances, + "cmh: dispatch for absent core type %u\n", type)) + return (struct core_dispatch){ .core_id =3D CORE_ID_NUM, .mbx_idx =3D -1= }; + + idx =3D (u32)atomic_inc_return(&ct->next_instance) % ct->num_instances; + inst =3D &ct->instances[idx]; + + d.core_id =3D inst->core_id; + + mbx =3D atomic_read(&inst->mbx_idx); + if (mbx >=3D 0) { + d.mbx_idx =3D mbx; + return d; + } + + /* Auto-assign on first use */ + count =3D tm.cfg->mbx_count; + new_mbx =3D (s32)((u32)atomic_inc_return(&affinity_next_mbx) % count); + old =3D atomic_cmpxchg(&inst->mbx_idx, -1, new_mbx); + + if (old >=3D 0) + d.mbx_idx =3D old; + else + d.mbx_idx =3D new_mbx; + + return d; +} + +/** + * cmh_core_num_instances() - Return instance count for a core type + * @type: Core type selector + * + * Return: Number of configured instances for @type, or 0 if the core is + * absent from the CORE_ENABLE register. + */ +u32 cmh_core_num_instances(enum cmh_core_type type) +{ + if (WARN_ON_ONCE(type >=3D CMH_NUM_CORE_TYPES)) + return 1; + + return core_types[type].num_instances; +} + +/** + * cmh_core_get_instance() - Get dispatch info for a specific instance + * @type: Core type selector + * @idx: Instance index within @type + * + * Returns (core_id, mbx_idx) for a specific instance by index, + * without advancing the round-robin counter. Triggers MBX auto-assign + * on first use if the instance has no MBX yet. + * + * Return: A core_dispatch with (core_id, mbx_idx) for instance @idx. + */ +struct core_dispatch cmh_core_get_instance(enum cmh_core_type type, u32 id= x) +{ + struct core_type_info *ct; + struct core_instance_info *inst; + struct core_dispatch d; + u32 count; + s32 mbx, new_mbx, old; + + if (WARN_ON_ONCE(type >=3D CMH_NUM_CORE_TYPES)) + return (struct core_dispatch){ .core_id =3D CORE_ID_NUM, .mbx_idx =3D -1= }; + + ct =3D &core_types[type]; + if (WARN_ON_ONCE(idx >=3D ct->num_instances)) + return (struct core_dispatch){ .core_id =3D CORE_ID_NUM, .mbx_idx =3D -1= }; + + inst =3D &ct->instances[idx]; + d.core_id =3D inst->core_id; + + mbx =3D atomic_read(&inst->mbx_idx); + if (mbx >=3D 0) { + d.mbx_idx =3D mbx; + return d; + } + + /* Auto-assign on first use */ + count =3D tm.cfg->mbx_count; + new_mbx =3D (s32)((u32)atomic_inc_return(&affinity_next_mbx) % count); + old =3D atomic_cmpxchg(&inst->mbx_idx, -1, new_mbx); + + if (old >=3D 0) + d.mbx_idx =3D old; + else + d.mbx_idx =3D new_mbx; + + return d; +} + +/** + * cmh_tm_txq_completion_notify() - Wake TM thread after RH completion + * + * Wakes the TM thread after the Response Handler completes a + * transaction. This unblocks the TM if it is waiting for a free MBX + * slot. The generation counter bump ensures the wait_event predicate + * evaluates to true on the next check. + */ +void cmh_tm_txq_completion_notify(void) +{ + atomic_inc(&txq_completion_gen); + wake_up_interruptible(&tm.cmq_waitq); +} + +/* Mailbox Selection */ + +/* + * Select a mailbox with at least @slots_needed free slots (round-robin). + * Returns mailbox index, or -EAGAIN if no mailbox qualifies. + * + * Note: the free-slot check here is advisory -- actual slot availability + * is enforced by the ring arithmetic under dispatch_lock in submit_vcq(). + * A TOCTOU gap exists between this check and the subsequent slot write, + * but it is safe: the worst case is a spurious -EAGAIN / backoff, never + * a ring overcommit. + */ +static int select_mailbox(u32 slots_needed) +{ + u32 count =3D tm.cfg->mbx_count; + u32 start =3D tm.next_mbx; + bool fits_anywhere =3D false; + u32 i; + + for (i =3D 0; i < count; i++) { + u32 idx =3D (start + i) % count; + struct cmh_mbx_config *m =3D &tm.cfg->mailboxes[idx]; + + if (cmh_rh_mbx_is_wedged(idx)) + continue; + + /* Could this ring ever hold the request, once drained? */ + if ((1U << m->slots_log2) >=3D slots_needed) + fits_anywhere =3D true; + + if (mbx_free_slots(m) >=3D slots_needed) { + tm.next_mbx =3D (idx + 1) % count; + return (int)idx; + } + cmh_stat_inc_mbx_queue_full(idx); + } + + /* + * No mailbox had room now. If no (non-wedged) mailbox's ring is + * even large enough to ever hold the request, it can never be + * placed -- report a permanent error so the caller fails it rather + * than spinning forever (head-of-line block). + */ + return fits_anywhere ? -EAGAIN : -EMSGSIZE; +} + +/* + * Resolve the target mailbox for a command message. + * + * If the message has a pinned MBX and it has enough free slots, use it. + * Otherwise fall back to round-robin selection. Returns mailbox index, + * or -EAGAIN when no MBX has enough free slots or all are wedged. + */ +static int resolve_mbx(struct command_msg *msg) +{ + u32 slots =3D msg->num_vcqs > 0 ? msg->num_vcqs : 1; + + if (msg->target_mbx >=3D 0 && + (u32)msg->target_mbx < tm.cfg->mbx_count) { + struct cmh_mbx_config *m =3D + &tm.cfg->mailboxes[msg->target_mbx]; + + if (cmh_rh_mbx_is_wedged((u32)msg->target_mbx)) + return -EAGAIN; + if ((1U << m->slots_log2) < slots) + return -EMSGSIZE; /* never fits this pinned ring */ + if (mbx_free_slots(m) >=3D slots) + return msg->target_mbx; + return -EAGAIN; /* pinned MBX full, retry */ + } + + return select_mailbox(slots); +} + +/* VCQ Submission */ + +/* + * Serialise all writes to R_MBX_QUEUE_TAIL. submit_vcq()'s doorbell + * advances TAIL while the response handler's cmh_tm_poke_tail() + * re-writes the current TAIL to wake the eSW. The latter is a + * read-modify-write running from softirq/timer context on another CPU, + * so without this lock it could read a stale TAIL and roll back a + * concurrent doorbell, dropping submitted VCQs. + */ +static DEFINE_SPINLOCK(cmh_mbx_tail_lock); + +/** + * cmh_tm_poke_tail() - Re-write R_MBX_QUEUE_TAIL to wake the eSW + * @base: Mailbox register base + * + * Writing the current TAIL back is a queue no-op but generates a fresh + * SIC interrupt edge so the eSW wakes from WFI. Serialised against the + * submit_vcq() doorbell so the read-modify-write cannot roll back a + * concurrent submission. Safe from softirq/timer context. + */ +void cmh_tm_poke_tail(void __iomem *base) +{ + unsigned long flags; + u32 tail; + + spin_lock_irqsave(&cmh_mbx_tail_lock, flags); + tail =3D cmh_reg_read32(base, R_MBX_QUEUE_TAIL); + cmh_reg_write32(tail, base, R_MBX_QUEUE_TAIL); + spin_unlock_irqrestore(&cmh_mbx_tail_lock, flags); +} + +/* + * Write VCQ(s) into consecutive DMA slots and ring the doorbell. + * + * A command_msg may carry one or more VCQs (num_vcqs field). For a + * multi-VCQ message the flat vcq_data array contains N VCQs laid out + * contiguously, each starting with its own header whose cmds field + * gives that VCQ's entry count. All VCQs are written to consecutive + * MBX slots and tracked by a single transaction_obj. + * + * Returns 0 on success, negative errno on failure. + */ +static int submit_vcq(struct command_msg *msg, u32 mbx_idx) +{ + struct cmh_mbx_config *mbx =3D &tm.cfg->mailboxes[mbx_idx]; + struct cmh_mbx_txq *txq =3D &tm.txqs[mbx_idx]; + struct transaction_obj *txn; + const struct vcq_cmd *cmds =3D msg->vcq_data; + u32 num_vcqs =3D msg->num_vcqs > 0 ? msg->num_vcqs : 1; + u32 tail, stride_bytes, offset =3D 0; + unsigned long flags; + u32 v; + + mutex_lock(&txq->dispatch_lock); + + /* Read current tail (first VCQ ID) */ + tail =3D cmh_reg_read32(mbx->reg_base, R_MBX_QUEUE_TAIL); + stride_bytes =3D 1U << mbx->stride_log2; + + /* Allocate transaction tracking object */ + txn =3D kzalloc_obj(*txn, GFP_KERNEL); + if (!txn) { + mutex_unlock(&txq->dispatch_lock); + return -ENOMEM; + } + + /* Write each VCQ into a consecutive DMA slot */ + for (v =3D 0; v < num_vcqs; v++) { + u32 vcq_cmds, copy_size; + void *slot; + + /* + * For single-VCQ messages (backward compat) use the + * msg-level vcq_count. For multi-VCQ, parse the per-VCQ + * header to find each VCQ's command count. + */ + if (num_vcqs =3D=3D 1) { + vcq_cmds =3D msg->vcq_count; + } else { + const struct vcq_hdr *hdr; + + /* Bound the per-VCQ header read within vcq_data. */ + if (offset >=3D msg->vcq_count) { + dev_err(cmh_dev(), + "tm: multi-VCQ %u offset OOB\n", v); + mutex_unlock(&txq->dispatch_lock); + kfree(txn); + return -EINVAL; + } + hdr =3D (const struct vcq_hdr *)&cmds[offset].hwc; + vcq_cmds =3D hdr->cmds; + } + + copy_size =3D vcq_cmds * sizeof(struct vcq_cmd); + if (copy_size > stride_bytes) { + dev_err(cmh_dev(), "tm: VCQ %u too large (%u bytes > stride %u)\n", + v, copy_size, stride_bytes); + mutex_unlock(&txq->dispatch_lock); + kfree(txn); + return -EMSGSIZE; + } + + if (vcq_cmds < MIN_VCQ_CMDS || vcq_cmds > MAX_VCQ_CMDS) { + dev_err(cmh_dev(), "tm: invalid vcq_count %u (range %u..%u)\n", + vcq_cmds, MIN_VCQ_CMDS, MAX_VCQ_CMDS); + mutex_unlock(&txq->dispatch_lock); + kfree(txn); + return -EINVAL; + } + + /* Bound the VCQ body within vcq_data. */ + if (vcq_cmds > msg->vcq_count - offset) { + dev_err(cmh_dev(), + "tm: VCQ %u body exceeds vcq_data\n", v); + mutex_unlock(&txq->dispatch_lock); + kfree(txn); + return -EINVAL; + } + + /* Copy pre-built VCQ into DMA slot */ + slot =3D mbx_slot_ptr(mbx, tail + v); + cmh_dma_write(slot, &cmds[offset], copy_size); + + /* Zero remaining slot bytes to avoid stale data */ + if (copy_size < stride_bytes) + cmh_dma_zero((u8 *)slot + copy_size, + stride_bytes - copy_size); + + offset +=3D vcq_cmds; + } + + /* Ensure VCQ data is visible in memory before advancing tail */ + wmb(); + /* FPGA: confirm DRAM accepted writes before SIC doorbell (cross-slave) */ + cmh_dma_fence(mbx_slot_ptr(mbx, tail + num_vcqs - 1)); + + /* Fill in transaction spanning all VCQs */ + txn->first_vcq_id =3D tail; + /* Expose the HW VCQ id so a sync timeout can check HEAD. */ + WRITE_ONCE(msg->first_vcq_id, tail); + txn->last_vcq_id =3D tail + num_vcqs - 1; + txn->mailbox_idx =3D mbx_idx; + txn->command_id =3D msg->command_id; + txn->error_code =3D 0; + txn->complete =3D msg->complete; + txn->completion_data =3D msg->completion_data; + atomic_set(&txn->state, TXN_INFLIGHT); + timer_setup(&txn->timeout_timer, txn_timeout_fn, 0); + INIT_LIST_HEAD(&txn->list); + + /* + * Set refcount: 2 if a per-txn timer will be armed (one ref for + * the TXQ owner that pops it, one for the timer callback), or 1 + * if no timer (sync paths, or async_timeout_ms =3D=3D 0). + */ + if (msg->timeout_jiffies) + refcount_set(&txn->refs, 2); + else + refcount_set(&txn->refs, 1); + + /* Enqueue transaction under spinlock */ + spin_lock_irqsave(&txq->lock, flags); + list_add_tail(&txn->list, &txq->head); + txq->depth++; + spin_unlock_irqrestore(&txq->lock, flags); + + /* + * Arm the per-request timeout BEFORE the doorbell (async only): once + * the doorbell rings, HW can complete and the TXQ can pop and free the + * txn concurrently, so arming afterwards races that free and can leave + * timer_delete() in cmh_txn_finish() unable to drop the timer ref. + */ + if (msg->timeout_jiffies) + mod_timer(&txn->timeout_timer, + jiffies + msg->timeout_jiffies); + + /* + * Ring doorbell: advance tail by number of VCQs submitted. + * Serialise against cmh_tm_poke_tail() so a concurrent re-poke + * cannot roll TAIL back over this advance. + */ + spin_lock_irqsave(&cmh_mbx_tail_lock, flags); + cmh_reg_write32(tail + num_vcqs, mbx->reg_base, R_MBX_QUEUE_TAIL); + spin_unlock_irqrestore(&cmh_mbx_tail_lock, flags); + + mutex_unlock(&txq->dispatch_lock); + + cmh_stat_record_vcq_submit(mbx_idx, num_vcqs, txq->depth); + + dev_dbg(cmh_dev(), "tm: submitted %u vcq(s) id=3D%u..%u to mbx[%u] tail_n= ow=3D%u\n", + num_vcqs, tail, tail + num_vcqs - 1, mbx_idx, + tail + num_vcqs); + + return 0; +} + +/* TM Thread */ + +static int cmh_tm_thread(void *data) +{ + struct command_msg *msg, *bl_promoted; + unsigned long flags; + int mbx_idx, ret; + + while (!kthread_should_stop()) { + /* Wait for work or stop signal */ + wait_event_interruptible(tm.cmq_waitq, + !list_empty(&tm.cmq) || kthread_should_stop()); + + if (kthread_should_stop()) + break; + + /* Dequeue one command message */ + spin_lock_irqsave(&tm.cmq_lock, flags); + if (list_empty(&tm.cmq)) { + spin_unlock_irqrestore(&tm.cmq_lock, flags); + continue; + } + msg =3D list_first_entry(&tm.cmq, struct command_msg, list); + list_del_init(&msg->list); + cmq_depth--; + + /* + * Promote one backlogged request into the CMQ now that + * there is room. The -EINPROGRESS notification is deferred + * until after cmq_lock is dropped (below). + */ + bl_promoted =3D NULL; + if (!list_empty(&tm.backlog)) { + bl_promoted =3D list_first_entry(&tm.backlog, + struct command_msg, list); + list_move_tail(&bl_promoted->list, &tm.cmq); + tm.backlog_depth--; + cmq_depth++; + cmh_stat_record_cmq_post(cmq_depth); + } + + spin_unlock_irqrestore(&tm.cmq_lock, flags); + + /* + * Signal -EINPROGRESS for the promoted backlog request with + * cmq_lock dropped: a consumer that resubmits from this + * callback would otherwise re-enter cmq_lock and self-deadlock. + * The promoted msg stays on the CMQ (only this thread dequeues + * it) until a later iteration, so the -EINPROGRESS still + * precedes its final completion. + */ + if (bl_promoted && bl_promoted->complete) + bl_promoted->complete(bl_promoted->completion_data, + -EINPROGRESS); + + /* Select a mailbox: pinned or round-robin */ + mbx_idx =3D resolve_mbx(msg); + + if (mbx_idx =3D=3D -EMSGSIZE) { + /* + * The request needs more ring slots than any mailbox + * can ever provide. Fail it rather than re-queuing: + * retrying would spin forever and head-of-line block + * every other request. + */ + dev_err(cmh_dev(), + "tm: cmd=3D0x%08x needs %u VCQs, exceeds ring capacity -- rejecting\n", + msg->command_id, + msg->num_vcqs > 0 ? msg->num_vcqs : 1); + if (msg->complete) + msg->complete(msg->completion_data, -EMSGSIZE); + command_msg_put(msg); + continue; + } + + if (mbx_idx < 0) { + /* + * Queue full -- re-enqueue at front and wait. + * + * Sleep on cmq_waitq with a short timeout. The RH + * calls cmh_tm_txq_completion_notify() after each + * completed transaction, which bumps the generation + * counter and wakes us immediately. The timeout is + * a safety net for missed wakeups. + */ + int gen =3D atomic_read(&txq_completion_gen); + unsigned long tmo; + + spin_lock_irqsave(&tm.cmq_lock, flags); + list_add(&msg->list, &tm.cmq); + cmq_depth++; + spin_unlock_irqrestore(&tm.cmq_lock, flags); + + tmo =3D usecs_to_jiffies(CMH_TM_BACKOFF_MAX_US); + wait_event_interruptible_timeout(tm.cmq_waitq, + kthread_should_stop() || + atomic_read(&txq_completion_gen) !=3D gen, + tmo ?: 1); + cmh_stat_inc_tm_backoff(); + continue; + } + + /* Submit VCQ to selected mailbox */ + WRITE_ONCE(msg->actual_mbx, mbx_idx); + ret =3D submit_vcq(msg, mbx_idx); + if (ret && msg->complete) + msg->complete(msg->completion_data, ret); + command_msg_put(msg); + } + + return 0; +} + +/* Public Interface */ + +/** + * cmh_tm_init() - Initialize the Transaction Manager subsystem + * @cfg: Hardware configuration describing mailboxes and core types + * + * Allocates per-mailbox transaction queues, applies core-type + * configuration, and starts the TM kthread. + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_tm_init(struct cmh_config *cfg) +{ + u32 i, j; + + tm.cfg =3D cfg; + tm.next_mbx =3D 0; + cmq_depth =3D 0; + + cmh_tm_affinity_reset(); + + /* Apply per-core-type config from DT child nodes */ + for (i =3D 0; i < CMH_NUM_CORE_TYPES; i++) { + struct cmh_core_type_cfg *src =3D &cfg->core_types[i]; + struct core_type_info *ct =3D &core_types[i]; + + ct->num_instances =3D src->num_instances; + for (j =3D 0; j < src->num_instances; j++) { + ct->instances[j].core_id =3D src->core_ids[j]; + if (src->mbx[j] >=3D 0) + atomic_set(&ct->instances[j].mbx_idx, + src->mbx[j]); + } + } + + /* Initialize CMQ and backlog */ + INIT_LIST_HEAD(&tm.cmq); + INIT_LIST_HEAD(&tm.backlog); + tm.backlog_depth =3D 0; + spin_lock_init(&tm.cmq_lock); + init_waitqueue_head(&tm.cmq_waitq); + + /* Allocate per-mailbox transaction queues */ + tm.txqs =3D kcalloc(cfg->mbx_count, sizeof(*tm.txqs), GFP_KERNEL); + if (!tm.txqs) + return -ENOMEM; + + for (i =3D 0; i < cfg->mbx_count; i++) { + INIT_LIST_HEAD(&tm.txqs[i].head); + spin_lock_init(&tm.txqs[i].lock); + mutex_init(&tm.txqs[i].dispatch_lock); + tm.txqs[i].depth =3D 0; + } + + /* Start TM thread */ + tm.thread =3D kthread_run(cmh_tm_thread, NULL, "cmh_tm"); + if (IS_ERR(tm.thread)) { + int ret =3D PTR_ERR(tm.thread); + + dev_err(cmh_dev(), "tm: failed to start thread (rc=3D%d)\n", ret); + tm.thread =3D NULL; + kfree(tm.txqs); + tm.txqs =3D NULL; + return ret; + } + + WRITE_ONCE(tm.running, true); + + return 0; +} + +/* + * cmh_tm_stop_and_drain_cmq() - Stop TM thread and drain CMQ/backlog + * + * Shared preamble for cmh_tm_cleanup() and cmh_tm_quiesce(): stops the + * kthread, marks the TM as not running, then splices the CMQ and backlog + * to local lists and cancels every pending command_msg outside the lock. + */ +static void cmh_tm_stop_and_drain_cmq(void) +{ + struct command_msg *msg; + unsigned long flags; + + if (tm.thread) { + kthread_stop(tm.thread); + tm.thread =3D NULL; + } + WRITE_ONCE(tm.running, false); + + /* + * Pop each queued command under cmq_lock and cancel it with the lock + * dropped. list_del_init() fully detaches the node before we release + * the lock, so a concurrent cmh_tm_try_cancel_command() either removes + * the node before we reach it or sees list_empty() and backs off -- it + * can no longer list_del a node that this drain is walking on a private + * list. Completion callbacks run with cmq_lock dropped (they may + * resubmit or take sleeping locks). + */ + for (;;) { + spin_lock_irqsave(&tm.cmq_lock, flags); + msg =3D list_first_entry_or_null(&tm.cmq, struct command_msg, + list); + if (msg) { + list_del_init(&msg->list); + cmq_depth--; + } + spin_unlock_irqrestore(&tm.cmq_lock, flags); + if (!msg) + break; + if (msg->complete) + msg->complete(msg->completion_data, -ECANCELED); + command_msg_put(msg); + } + + for (;;) { + spin_lock_irqsave(&tm.cmq_lock, flags); + msg =3D list_first_entry_or_null(&tm.backlog, struct command_msg, + list); + if (msg) { + list_del_init(&msg->list); + tm.backlog_depth--; + } + spin_unlock_irqrestore(&tm.cmq_lock, flags); + if (!msg) + break; + if (msg->complete) + msg->complete(msg->completion_data, -ECANCELED); + command_msg_put(msg); + } +} + +/** + * cmh_tm_cleanup() - Tear down the Transaction Manager subsystem + * + * Stops the TM kthread, drains the CMQ, backlog, and all per-mailbox + * transaction queues, notifying waiters with -ECANCELED or -ETIMEDOUT. + * Frees all TM-owned resources. + * + * The eSW/hardware exposes no "stop engines" or global-abort primitive: + * MBX_COMMAND_FLUSH discards VCQs still queued in a mailbox but does not + * abort a command the engine is already executing. Teardown therefore + * force-completes any residual in-flight transaction on the driver side + * rather than waiting on hardware. This is ordered safely: the caller + * (cmh_remove) invokes cmh_rh_cleanup() first, which cancels the + * watchdog, masks the MBX interrupts and frees the IRQ handlers, so the + * Response Handler can no longer deliver a completion for a transaction + * this drain is finishing -- the drain owns every remaining transaction + * exclusively. Consumers are expected to have quiesced by this point + * (the crypto core blocks module unload until all transforms are freed, + * and the synchronous /dev/cmh_mgmt ioctls complete before returning), + * so the residual queue normally holds only already-abandoned entries + * such as timed-out transactions. + */ +void cmh_tm_cleanup(void) +{ + struct transaction_obj *txn, *tmp_txn; + unsigned long flags; + u32 i; + + cmh_tm_stop_and_drain_cmq(); + + /* Drain per-mailbox transaction queues */ + if (tm.txqs) { + for (i =3D 0; i < tm.cfg->mbx_count; i++) { + LIST_HEAD(drain); + int old; + + spin_lock_irqsave(&tm.txqs[i].lock, flags); + list_splice_init(&tm.txqs[i].head, &drain); + tm.txqs[i].depth =3D 0; + spin_unlock_irqrestore(&tm.txqs[i].lock, flags); + + list_for_each_entry_safe(txn, tmp_txn, &drain, list) { + list_del(&txn->list); + + if (timer_delete_sync(&txn->timeout_timer)) + txn_put(txn); + + old =3D atomic_cmpxchg(&txn->state, + TXN_INFLIGHT, + TXN_COMPLETE); + if (txn->complete) { + if (old =3D=3D TXN_INFLIGHT) + txn->complete(txn->completion_data, + -ECANCELED); + else if (old =3D=3D TXN_TIMED_OUT) + txn->complete(txn->completion_data, + -ETIMEDOUT); + } + + txn_put(txn); + } + } + kfree(tm.txqs); + tm.txqs =3D NULL; + } +} + +/* + * Default drain timeout for suspend/quiesce (milliseconds). + * Covers all symmetric + PKE operations. PQC callers (SLH-DSA sign + * at up to 120 s) should complete before system suspend is requested. + */ +static unsigned int drain_timeout_ms =3D 10000; + +/** + * cmh_tm_quiesce() - Quiesce the TM for suspend or shutdown + * + * Stops the TM kthread, drains the CMQ and backlog, then waits up to + * drain_timeout_ms for in-flight transactions to complete via the + * Response Handler. Any remaining transactions after the deadline + * are force-cancelled. + */ +void cmh_tm_quiesce(void) +{ + struct transaction_obj *txn, *tmp_txn; + unsigned long deadline; + unsigned long flags; + u32 i; + bool drained =3D true; + + cmh_tm_stop_and_drain_cmq(); + + /* Wait for in-flight TXQ transactions to complete via RH */ + if (!tm.txqs) + return; + + deadline =3D jiffies + msecs_to_jiffies(drain_timeout_ms); + do { + drained =3D true; + for (i =3D 0; i < tm.cfg->mbx_count; i++) { + if (READ_ONCE(tm.txqs[i].depth)) { + drained =3D false; + break; + } + } + if (drained) + break; + usleep_range(1000, 2000); + } while (time_before(jiffies, deadline)); + + if (!drained) { + dev_warn(cmh_dev(), + "tm: quiesce drain timeout (%u ms), cancelling remaining transactions\= n", + drain_timeout_ms); + /* + * The drain timed out, so the RH is no longer making + * progress. Quiesce it first -- cmh_rh_suspend() stops the + * watchdog, masks the MBX interrupts and synchronize_irq()s + * out any in-flight handler -- so the force-cancel below owns + * every remaining transaction exclusively. Without this a + * live RH could concurrently finish/txn_put a transaction we + * are cancelling here (suspend runs cmh_tm_quiesce() before + * cmh_rh_suspend()). cmh_rh_suspend() is idempotent, so the + * caller's later call is a no-op. + */ + cmh_rh_suspend(tm.cfg); + + for (i =3D 0; i < tm.cfg->mbx_count; i++) { + LIST_HEAD(drain); + int old; + + spin_lock_irqsave(&tm.txqs[i].lock, flags); + list_splice_init(&tm.txqs[i].head, &drain); + tm.txqs[i].depth =3D 0; + spin_unlock_irqrestore(&tm.txqs[i].lock, flags); + + list_for_each_entry_safe(txn, tmp_txn, &drain, list) { + list_del(&txn->list); + + if (timer_delete_sync(&txn->timeout_timer)) + txn_put(txn); + + old =3D atomic_cmpxchg(&txn->state, + TXN_INFLIGHT, + TXN_COMPLETE); + if (txn->complete) { + if (old =3D=3D TXN_INFLIGHT) + txn->complete(txn->completion_data, + -ECANCELED); + else if (old =3D=3D TXN_TIMED_OUT) + txn->complete(txn->completion_data, + -ETIMEDOUT); + } + + txn_put(txn); + } + } + } +} + +/** + * cmh_tm_resume() - Resume the TM after suspend + * + * Restarts the TM kthread after a prior cmh_tm_quiesce(). + * + * Return: 0 on success, negative errno if kthread creation fails. + */ +int cmh_tm_resume(void) +{ + if (tm.thread || !tm.cfg) + return 0; + + tm.thread =3D kthread_run(cmh_tm_thread, NULL, "cmh_tm"); + if (IS_ERR(tm.thread)) { + int ret =3D PTR_ERR(tm.thread); + + dev_err(cmh_dev(), "tm: resume kthread_run failed (%d)\n", + ret); + tm.thread =3D NULL; + return ret; + } + WRITE_ONCE(tm.running, true); + return 0; +} + +/** + * cmh_tm_try_cancel_command() - Cancel a queued command message + * @msg: Command message to cancel + * + * Attempts to remove @msg from the CMQ before the TM thread dequeues + * it. Must be called while @msg is still valid (before the caller's + * stack frame that owns it is freed). + * + * Return: true if @msg was removed, false if already consumed by TM. + */ +bool cmh_tm_try_cancel_command(struct command_msg *msg) +{ + unsigned long flags; + bool cancelled =3D false; + + spin_lock_irqsave(&tm.cmq_lock, flags); + if (!list_empty(&msg->list)) { + list_del_init(&msg->list); + cmq_depth--; + cancelled =3D true; + } + spin_unlock_irqrestore(&tm.cmq_lock, flags); + + return cancelled; +} + +/** + * cmh_tm_post_command() - Post a command message to the CMQ + * @msg: Pre-built command message to enqueue + * + * Enqueues @msg on the Command Message Queue and wakes the TM thread. + * If the CMQ is full, the message may be placed on the backlog queue + * (returning -EBUSY) if @msg->backlog_ok is set, or rejected with + * -EAGAIN. + * + * Return: 0 on success, -EBUSY if backlogged, -EAGAIN if full, + * -ENODEV if TM is not running. + */ +int cmh_tm_post_command(struct command_msg *msg) +{ + unsigned long flags; + + spin_lock_irqsave(&tm.cmq_lock, flags); + /* + * Re-check tm.running under cmq_lock. cmh_tm_stop_and_drain_cmq() + * clears it and then splices the CMQ under this same lock, so a post + * that wins the lock after the drain sees !running and bails, while + * one that wins before is caught by the splice -- neither leaks a + * message onto a queue nobody will service. + */ + if (!READ_ONCE(tm.running)) { + spin_unlock_irqrestore(&tm.cmq_lock, flags); + return -ENODEV; + } + if (cmq_depth >=3D cmq_max_depth) { + if (msg->backlog_ok && + tm.backlog_depth < backlog_max_depth) { + list_add_tail(&msg->list, &tm.backlog); + tm.backlog_depth++; + spin_unlock_irqrestore(&tm.cmq_lock, flags); + return -EBUSY; + } + spin_unlock_irqrestore(&tm.cmq_lock, flags); + cmh_stat_inc_cmq_eagain(); + return -EAGAIN; + } + INIT_LIST_HEAD(&msg->list); + list_add_tail(&msg->list, &tm.cmq); + cmq_depth++; + cmh_stat_record_cmq_post(cmq_depth); + spin_unlock_irqrestore(&tm.cmq_lock, flags); + + wake_up_interruptible(&tm.cmq_waitq); + return 0; +} + +/* Synchronous Submit (refcounted completion + timeout) */ + +/* + * Heap-allocated sync context with refcounting. + * + * The completion callback may fire after the waiter has timed out and + * returned (e.g. during cmh_tm_cleanup on rmmod). If the struct lived + * on the waiter's stack, the callback would touch freed memory -- + * triggering a "BUG: spinlock bad magic" on the completion's spinlock. + * + * Two references are held: one by the waiter, one by the callback. + * Whichever runs last frees the struct. + */ +struct cmh_sync_ctx { + struct completion done; + int error; + refcount_t refs; /* 2: waiter + callback */ + + /* Optional orphan cleanup -- called when the last ref drops after + * the waiter abandoned an in-flight VCQ (noabort path). Lets the + * caller defer DMA-buffer cleanup until the eSW finishes writing. + */ + void (*orphan_cb)(void *data); + void *orphan_data; +}; + +static void cmh_sync_ctx_put(struct cmh_sync_ctx *ctx) +{ + if (refcount_dec_and_test(&ctx->refs)) { + if (ctx->orphan_cb) + ctx->orphan_cb(ctx->orphan_data); + kfree(ctx); + } +} + +static void cmh_sync_complete(void *data, int error) +{ + struct cmh_sync_ctx *ctx =3D data; + + ctx->error =3D error; + complete(&ctx->done); + cmh_sync_ctx_put(ctx); +} + +/* + * Default VCQ completion timeout (milliseconds), tunable via debugfs + * config/vcq_timeout_ms. Only affects the default timeout used by cmh_tm= _submit_sync() + * and cmh_tm_submit_sync_mbx(); callers that pass an explicit timeout_hz + * (e.g. RSA keygen) are not affected. + */ +static unsigned int vcq_timeout_ms =3D 2000; + +/* + * Extended timeout for slow crypto operations: RSA keygen, PQC + * keygen/sign/verify. Tunable via debugfs config/slow_op_timeout_ms. + */ +static unsigned int slow_op_timeout_ms =3D 300000; + +/** + * cmh_tm_submit_sync_tmo() - Synchronous VCQ submit with timeout + * @vcq_cmds: Array of pre-built VCQ command entries + * @vcq_count: Total number of entries in @vcq_cmds + * @num_vcqs: Number of VCQs packed in @vcq_cmds + * @target_mbx: Pinned mailbox index, or -1 for round-robin + * @timeout_hz: Completion timeout in jiffies + * + * Posts a VCQ command to the TM, waits for completion up to + * @timeout_hz. On timeout, issues MBX_COMMAND_ABORT if the VCQ is + * already in-flight and waits for the RH to drain the aborted + * transaction before returning. Consequently, once this returns + * -ETIMEDOUT the HW is no longer accessing the caller's DMA buffers, + * so the caller may unmap/free them without orphaning -- there is no + * post-return DMA to race. (The sole exception is the wedged-HW + * residual documented at the 5 s abort-timeout branch below.) Must be + * called from process context. + * + * Return: 0 on success, -ETIMEDOUT, or negative errno. + */ +int cmh_tm_submit_sync_tmo(struct vcq_cmd *vcq_cmds, u32 vcq_count, + u32 num_vcqs, s32 target_mbx, + unsigned long timeout_hz) +{ + struct cmh_sync_ctx *sync; + struct command_msg *msg; + unsigned long left; + int ret; + + /* + * This path sleeps (GFP_KERNEL allocations + wait_for_completion) + * and is not safe from atomic / non-sleepable contexts. All + * current callers run in process context (crypto API userspace or + * ioctl), so this is never violated today. Catch it loudly if + * a future caller gets this wrong. + */ + WARN_ON_ONCE(!in_task()); + + sync =3D kzalloc_obj(*sync, GFP_KERNEL); + if (!sync) + return -ENOMEM; + + msg =3D kzalloc_obj(*msg, GFP_KERNEL); + if (!msg) { + kfree(sync); + return -ENOMEM; + } + + init_completion(&sync->done); + sync->error =3D 0; + refcount_set(&sync->refs, 2); /* waiter + callback */ + + /* + * Heap-copy the caller's VCQ array so the msg owns its data. + * This decouples VCQ lifetime from the caller's stack frame, + * which matters when the TM thread backs off (resolve_mbx + * returns -1) and re-enqueues the msg after the caller's + * wait_for_completion_timeout expires. + */ + msg->vcq_data =3D kmemdup(vcq_cmds, vcq_count * sizeof(*vcq_cmds), + GFP_KERNEL); + if (!msg->vcq_data) { + kfree(msg); + kfree(sync); + return -ENOMEM; + } + + INIT_LIST_HEAD(&msg->list); + if (WARN_ON_ONCE(vcq_count < MIN_VCQ_CMDS)) { + ret =3D -EINVAL; + goto err_free; + } + msg->command_id =3D vcq_cmds[1].id; /* first real command's ID */ + msg->vcq_count =3D vcq_count; + msg->num_vcqs =3D num_vcqs; + msg->target_mbx =3D target_mbx; + msg->actual_mbx =3D -1; + msg->first_vcq_id =3D U32_MAX; + msg->complete =3D cmh_sync_complete; + msg->completion_data =3D sync; + refcount_set(&msg->refs, 2); /* waiter + TM subsystem */ + + ret =3D cmh_tm_post_command(msg); + if (ret) { +err_free: + kfree(msg->vcq_data); + kfree(msg); + kfree(sync); /* callback will never fire */ + return ret; + } + + dev_dbg(cmh_dev(), "tm: submit_sync posted cmd 0x%08x, waiting...\n", + msg->command_id); + + left =3D wait_for_completion_timeout(&sync->done, timeout_hz); + if (!left) { + dev_err(cmh_dev(), + "tm: submit_sync timeout (%lums) cmd=3D0x%08x\n", + timeout_hz * 1000 / HZ, msg->command_id); + if (cmh_tm_try_cancel_command(msg)) { + /* + * Msg was still queued -- TM never saw it. + * Drop the callback ref (no txn will fire it) + * and free msg directly (sole owner). + */ + cmh_sync_ctx_put(sync); /* no txn -> drop cb ref */ + cmh_sync_ctx_put(sync); /* drop waiter ref */ + command_msg_put(msg); /* matches refcount_set(2) */ + command_msg_put(msg); + } else { + /* + * TM has dequeued msg and the VCQ is in-flight. + * Issue MBX_COMMAND_ABORT to force-stop the VCQ; + * the RH will fire MBX_ERROR_IRQ, complete the + * transaction with -EIO, and issue RESTART. + * + * cmh_rh_abort_mbx() serialises the write under + * rh_process_lock, preventing clobber of a + * concurrent RESTART/FLUSH from the watchdog. + */ + s32 abrt_mbx =3D READ_ONCE(msg->actual_mbx); + + /* + * Only ABORT if our VCQ is the one currently at the + * mailbox HEAD. ABORT stops whatever executes at + * HEAD, so aborting while a different (older, + * possibly long-running) transaction is at HEAD would + * kill an unrelated request. If ours is queued behind + * it, the HW has not started ours yet -- leave it to + * the completion wait / watchdog rather than abort + * someone else. + */ + if (abrt_mbx >=3D 0 && + (u32)abrt_mbx < tm.cfg->mbx_count) { + void __iomem *abase =3D + tm.cfg->mailboxes[abrt_mbx].reg_base; + u32 head =3D cmh_reg_read32(abase, + R_MBX_QUEUE_HEAD); + + if (READ_ONCE(msg->first_vcq_id) =3D=3D head) { + dev_warn(cmh_dev(), + "tm: aborting mbx[%d] cmd=3D0x%08x\n", + abrt_mbx, msg->command_id); + cmh_rh_abort_mbx((u32)abrt_mbx); + } + } + + /* + * Wait for the RH completion (ABORT triggers + * MBX_ERROR_IRQ within microseconds). Fixed + * 5 s ceiling -- not configurable because if + * ABORT doesn't complete in this window the + * HW is wedged and more waiting won't help. + */ + left =3D wait_for_completion_timeout(&sync->done, + 5 * HZ); + if (!left) { + /* + * ABORT did not complete within 5 s -- the HW + * is wedged. Unlike cmh_tm_submit_sync_noabort, + * this path sets no orphan_cb, so the caller's + * DMA buffers are NOT orphaned here: the caller + * frees them on the -ETIMEDOUT below. If the + * eSW later recovers and DMAs into them that is + * a use-after-free -- an accepted residual for a + * path that cannot be reached unless ABORT + * itself hangs for 5 s (HW already wedged). + */ + dev_err(cmh_dev(), + "tm: abort timeout (5s) cmd=3D0x%08x - HW wedged\n", + msg->command_id); + } + cmh_sync_ctx_put(sync); /* drop waiter ref */ + command_msg_put(msg); /* drop waiter ref on msg */ + } + return -ETIMEDOUT; + } + + ret =3D sync->error; + cmh_sync_ctx_put(sync); /* drop waiter ref */ + command_msg_put(msg); /* drop waiter ref on msg */ + return ret; +} + +/** + * cmh_tm_submit_sync_mbx() - Synchronous VCQ submit on a target MBX + * @vcq_cmds: Array of pre-built VCQ command entries + * @vcq_count: Total number of entries in @vcq_cmds + * @num_vcqs: Number of VCQs packed in @vcq_cmds + * @target_mbx: Pinned mailbox index, or -1 for round-robin + * + * Convenience wrapper around cmh_tm_submit_sync_tmo() using the + * default vcq_timeout_ms module parameter. + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_tm_submit_sync_mbx(struct vcq_cmd *vcq_cmds, u32 vcq_count, + u32 num_vcqs, s32 target_mbx) +{ + return cmh_tm_submit_sync_tmo(vcq_cmds, vcq_count, num_vcqs, + target_mbx, + msecs_to_jiffies(vcq_timeout_ms)); +} + +/** + * cmh_tm_async_timeout_jiffies() - Default async per-request timeout + * + * Return: Timeout in jiffies from the async_timeout_ms module param, + * or 0 if async timeouts are disabled. + */ +unsigned long cmh_tm_async_timeout_jiffies(void) +{ + return async_timeout_ms ? msecs_to_jiffies(async_timeout_ms) : 0; +} + +/** + * cmh_tm_slow_op_timeout_jiffies() - Timeout for slow crypto ops + * + * Returns the extended timeout used for RSA keygen, PQC keygen/sign, + * and similar long-running operations. + * + * Return: Timeout in jiffies from the slow_op_timeout_ms module param. + */ +unsigned long cmh_tm_slow_op_timeout_jiffies(void) +{ + return msecs_to_jiffies(slow_op_timeout_ms); +} + +/** + * cmh_tm_submit_async() - Asynchronous VCQ submission + * @vcq_cmds: Array of pre-built VCQ command entries + * @vcq_count: Total number of entries in @vcq_cmds + * @num_vcqs: Number of VCQs packed in @vcq_cmds + * @target_mbx: Pinned mailbox index, or -1 for round-robin + * @callback: Completion callback (see context note below) + * @callback_data: Opaque data passed to @callback + * @backlog_ok: Allow backlogging if CMQ is full + * @timeout_jiffies: Per-request timeout (0 =3D no timeout) + * + * Builds a command_msg, heap-copies the VCQ data, and posts it to the + * CMQ via cmh_tm_post_command(). + * + * Callback context guarantee: + * The @callback may be invoked from one of: + * - RH threaded IRQ handler (process context, BH disabled) + * - RH watchdog timer (softirq / timer context) + * - TM kthread if submit_vcq() fails post-dequeue + * - cmh_tm_cleanup()/cmh_tm_quiesce() during drain (process context) + * It is NEVER invoked from hardirq context. + * + * Because the watchdog path runs from timer softirq, callbacks + * MUST be safe in atomic/softirq context: no mutex, no GFP_KERNEL, + * no sleeping locks. crypto_request_complete() is safe (documented + * callable from any context). kfree_sensitive() and + * scatterwalk_map_and_copy() are also safe (non-sleeping). + * Callers must not assume thread affinity (callback may run on any CPU). + * + * Unlike the _sync variants, this function: + * - Does NOT allocate a cmh_sync_ctx or wait for completion + * - Uses GFP_ATOMIC for internal allocations because the crypto API + * may call ->encrypt/->decrypt/->hash_final from softirq context + * (e.g. network stack via IPsec/TLS); GFP_KERNEL would deadlock. + * + * The command_msg is single-owner (refcount 1) -- the TM subsystem + * owns it after post and frees it after dispatching to the HW. + * + * DMA buffer ownership: the caller transfers ownership to the callback + * on return of 0 or -EBUSY. On any other return, the caller must + * clean up DMA buffers itself -- the callback will never fire. + * + * Return: 0 on successful post, -EBUSY if backlogged, -ENOMEM, + * -EINVAL, -EAGAIN, or -ENODEV on failure. + */ +int cmh_tm_submit_async(struct vcq_cmd *vcq_cmds, u32 vcq_count, + u32 num_vcqs, s32 target_mbx, + cmh_completion_fn callback, void *callback_data, + bool backlog_ok, unsigned long timeout_jiffies) +{ + struct command_msg *msg; + int ret; + + msg =3D kzalloc_obj(*msg, GFP_ATOMIC); + if (!msg) + return -ENOMEM; + + msg->vcq_data =3D kmemdup(vcq_cmds, + array_size(vcq_count, sizeof(*vcq_cmds)), + GFP_ATOMIC); + if (!msg->vcq_data) { + kfree(msg); + return -ENOMEM; + } + + INIT_LIST_HEAD(&msg->list); + if (WARN_ON_ONCE(vcq_count < MIN_VCQ_CMDS)) { + kfree(msg->vcq_data); + kfree(msg); + return -EINVAL; + } + msg->command_id =3D vcq_cmds[1].id; + msg->vcq_count =3D vcq_count; + msg->num_vcqs =3D num_vcqs; + msg->target_mbx =3D target_mbx; + msg->actual_mbx =3D -1; + msg->complete =3D callback; + msg->completion_data =3D callback_data; + msg->backlog_ok =3D backlog_ok; + msg->timeout_jiffies =3D timeout_jiffies; + refcount_set(&msg->refs, 1); /* sole owner: TM subsystem */ + + ret =3D cmh_tm_post_command(msg); + if (ret && ret !=3D -EBUSY) { + kfree(msg->vcq_data); + kfree(msg); + } + return ret; +} + +/** + * cmh_tm_submit_sync_noabort() - Sync submit without MBX abort on timeout + * @vcq_cmds: Array of pre-built VCQ command entries + * @vcq_count: Total number of entries in @vcq_cmds + * @num_vcqs: Number of VCQs packed in @vcq_cmds + * @timeout_hz: Completion timeout in jiffies + * @orphan_cb: Optional cleanup callback for abandoned DMA buffers + * @orphan_data: Opaque data passed to @orphan_cb + * + * On timeout, if the command was still queued it is cancelled and + * -EAGAIN is returned (caller may free all resources). If the VCQ is + * already in-flight, the waiter drops its refs and returns -EINPROGRESS + * -- the RH callback will fire when the eSW finishes the VCQ and free + * the sync_ctx / msg via the refcount mechanism. + * + * @orphan_cb is invoked when the last ref on the sync_ctx drops after + * the waiter abandoned an in-flight VCQ, allowing the caller to defer + * DMA-buffer cleanup until the eSW finishes writing. + * + * This prevents a short-timeout command (e.g. DRBG GENERATE from the + * hwrng kthread) from aborting the entire MBX and killing unrelated + * long-running operations (e.g. SLH-DSA sign at 120 s). + * + * Return: 0 on success, -EAGAIN if cancelled from queue, + * -EINPROGRESS if left in-flight, or negative errno. + */ +int cmh_tm_submit_sync_noabort(struct vcq_cmd *vcq_cmds, u32 vcq_count, + u32 num_vcqs, unsigned long timeout_hz, + void (*orphan_cb)(void *), void *orphan_data) +{ + struct cmh_sync_ctx *sync; + struct command_msg *msg; + unsigned long left; + int ret; + + WARN_ON_ONCE(!in_task()); + + sync =3D kzalloc_obj(*sync, GFP_KERNEL); + if (!sync) + return -ENOMEM; + + msg =3D kzalloc_obj(*msg, GFP_KERNEL); + if (!msg) { + kfree(sync); + return -ENOMEM; + } + + init_completion(&sync->done); + sync->error =3D 0; + refcount_set(&sync->refs, 2); + + INIT_LIST_HEAD(&msg->list); + if (WARN_ON_ONCE(vcq_count < MIN_VCQ_CMDS)) { + kfree(msg); + kfree(sync); + return -EINVAL; + } + msg->command_id =3D vcq_cmds[1].id; + msg->vcq_data =3D kmemdup(vcq_cmds, vcq_count * sizeof(*vcq_cmds), + GFP_KERNEL); + if (!msg->vcq_data) { + kfree(msg); + kfree(sync); + return -ENOMEM; + } + msg->vcq_count =3D vcq_count; + msg->num_vcqs =3D num_vcqs; + msg->target_mbx =3D -1; + msg->actual_mbx =3D -1; + msg->complete =3D cmh_sync_complete; + msg->completion_data =3D sync; + refcount_set(&msg->refs, 2); + + ret =3D cmh_tm_post_command(msg); + if (ret) { + kfree(msg->vcq_data); + kfree(msg); + kfree(sync); + return ret; + } + + left =3D wait_for_completion_timeout(&sync->done, timeout_hz); + if (!left) { + if (cmh_tm_try_cancel_command(msg)) { + /* Still queued -- TM never saw it, clean up fully */ + cmh_sync_ctx_put(sync); /* drop cb ref */ + cmh_sync_ctx_put(sync); /* drop waiter ref */ + command_msg_put(msg); /* matches refcount_set(2) */ + command_msg_put(msg); + return -EAGAIN; + } + + /* + * In-flight: skip ABORT. Transfer orphan cleanup + * ownership to sync_ctx -- the RH callback will + * eventually complete this VCQ, and when the last + * ref drops, orphan_cb frees any DMA buffers the + * eSW was still writing to. + */ + dev_dbg_ratelimited(cmh_dev(), + "tm: noabort timeout (%lums) cmd=3D0x%08x, leaving in-flight\n", + timeout_hz * 1000 / HZ, + msg->command_id); + sync->orphan_cb =3D orphan_cb; + sync->orphan_data =3D orphan_data; + cmh_sync_ctx_put(sync); + command_msg_put(msg); + return -EINPROGRESS; + } + + ret =3D sync->error; + cmh_sync_ctx_put(sync); + command_msg_put(msg); + return ret; +} + +/** + * cmh_tm_submit_sync() - Synchronous VCQ submit with default timeout + * @vcq_cmds: Array of pre-built VCQ command entries + * @vcq_count: Total number of entries in @vcq_cmds + * @num_vcqs: Number of VCQs packed in @vcq_cmds + * + * Convenience wrapper: submits via round-robin MBX selection with the + * default vcq_timeout_ms. + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_tm_submit_sync(struct vcq_cmd *vcq_cmds, u32 vcq_count, + u32 num_vcqs) +{ + return cmh_tm_submit_sync_mbx(vcq_cmds, vcq_count, num_vcqs, -1); +} + +#define MBX_FLUSH_TIMEOUT_MS 1000 +#define MBX_FLUSH_POLL_MIN_US 10 +#define MBX_FLUSH_POLL_MAX_US 50 + +/** + * cmh_tm_flush_mbx() - Issue MBX_COMMAND_FLUSH and wait for completion + * @mbx_idx: Mailbox index to flush + * + * Resets the eSW child mailbox state: clears the VCQ command queue, + * resets head/tail, and -- critically -- resets the child temp stack + * via mbx_hdr_init() (sets hdr->temp back to &cmds[MAX_VCQ_CMDS]). + * + * Why this is needed: + * KIC derivation commands that output to SYS_REF_TEMP allocate on the + * per-MBX child temp LIFO stack (mbx_alloc_temp, each costing + * ROUND_UP(len,4)+56 bytes). These allocations persist across VCQ + * completions because mbx_vcq_done() does NOT reset the temp stack. + * Without an explicit flush, sequential KIC-TEMP ioctls exhaust the + * ~960-byte temp area and subsequent derives fail with ENOMEM. + * + * What is NOT affected: + * KIC HW keys, datastore objects, DRBG state -- these survive the + * flush. Only the queue pointers and temp stack are reset. + * + * Concurrency: + * Acquires the per-MBX dispatch_lock mutex to serialise with VCQ + * dispatch in submit_vcq(). This prevents the flush from resetting + * head/tail while the TM kthread is writing a VCQ to a DMA slot on + * the same MBX. The eSW clears R_MBX_COMMAND to zero once the flush + * completes. + * + * Return: 0 on success, -EINVAL, -ENODEV, -EBUSY, or -ETIMEDOUT. + */ +int cmh_tm_flush_mbx(s32 mbx_idx) +{ + struct cmh_mbx_config *mbx; + struct cmh_mbx_txq *txq; + struct transaction_obj *txn; + void __iomem *base; + u32 reg; + int ret; + + if (!tm.cfg || mbx_idx < 0 || (u32)mbx_idx >=3D tm.cfg->mbx_count) + return -EINVAL; + + mbx =3D &tm.cfg->mailboxes[mbx_idx]; + base =3D mbx->reg_base; + if (!base) + return -ENODEV; + + txq =3D &tm.txqs[mbx_idx]; + mutex_lock(&txq->dispatch_lock); + + /* Ensure no command is already pending */ + if (cmh_reg_read32(base, R_MBX_COMMAND) !=3D 0) { + mutex_unlock(&txq->dispatch_lock); + return -EBUSY; + } + + cmh_reg_write32(MBX_COMMAND_FLUSH, base, R_MBX_COMMAND); + + /* Poll until eSW clears the command register */ + ret =3D read_poll_timeout(cmh_reg_read32, reg, reg =3D=3D 0, + MBX_FLUSH_POLL_MIN_US, + MBX_FLUSH_TIMEOUT_MS * 1000, + true, base, R_MBX_COMMAND); + if (ret) + dev_err(cmh_dev(), "mbx %u flush timeout (cmd=3D0x%08x)\n", + mbx->instance, + cmh_reg_read32(base, R_MBX_COMMAND)); + + /* + * FLUSH discarded every queued VCQ (HEAD=3DTAIL). The caller + * contract is that the mailbox is idle here, but defensively + * complete any straggler transaction with -ECANCELED: otherwise + * the response handler would see HEAD jump to TAIL and report the + * discarded request as a successful completion. dispatch_lock is + * held, so no new submit_vcq() can race this drain. + */ + while ((txn =3D cmh_tm_pop_transaction((u32)mbx_idx)) !=3D NULL) { + cmh_txn_finish(txn, -ECANCELED); + cmh_tm_txq_completion_notify(); + } + + mutex_unlock(&txq->dispatch_lock); + return ret; +} + +/** + * cmh_vcq_pack_and_submit() - Pack payload into VCQs and submit sync + * @payload: Array of VCQ command entries (without headers) + * @count: Number of entries in @payload + * @packed: Caller-provided output buffer for packed VCQ data + * @max_packed: Size of @packed buffer in vcq_cmd entries + * @target_mbx: Pinned mailbox index, or -1 for round-robin + * + * Splits @payload into VCQ-sized chunks, prepends headers, and submits + * synchronously. + * + * Return: 0 on success, -EMSGSIZE if @packed is too small, or + * negative errno from submit. + */ +int cmh_vcq_pack_and_submit(const struct vcq_cmd *payload, u32 count, + struct vcq_cmd *packed, u32 max_packed, + s32 target_mbx) +{ + u32 max_per_vcq =3D cmh_tm_max_cmds_per_vcq(); + u32 max_payload_per =3D max_per_vcq - 1; + u32 num_vcqs =3D 0, total =3D 0, i =3D 0; + + while (i < count) { + u32 chunk =3D min_t(u32, count - i, max_payload_per); + u32 vcq_cmds =3D chunk + 1; + + if (total + vcq_cmds > max_packed) + return -EMSGSIZE; + + vcq_set_header(&packed[total], vcq_cmds); + memcpy(&packed[total + 1], &payload[i], + chunk * sizeof(struct vcq_cmd)); + + total +=3D vcq_cmds; + i +=3D chunk; + num_vcqs++; + } + + return cmh_tm_submit_sync_mbx(packed, total, num_vcqs, target_mbx); +} + +/** + * cmh_vcq_pack_and_submit_async() - Pack payload and submit async + * @payload: Array of VCQ command entries (without headers) + * @count: Number of entries in @payload + * @packed: Caller-provided output buffer for packed VCQ data + * @max_packed: Size of @packed buffer in vcq_cmd entries + * @target_mbx: Pinned mailbox index, or -1 for round-robin + * @callback: Completion callback + * @callback_data: Opaque data passed to @callback + * @backlog_ok: Allow backlogging if CMQ is full + * @timeout_jiffies: Per-request timeout (0 =3D no timeout) + * + * Asynchronous variant of cmh_vcq_pack_and_submit(). Splits @payload + * into VCQ-sized chunks, prepends headers, and submits via + * cmh_tm_submit_async(). + * + * Return: 0 on success, -EBUSY if backlogged, -EMSGSIZE if @packed + * is too small, or negative errno from submit. + */ +int cmh_vcq_pack_and_submit_async(const struct vcq_cmd *payload, u32 count, + struct vcq_cmd *packed, u32 max_packed, + s32 target_mbx, + cmh_completion_fn callback, + void *callback_data, + bool backlog_ok, + unsigned long timeout_jiffies) +{ + u32 max_per_vcq =3D cmh_tm_max_cmds_per_vcq(); + u32 max_payload_per =3D max_per_vcq - 1; + u32 num_vcqs =3D 0, total =3D 0, i =3D 0; + + while (i < count) { + u32 chunk =3D min_t(u32, count - i, max_payload_per); + u32 vcq_cmds =3D chunk + 1; + + if (total + vcq_cmds > max_packed) + return -EMSGSIZE; + + vcq_set_header(&packed[total], vcq_cmds); + memcpy(&packed[total + 1], &payload[i], + chunk * sizeof(struct vcq_cmd)); + + total +=3D vcq_cmds; + i +=3D chunk; + num_vcqs++; + } + + return cmh_tm_submit_async(packed, total, num_vcqs, target_mbx, + callback, callback_data, backlog_ok, + timeout_jiffies); +} + +/** + * cmh_tm_peek_transaction() - Peek at the head of a mailbox TXQ + * @mbx_idx: Mailbox index to inspect + * + * Returns a pointer to the oldest in-flight transaction without + * removing it from the queue. The caller must not free the returned + * object. + * + * Return: Pointer to the head transaction, or NULL if empty. + */ +struct transaction_obj *cmh_tm_peek_transaction(u32 mbx_idx) +{ + struct cmh_mbx_txq *txq; + struct transaction_obj *txn =3D NULL; + unsigned long flags; + + if (!tm.txqs || mbx_idx >=3D tm.cfg->mbx_count) + return NULL; + + txq =3D &tm.txqs[mbx_idx]; + + spin_lock_irqsave(&txq->lock, flags); + if (!list_empty(&txq->head)) + txn =3D list_first_entry(&txq->head, struct transaction_obj, + list); + spin_unlock_irqrestore(&txq->lock, flags); + + return txn; +} + +/** + * cmh_tm_pop_transaction() - Remove and return the head of a MBX TXQ + * @mbx_idx: Mailbox index to pop from + * + * Dequeues the oldest in-flight transaction from the per-mailbox + * transaction queue. The caller takes ownership and must eventually + * call cmh_txn_finish() or txn_put(). + * + * Return: Pointer to the dequeued transaction, or NULL if empty. + */ +struct transaction_obj *cmh_tm_pop_transaction(u32 mbx_idx) +{ + struct cmh_mbx_txq *txq; + struct transaction_obj *txn; + unsigned long flags; + + if (!tm.txqs || mbx_idx >=3D tm.cfg->mbx_count) + return NULL; + + txq =3D &tm.txqs[mbx_idx]; + + spin_lock_irqsave(&txq->lock, flags); + if (list_empty(&txq->head)) { + spin_unlock_irqrestore(&txq->lock, flags); + return NULL; + } + txn =3D list_first_entry(&txq->head, struct transaction_obj, list); + list_del_init(&txn->list); + txq->depth--; + spin_unlock_irqrestore(&txq->lock, flags); + + return txn; +} + +/* -- debugfs timeout accessors ----------------------------------------- = */ + +#ifdef CONFIG_CRYPTO_DEV_CMH_DEBUG +/** + * cmh_tm_timeout_async_ptr() - Return pointer to async_timeout_ms for deb= ugfs + * + * Return: pointer to the static async_timeout_ms variable. + */ +unsigned int *cmh_tm_timeout_async_ptr(void) { return &async_timeout_ms= ; } + +/** + * cmh_tm_timeout_vcq_ptr() - Return pointer to vcq_timeout_ms for debugfs + * + * Return: pointer to the static vcq_timeout_ms variable. + */ +unsigned int *cmh_tm_timeout_vcq_ptr(void) { return &vcq_timeout_ms; } + +/** + * cmh_tm_timeout_slow_op_ptr() - Return pointer to slow_op_timeout_ms for= debugfs + * + * Return: pointer to the static slow_op_timeout_ms variable. + */ +unsigned int *cmh_tm_timeout_slow_op_ptr(void) { return &slow_op_timeout_= ms; } + +/** + * cmh_tm_timeout_drain_ptr() - Return pointer to drain_timeout_ms for deb= ugfs + * + * Return: pointer to the static drain_timeout_ms variable. + */ +unsigned int *cmh_tm_timeout_drain_ptr(void) { return &drain_timeout_ms= ; } + +/** + * cmh_tm_cmq_max_depth_ptr() - Return pointer to cmq_max_depth for debugfs + * + * Return: pointer to the static cmq_max_depth variable. + */ +unsigned int *cmh_tm_cmq_max_depth_ptr(void) { return &cmq_max_depth; } + +/** + * cmh_tm_backlog_max_depth_ptr() - Return pointer to backlog_max_depth fo= r debugfs + * + * Return: pointer to the static backlog_max_depth variable. + */ +unsigned int *cmh_tm_backlog_max_depth_ptr(void) { return &backlog_max_dep= th; } +#endif diff --git a/drivers/crypto/cmh/include/cmh.h b/drivers/crypto/cmh/include/= cmh.h new file mode 100644 index 000000000000..97a7bc764d9b --- /dev/null +++ b/drivers/crypto/cmh/include/cmh.h @@ -0,0 +1,24 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Top-level Device Structure + */ + +#ifndef CMH_H +#define CMH_H + +#include + +#include "cmh_config.h" + +/** + * struct cmh_device - Top-level driver state for a CMH hardware instance + * @config: Hardware configuration (core mappings, MBX layout, feature fla= gs) + * @dev: Platform or parent device used for DMA and logging + */ +struct cmh_device { + struct cmh_config config; + struct device *dev; +}; + +#endif /* CMH_H */ diff --git a/drivers/crypto/cmh/include/cmh_aes_abi.h b/drivers/crypto/cmh/= include/cmh_aes_abi.h new file mode 100644 index 000000000000..0b876dd67773 --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_aes_abi.h @@ -0,0 +1,98 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- AES Core ABI Definitions + * + * Kernel-side definitions for the CMH AES ABI. + * All constants and layouts derived from the CMH eSW ABI. + */ + +#ifndef CMH_AES_ABI_H +#define CMH_AES_ABI_H + +#include + +/* AES Block Size */ + +#define CMH_AES_BLOCK_SIZE 16U +#define CMH_AES_IV_SIZE 16U + +/* AES Modes (per CMH AES ABI) */ + +#define AES_MODE_ECB 1U +#define AES_MODE_CBC 2U +#define AES_MODE_CTR 3U +#define AES_MODE_CFB 4U +#define AES_MODE_GCM 5U +#define AES_MODE_CMAC 6U +#define AES_MODE_CCM 7U +#define AES_MODE_XTS 8U + +/* AES Operations (per CMH AES ABI) */ + +#define AES_OP_DECRYPT 1U +#define AES_OP_ENCRYPT 2U + +/* AES Command IDs */ + +#define AES_CMD_INIT 0x01U +#define AES_CMD_AAD_UPDATE 0x02U +#define AES_CMD_AAD_FINAL 0x03U +#define AES_CMD_UPDATE 0x04U +#define AES_CMD_FINAL 0x05U +#define AES_CMD_SCATTERGATHER 0x06U +#define AES_CMD_CCM_INIT 0x0AU +#define AES_CMD_AAD_FINAL_AUTH 0x0EU + +/* AES Command Structures */ + +struct aes_cmd_init { + u64 key; /* datastore reference for the key */ + u64 iv; /* DMA address of the IV (or nonce in CCM) */ + u32 keylen; /* key length in bytes */ + u32 ivlen; /* IV length in bytes (0..16) */ + u32 mode; /* AES mode (AES_MODE_*) */ + u32 op; /* AES operation (AES_OP_*) */ + u32 aadlen; /* AAD length or 0 */ + u32 iolen; /* plaintext/ciphertext length */ + u32 taglen; /* tag length or 0 */ + u32 xts_offset; /* XTS block index j; 0 for the skcipher path */ +}; + +struct aes_cmd_aad_final { + u64 data; /* DMA address of AAD data */ + u32 datalen; /* AAD data length */ +}; + +struct aes_cmd_aad_final_auth { + u64 data; /* DMA address of final AAD data */ + u32 datalen; /* final AAD data length */ + u64 tag; /* DMA address of tag */ + u32 taglen; /* tag length */ +}; + +struct aes_cmd_update { + u64 input; /* DMA address of input data */ + u64 output; /* DMA address of output data */ + u32 iolen; /* input/output data length */ +}; + +struct aes_cmd_final { + u64 input; /* DMA address of last input data */ + u64 output; /* DMA address of last output data */ + u64 tag; /* DMA address of tag (AEAD only) */ + u32 iolen; /* last input/output data length */ + u32 taglen; /* tag length (AEAD only) */ +}; + +/* AES Command Union */ + +union aes_cmd { + struct aes_cmd_init cmd_init; + struct aes_cmd_update cmd_update; + struct aes_cmd_final cmd_final; + struct aes_cmd_aad_final cmd_aad_final; + struct aes_cmd_aad_final_auth cmd_aad_final_auth; +}; + +#endif /* CMH_AES_ABI_H */ diff --git a/drivers/crypto/cmh/include/cmh_ccp_abi.h b/drivers/crypto/cmh/= include/cmh_ccp_abi.h new file mode 100644 index 000000000000..4e3eb9feaec9 --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_ccp_abi.h @@ -0,0 +1,108 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- CCP Core ABI Definitions + * + * Kernel-side definitions for the CMH CCP ABI. + * All constants and layouts derived from the CMH eSW ABI. + * + * The CCP core provides three modes: + * - ChaCha20 stream cipher (skcipher) + * - Poly1305 one-time authenticator (shash) + * - ChaCha20-Poly1305 AEAD (RFC 7539) + */ + +#ifndef CMH_CCP_ABI_H +#define CMH_CCP_ABI_H + +#include + +/* CCP Block Sizes */ + +#define CCP_CHACHA_BLOCK_SIZE 64U /* ChaCha20 block =3D 512 bits */ +#define CCP_POLY_BLOCK_SIZE 16U /* Poly1305 block =3D 128 bits */ +#define CCP_CTRNONCE_SIZE 16U /* 4-byte LE counter + 12-byte nonce */ +#define CCP_POLY_KEY_SIZE 16U /* r_key and s_key each 16 bytes */ +#define CCP_POLY_TAG_SIZE 16U /* Poly1305 tag =3D 128 bits */ +#define CCP_CHACHA_CTR_LEN 4U /* 32-bit counter */ + +/* CCP Operations (per CMH CCP ABI) */ + +#define CCP_OP_DECRYPT 1U +#define CCP_OP_ENCRYPT 2U + +/* CCP Command IDs */ + +#define CCP_CMD_CHACHA20_INIT 0x01U +#define CCP_CMD_POLY1305_INIT 0x02U +#define CCP_CMD_AEAD_INIT 0x03U +#define CCP_CMD_AAD_UPDATE 0x04U +#define CCP_CMD_AAD_FINAL 0x05U +#define CCP_CMD_UPDATE 0x06U +#define CCP_CMD_FINAL 0x07U +#define CCP_CMD_SCATTERGATHER 0x08U +/* CCP_CMD_FLUSH =3D VCQ_CMD_FLUSH (0xFF) -- defined in cmh_vcq.h */ + +/* CCP Command Structures */ + +struct ccp_cmd_chacha { + u64 key; /* datastore reference for the key */ + u64 ctrnonce; /* DMA address of the 16-byte counter+nonce */ + u32 keylen; /* key length: 16 or 32 bytes */ + u32 ctrnoncelen; /* always 16 */ + u32 ctrlen; /* counter length: 4 bytes */ + u32 op; /* CCP_OP_ENCRYPT or CCP_OP_DECRYPT */ +}; + +struct ccp_cmd_poly { + u64 rkey; /* datastore reference for the r key */ + u64 skey; /* datastore reference for the s key */ + u32 rkeylen; /* always 16 */ + u32 skeylen; /* always 16 */ +}; + +struct ccp_cmd_aead { + u64 key; /* datastore reference for the key */ + u64 ctrnonce; /* DMA address of the 16-byte counter+nonce */ + u32 keylen; /* key length: 32 bytes */ + u32 ctrnoncelen; /* always 16 */ + u32 op; /* CCP_OP_ENCRYPT or CCP_OP_DECRYPT */ +}; + +struct ccp_cmd_aad_update { + u64 aad; /* DMA address of AAD data */ + u32 aadlen; /* AAD length (must be multiple of 16) */ +}; + +struct ccp_cmd_aad_final { + u64 aad; /* DMA address of last AAD data */ + u32 aadlen; /* last AAD length (any size) */ +}; + +struct ccp_cmd_update { + u64 input; /* DMA address of input data */ + u64 output; /* DMA address of output data */ + u32 iolen; /* input/output length */ +}; + +struct ccp_cmd_final { + u64 input; /* DMA address of last input data */ + u64 output; /* DMA address of last output data */ + u64 tag; /* DMA address of the 16-byte tag */ + u32 iolen; /* last input/output data length */ + u32 taglen; /* tag length (always 16) */ +}; + +/* CCP Command Union */ + +union ccp_cmd { + struct ccp_cmd_chacha cmd_chacha; + struct ccp_cmd_poly cmd_poly; + struct ccp_cmd_aead cmd_aead; + struct ccp_cmd_aad_update cmd_aad_update; + struct ccp_cmd_aad_final cmd_aad_final; + struct ccp_cmd_update cmd_update; + struct ccp_cmd_final cmd_final; +}; + +#endif /* CMH_CCP_ABI_H */ diff --git a/drivers/crypto/cmh/include/cmh_config.h b/drivers/crypto/cmh/i= nclude/cmh_config.h new file mode 100644 index 000000000000..d3315b0736ba --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_config.h @@ -0,0 +1,105 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Configuration Structures and Defaults + */ + +#ifndef CMH_CONFIG_H +#define CMH_CONFIG_H + +#include +#include + +#include "cmh_registers.h" +#include "cmh_vcq.h" + +/* Limits */ + +/* + * Max mailboxes the driver manages simultaneously. The hardware address + * space supports CMH_MAX_MBX_INSTANCES (64) instance indices, but this + * compile-time constant caps how many the driver allocates DMA queues, + * IRQ slots, and per-transform cache entries for. To manage more + * mailboxes (up to the HW max), increase this value and rebuild the LKM + * -- it cannot be changed via module parameters at runtime. + */ +#define CMH_MAX_CONFIGURED_MBX 16 +#define CMH_MAX_CORE_INSTANCES 8 + +/* MBX setup parameter ranges (per CMH hardware specification) */ +#define CMH_MBX_SLOTS_LOG2_MIN 1 +#define CMH_MBX_SLOTS_LOG2_MAX 15 +#define CMH_MBX_STRIDE_LOG2_MIN 7 +#define CMH_MBX_STRIDE_LOG2_MAX 10 + +/* Default Configuration Values */ + +#define CMH_DEFAULT_MBX_COUNT 2 +#define CMH_DEFAULT_SLOTS_LOG2 6 /* 2^6 =3D 64 slots */ +#define CMH_DEFAULT_STRIDE_LOG2 9 /* 2^9 =3D 512 bytes per slot */ +#define CMH_DEFAULT_IRQ (-1) /* polling mode */ +#define CMH_FW_READY_TIMEOUT_MS 5000 /* built-in; raise here for a slower= eSW boot */ + +/* Per-Core-Type Instance Configuration */ + +struct cmh_core_type_cfg { + u32 num_instances; + u32 core_ids[CMH_MAX_CORE_INSTANCES]; + s32 mbx[CMH_MAX_CORE_INSTANCES]; /* -1 =3D auto-assign */ +}; + +/* Per-Mailbox Configuration */ + +struct cmh_mbx_config { + u32 instance; /* 0-based MBX instance index (0..63) */ + u32 slots_log2; /* log2(slot count), range 1..15 */ + u32 stride_log2; /* log2(bytes per slot), range 7..10 */ + int irq; /* per-mailbox virq, -1 if none (poll) */ + u32 cores[CMH_NUM_CORE_TYPES]; /* rambus,cores affinity IDs */ + u32 num_cores; /* number of entries in cores[] */ + u32 lock_val; /* MBX lock token (non-zero while held) */ + dma_addr_t dma_handle; /* DMA bus address from dma_alloc_coheren= t */ + void *virt_addr; /* kernel virtual address of MBXQ buffer = */ + size_t queue_size; /* total queue buffer size in bytes */ + void __iomem *reg_base; /* ioremap'd register base for this insta= nce */ +}; + +/* Global Device Configuration */ + +struct cmh_config { + phys_addr_t sic_base; + size_t sic_size; + void __iomem *sic_mapped; /* ioremap'd SIC region */ + u32 mbx_count; + struct cmh_mbx_config mailboxes[CMH_MAX_CONFIGURED_MBX]; + unsigned int fw_ready_timeout_ms; /* FW mission-mode timeo= ut */ + struct cmh_core_type_cfg core_types[CMH_NUM_CORE_TYPES]; +}; + +/* Module Parameter Interface */ + +struct platform_device; + +/** + * cmh_config_init() - Populate config from module params and device-tree + * @cfg: Configuration structure to fill + * @pdev: Platform device (for DT properties and IRQ lookup) + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_config_init(struct cmh_config *cfg, struct platform_device *pdev); + +/** + * cmh_config_discover_cores() - Enumerate crypto cores from the hardware + * @cfg: Configuration structure (SIC must already be mapped) + * + * Reads the SIC CORE_ENABLE register to determine which crypto core types + * the silicon build provides, populates cfg->core_types[], applies the + * per-mailbox rambus,cores affinity, and validates the result. Must be + * called after cfg->sic_mapped is valid (i.e. after the SIC ioremap). + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_config_discover_cores(struct cmh_config *cfg); + +#endif /* CMH_CONFIG_H */ diff --git a/drivers/crypto/cmh/include/cmh_debugfs.h b/drivers/crypto/cmh/= include/cmh_debugfs.h new file mode 100644 index 000000000000..d85158e23bcf --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_debugfs.h @@ -0,0 +1,90 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- debugfs Per-MBX and TM Counters + * + * Exposes diagnostic counters under /sys/kernel/debug/cmh/: + * + * mbxN/vcqs_submitted Total VCQs sent to MBX N + * mbxN/vcqs_completed Total completions received + * mbxN/vcqs_errors Total error completions + * mbxN/queue_full_count Times select_mailbox() skipped this MBX + * mbxN/max_queue_depth High-water mark of in-flight transactions + * + * tm/cmq_posts Total cmh_tm_post_command() calls + * tm/cmq_depth_max High-water mark of CMQ length + * tm/cmq_eagain_count Times CMQ was full (-EAGAIN) + * tm/backoff_count Times TM backed off (all MBX queues full) + * tm/async_timeout_count Async requests that timed out + * + * Counters are atomic64_t -- safe to read from any context. + * When CONFIG_CRYPTO_DEV_CMH_DEBUG is off, all functions become no-ops an= d the + * compiler eliminates the counter code entirely. + */ + +#ifndef CMH_DEBUGFS_H +#define CMH_DEBUGFS_H + +#include +#include + +/* Per-Mailbox Statistics */ + +struct cmh_mbx_stats { + atomic64_t vcqs_submitted; + atomic64_t vcqs_completed; + atomic64_t vcqs_errors; + atomic64_t queue_full_count; + atomic64_t max_queue_depth; +}; + +/* TM-Level Statistics */ + +struct cmh_tm_stats { + atomic64_t cmq_posts; + atomic64_t cmq_depth_max; + atomic64_t cmq_eagain_count; + atomic64_t backoff_count; + atomic64_t async_timeout_count; +}; + +/** + * cmh_stat_update_max() - Atomically update a high-water mark counter + * @counter: atomic64_t counter to update + * @val: New candidate value + * + * Updates @counter to @val if @val exceeds the current maximum. + * Lock-free via atomic cmpxchg loop. + */ +static inline void cmh_stat_update_max(atomic64_t *counter, s64 val) +{ + s64 cur; + + do { + cur =3D atomic64_read(counter); + if (val <=3D cur) + return; + } while (atomic64_cmpxchg(counter, cur, val) !=3D cur); +} + +/* Interface (stub when CONFIG_CRYPTO_DEV_CMH_DEBUG is off) */ + +struct cmh_config; + +#ifdef CONFIG_CRYPTO_DEV_CMH_DEBUG + +void cmh_debugfs_init(struct cmh_config *cfg); +void cmh_debugfs_cleanup(void); + +struct cmh_mbx_stats *cmh_debugfs_mbx_stats(u32 mbx_idx); +struct cmh_tm_stats *cmh_debugfs_tm_stats(void); + +#else /* !CONFIG_CRYPTO_DEV_CMH_DEBUG */ + +static inline void cmh_debugfs_init(struct cmh_config *c) { } +static inline void cmh_debugfs_cleanup(void) {} +static inline struct cmh_mbx_stats *cmh_debugfs_mbx_stats(u32 i) { return = NULL; } +static inline struct cmh_tm_stats *cmh_debugfs_tm_stats(void) { return = NULL; } + +#endif /* CONFIG_CRYPTO_DEV_CMH_DEBUG */ +#endif /* CMH_DEBUGFS_H */ diff --git a/drivers/crypto/cmh/include/cmh_dma.h b/drivers/crypto/cmh/incl= ude/cmh_dma.h new file mode 100644 index 000000000000..7dd0d8311785 --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_dma.h @@ -0,0 +1,219 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- DMA Interface + * + * Platform-independent DMA operations for the CMH crypto accelerator. + * All functions are implemented in cmh_dma.c (standard kernel DMA API). + * + * Alternate backends may be linked in place of cmh_dma.c for + * non-standard platforms. Such backends must implement the same + * symbol set and may use different allocation and mapping semantics + * (e.g. pool-based alloc/free instead of address translation). + */ + +#ifndef CMH_DMA_H +#define CMH_DMA_H + +#include +#include + +#include "cmh_vcq.h" + +struct platform_device; + +/** + * cmh_dma_init() - Initialize the DMA backend + * @pdev: Platform device (provides struct device for DMA ops) + * + * Called early in .probe(). The standard backend stores the device + * pointer; alternate backends may set up additional resources. + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_dma_init(struct platform_device *pdev); + +/** + * cmh_dma_cleanup() - Tear down the DMA backend + * + * Called in .remove() and error paths. Releases any resources + * allocated by cmh_dma_init(). + */ +void cmh_dma_cleanup(void); + +/** + * cmh_dev() - Global device accessor + * + * Returns the struct device * associated with the platform_driver instanc= e. + * Valid only between cmh_dma_init() and cmh_dma_cleanup(). + * + * Return: Platform device pointer, or NULL outside lifecycle. + */ +struct device *cmh_dev(void); + +/* Streaming DMA map / unmap (short-lived per-request buffers) */ + +dma_addr_t cmh_dma_map_single(void *buf, size_t size, + enum dma_data_direction dir); +void cmh_dma_unmap_single(dma_addr_t addr, size_t size, + enum dma_data_direction dir); + +/* + * Sync a DMA_FROM_DEVICE buffer so the CPU sees device-written data. + * + * Required before reading *buf when SWIOTLB bounce buffering is active + * (e.g. arm64 without IOMMU): the device writes to the bounce buffer, + * not the original allocation, so the CPU must sync before access. + * On architectures without bounce buffers (e.g. rv64) this is a no-op. + * + * Call between cmh_tm_submit_sync() and the first CPU read of the buffer, + * while the mapping is still live (before cmh_dma_unmap_single). + */ +void cmh_dma_sync_for_cpu(dma_addr_t addr, size_t size, + enum dma_data_direction dir); + +/* + * Sync a DMA_TO_DEVICE buffer so the device sees CPU-written data. + * + * Required after CPU writes to a mapped streaming buffer (e.g. SG + * descriptor arrays that need items_dma for .lli pointer calculation + * before content is written). Must be called before the device reads. + */ +void cmh_dma_sync_for_device(dma_addr_t addr, size_t size, + enum dma_data_direction dir); + +int cmh_dma_map_error(dma_addr_t addr); + +/* Coherent DMA alloc / free (long-lived MBX queue buffers) */ + +void *cmh_dma_alloc(size_t size, dma_addr_t *handle, gfp_t gfp); +void cmh_dma_free(size_t size, void *virt, dma_addr_t handle); + +/** + * cmh_dma_write() - Copy data into a DMA-allocated buffer + * @dst: Destination pointer (from cmh_dma_alloc) + * @src: Source kernel buffer + * @len: Number of bytes to copy + * + * Copies @len bytes from @src to @dst. @dst must have been obtained + * from cmh_dma_alloc(). Abstracted to allow platforms with non-standard + * DMA buffer access semantics. + */ +void cmh_dma_write(void *dst, const void *src, size_t len); + +/** + * cmh_dma_fence() - Fence preceding writes to DMA-allocated memory + * @ptr: Any pointer into the region that was written + * + * Ensures all preceding CPU writes to DMA memory are committed to the + * target memory controller before subsequent MMIO register writes. + * + * Required on FPGA platforms where DMA memory and device control + * registers reside on different AXI slaves -- a CPU-side wmb() only + * orders store dispatch, not arrival at the target. A read from the + * DMA memory slave forces the memory controller to serialize behind + * all preceding writes from this CPU before responding, guaranteeing + * the data is committed before the doorbell register write is issued. + * + * On standard DMA API platforms (cache-coherent), this is a no-op. + */ +void cmh_dma_fence(void *ptr); + +/** + * cmh_dma_zero() - Zero a DMA-allocated buffer + * @dst: Destination pointer (from cmh_dma_alloc) + * @len: Number of bytes to zero + */ +void cmh_dma_zero(void *dst, size_t len); + +/* + * CMH eSW scatter-gather chain -- built with proper DMA mappings. + * + * The CMH eSW DMAC walks a linked list of dma_scattergather_item + * descriptors. Each .src is the DMA address of an input buffer; + * each .lli is the DMA address of the next descriptor (0 =3D end). + * + * The descriptor array uses streaming DMA (kmalloc + dma_map_single) + * so that cmh_dma_free_sg() is safe from any context -- including + * BH-disabled completion callbacks where dma_free_coherent's + * vunmap() path would crash on non-coherent architectures. + */ + +/* Input descriptor for cmh_dma_build_sg() -- one per data buffer */ +struct cmh_dma_buf { + void *data; + u32 len; +}; + +/* Opaque handle returned by cmh_dma_build_sg(); pass to cmh_dma_free_sg()= */ +struct cmh_sg_map { + struct dma_scattergather_item *items; /* CPU virtual address */ + dma_addr_t items_dma; /* DMA address (pass to GATHER cmd) */ + size_t items_size; /* allocation size */ + u32 count; + struct { + dma_addr_t dma; + u32 len; + } bufs[]; /* per-entry source DMA handles */ +}; + +/** + * cmh_dma_build_sg() - Build a DMA-mapped CMH eSW SG chain + * @bufs: Array of kernel buffer descriptors (data pointer + length) + * @count: Number of entries in @bufs (must be > 0; returns NULL for 0) + * @gfp: Allocation flags (GFP_KERNEL or GFP_ATOMIC) + * + * Allocates a dma_scattergather_item chain using streaming DMA + * (kmalloc + dma_map_single), DMA-maps each source buffer, and + * links the descriptors. + * The returned cmh_sg_map->items_dma is the address to pass to + * vcq_add_hc_gather() (or any core's scatter-gather command). + * + * Caller contract: + * - Each bufs[i].data must point to DMA-mappable memory (kmalloc, + * page-allocated, or vmalloc with DMA support). Stack buffers + * are NOT safe. + * - Each bufs[i].len must be > 0. + * - The returned cmh_sg_map must remain alive (not freed) until + * the hardware completes the scatter-gather operation. Only then + * may cmh_dma_free_sg() be called. + * - There is no hardware-imposed limit on @count, but callers are + * responsible for bounding it to avoid excessive DMA mappings. + * In practice, hash uses <=3D 2 entries (partial + new data). + * + * Return: Opaque cmh_sg_map handle, or NULL on allocation/mapping failure. + */ +struct cmh_sg_map *cmh_dma_build_sg(const struct cmh_dma_buf *bufs, u32 co= unt, + gfp_t gfp); + +/** + * cmh_dma_free_sg() - Unmap all buffers and free the SG chain + * @sgm: Handle from cmh_dma_build_sg(), or NULL (no-op) + */ +void cmh_dma_free_sg(struct cmh_sg_map *sgm); + +/* + * Orphan-DMA context -- generic helper for the noabort submit path. + * + * When cmh_tm_submit_sync_noabort() times out with a VCQ still + * in-flight, the eSW will continue writing to DMA buffers after the + * caller returns. Callers wrap their DMA state in this struct and + * pass cmh_dma_orphan_free as the orphan_cb -- the RH callback frees + * the mapping + buffer when the VCQ eventually completes. + * + * Drain guarantee: cmh_tm_cleanup() calls timer_delete_sync() on each + * TXN timeout timer and splices all TXQ entries before invoking their + * completion callbacks. This ensures no orphan callback can race with + * or run after TM cleanup completes -- by that point every in-flight + * transaction has been force-completed and its orphan_cb invoked. + */ +struct cmh_dma_orphan { + void *buf; + dma_addr_t addr; + size_t len; + enum dma_data_direction dir; +}; + +void cmh_dma_orphan_free(void *data); + +#endif /* CMH_DMA_H */ diff --git a/drivers/crypto/cmh/include/cmh_drbg_abi.h b/drivers/crypto/cmh= /include/cmh_drbg_abi.h new file mode 100644 index 000000000000..d4cebfe83d4b --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_drbg_abi.h @@ -0,0 +1,67 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- DRBG Core ABI Definitions + * + * Kernel-side definitions for the CMH DRBG ABI. + * All constants and layouts derived from the CMH eSW ABI. + */ + +#ifndef CMH_DRBG_ABI_H +#define CMH_DRBG_ABI_H + +#include + +/* DRBG Commands */ + +#define DRBG_CMD_CONFIG 0x01U +#define DRBG_CMD_GENERATE 0x02U +#define DRBG_CMD_DATASTORE 0x03U +#define DRBG_CMD_RESET 0x04U + +/* DRBG Entropy Ratio (per CMH DRBG ABI) */ + +#define DRBG_ENTROPY_RATIO_ONE 0U +#define DRBG_ENTROPY_RATIO_ONE_HALF 1U +#define DRBG_ENTROPY_RATIO_ONE_THIRD 2U +#define DRBG_ENTROPY_RATIO_ONE_FOURTH 3U + +/* DRBG Security Strength (per CMH DRBG ABI) */ + +#define DRBG_SECURITY_STRENGTH_128 0x00U +#define DRBG_SECURITY_STRENGTH_256 0x10U + +/* DRBG Personalization Data Length */ + +#define DRBG_PADATA_LEN 16U + +/* DRBG Command Structures */ + +struct drbg_cmd_config { + u32 entropy_ratio; /* drbg_entropy_ratio value */ + u32 security_strength; /* drbg_security_strength value */ + u8 padata[DRBG_PADATA_LEN]; +}; + +struct drbg_cmd_generate { + u64 dst; /* DMA physical address for output */ + u32 len; /* requested output length in bytes */ + u8 padata[DRBG_PADATA_LEN]; +}; + +struct drbg_cmd_datastore { + u64 ref; /* datastore reference */ + u32 len; /* data length in bytes */ + u32 type; /* datastore type */ + u8 padata[DRBG_PADATA_LEN]; +}; + +/* DRBG Command Union */ + +union drbg_cmd { + struct drbg_cmd_config cmd_config; + struct drbg_cmd_generate cmd_generate; + struct drbg_cmd_datastore cmd_datastore; +}; + +#endif /* CMH_DRBG_ABI_H */ diff --git a/drivers/crypto/cmh/include/cmh_eac_abi.h b/drivers/crypto/cmh/= include/cmh_eac_abi.h new file mode 100644 index 000000000000..f0ebd3de1fb4 --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_eac_abi.h @@ -0,0 +1,44 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- EAC (Error and Alarm Controller) ABI Definitions + * + * Kernel-side definitions for the CMH EAC ABI. + * All constants and layouts derived from the CMH eSW ABI. + */ + +#ifndef CMH_EAC_ABI_H +#define CMH_EAC_ABI_H + +#include + +/* EAC Commands */ + +#define EAC_CMD_READ 0x01U + +/* EAC Read Response -- eSW writes this to the DMA destination buffer */ + +struct eac_read_rsp { + u64 mailbox_notification; /* bitmask: MBX that raised safety notif */ + u32 hw_error; /* bitmask: HWC that raised error */ + u32 hw_nmi; /* bitmask: HWC that raised NMI */ + u32 hw_panic; /* bitmask: HWC that raised HW panic */ + u32 safety_fatal; /* bitmask: HWC that raised fatal safety */ + u32 safety_notification; /* bitmask: HWC that raised safety notif */ + u32 sw_info0; /* eSW tracing information */ + u32 sw_info1; /* eSW tracing information */ + u32 sram_bank_errors[4]; /* correctable ECC error counts per bank */ +}; + +/* EAC Command Structures */ + +struct eac_cmd_read { + u64 dst; /* DMA destination for eac_read_rsp */ + u32 len; /* must be >=3D sizeof(struct eac_read_rsp) */ +}; + +union eac_cmd { + struct eac_cmd_read cmd_read; +}; + +#endif /* CMH_EAC_ABI_H */ diff --git a/drivers/crypto/cmh/include/cmh_hc_abi.h b/drivers/crypto/cmh/i= nclude/cmh_hc_abi.h new file mode 100644 index 000000000000..4e8c5ea3c69c --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_hc_abi.h @@ -0,0 +1,162 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Hash Core (HC) ABI Definitions + * + * Kernel-side definitions for the CMH HC (Hash Core) ABI. + * All constants and layouts derived from the CMH eSW ABI. + */ + +#ifndef CMH_HC_ABI_H +#define CMH_HC_ABI_H + +#include +#include + +/* HC Commands */ + +#define HC_CMD_INIT 0x01U +#define HC_CMD_HMAC 0x02U +#define HC_CMD_UPDATE 0x03U +#define HC_CMD_FINAL 0x04U +#define HC_CMD_UPDATE2D 0x05U +#define HC_CMD_SQUEEZE 0x07U +#define HC_CMD_GATHER 0x08U +#define HC_CMD_CSHAKE 0x09U +#define HC_CMD_KMAC 0x0AU +#define HC_CMD_SAVE 0x0BU +#define HC_CMD_RESTORE 0x0CU + +/* HC Algorithms (per CMH HC ABI) */ + +#define HC_ALGO_SHA2_224 1U +#define HC_ALGO_SHA2_256 2U +#define HC_ALGO_SHA2_384 3U +#define HC_ALGO_SHA2_512 4U +#define HC_ALGO_SHA3_224 5U +#define HC_ALGO_SHA3_256 6U +#define HC_ALGO_SHA3_384 7U +#define HC_ALGO_SHA3_512 8U +#define HC_ALGO_SHAKE128 9U +#define HC_ALGO_SHAKE256 10U + +/* HC Algo Flags */ + +#define HC_ALGO_FLAG_SCA_KEY BIT(18) /* SCA key in 2 shares */ +#define HC_ALGO_FLAG_SCA_OUT BIT(19) /* SCA output in 2 shares */ + +#define HC_ALGO_SET(flags, algo) (((flags) & 0xFF0000UL) | ((algo) & 0xFF= UL)) +#define HC_ALGO_GET(algo) ((algo) & 0xFFU) + +/* Hash Digest Sizes */ + +#define CMH_SHA224_DIGEST_SIZE 28U +#define CMH_SHA256_DIGEST_SIZE 32U +#define CMH_SHA384_DIGEST_SIZE 48U +#define CMH_SHA512_DIGEST_SIZE 64U + +/* SHA-3 digest sizes are the same as SHA-2 for matching output widths */ +#define CMH_SHA3_224_DIGEST_SIZE 28U +#define CMH_SHA3_256_DIGEST_SIZE 32U +#define CMH_SHA3_384_DIGEST_SIZE 48U +#define CMH_SHA3_512_DIGEST_SIZE 64U + +/* SHAKE default output lengths (fixed-output ahash registration) */ +#define CMH_SHAKE128_DIGEST_SIZE 32U /* 128-bit security -> 32 bytes */ +#define CMH_SHAKE256_DIGEST_SIZE 64U /* 256-bit security -> 64 bytes */ + +/* HC Context (for SAVE/RESTORE) */ + +#define HC_CONTEXT_WORDS 149U +#define HC_CONTEXT_SIZE (HC_CONTEXT_WORDS * 4 + 4) /* ctx[149] + = crc */ + +/* cSHAKE function name max length */ + +#define HC_CSHAKE_MAX_NAMELEN 36U + +/* + * Maximum customization string (S) length for cSHAKE / KMAC. + * + * S is packed as inline VCQ data after the CSHAKE/KMAC command slot. + * The worst-case VCQ layout (KMAC with raw key + GATHER) uses 5 fixed + * slots out of CMH_KMAC_MAX_PAYLOAD (9), leaving 4 inline slots. + * Each VCQ slot is 64 bytes, so the safe limit is 4 * 64 =3D 256 bytes. + */ +#define HC_CSHAKE_MAX_CUSTOMLEN 256U + +/* HC Command Structures */ + +struct hc_cmd_init { + u32 algo; /* hc_algo value, optionally ORed with HC_ALGO_FLAG_* */ +}; + +struct hc_cmd_hmac { + u64 key; /* datastore reference for HMAC key */ + u32 keylen; /* key length in bytes */ + u32 algo; /* hc_algo value */ +}; + +struct hc_cmd_update { + u64 input; /* DMA physical address of input data */ + u32 inlen; /* input data length in bytes */ +}; + +struct hc_cmd_final { + u64 digest; /* DMA physical address for output digest */ + u32 outlen; /* digest length in bytes */ +}; + +struct hc_cmd_update2d { + u64 input; /* DMA source address for input data */ + u64 output; /* DMA destination address for pass-through data */ + u32 iolen; /* input/pass-through data length in bytes */ +}; + +struct hc_cmd_gather { + u64 lista; /* DMA address of dma_scattergather_item chain */ + u32 sgcmd; /* HC sub-command: HC_CMD_UPDATE or HC_CMD_UPDATE2D */ +}; + +struct hc_cmd_cshake { + u64 custom; /* DMA address for the customization string */ + u32 customlen; /* length of the customization string */ + u32 algo; /* HC_ALGO_SHAKE128 or HC_ALGO_SHAKE256 */ + u32 namelen; /* length of the function name string */ + u8 name[HC_CSHAKE_MAX_NAMELEN]; /* function name string (inline) */ +}; + +struct hc_cmd_kmac { + u64 key; /* datastore reference for KMAC key */ + u64 custom; /* DMA address for the customization string */ + u32 keylen; /* key length in bytes */ + u32 customlen; /* length of the customization string */ + u32 algo; /* HC_ALGO_SHAKE128 or HC_ALGO_SHAKE256 */ + u32 outlen; /* requested output digest length */ +}; + +struct hc_cmd_save { + u64 output; /* DMA physical address for saved context */ + u32 outlen; /* must be HC_CONTEXT_SIZE */ +}; + +struct hc_cmd_restore { + u64 input; /* DMA physical address of saved context */ + u32 inlen; /* must be HC_CONTEXT_SIZE */ +}; + +/* HC Command Union */ + +union hc_cmd { + struct hc_cmd_init cmd_init; + struct hc_cmd_hmac cmd_hmac; + struct hc_cmd_cshake cmd_cshake; + struct hc_cmd_kmac cmd_kmac; + struct hc_cmd_update cmd_update; + struct hc_cmd_final cmd_final; + struct hc_cmd_update2d cmd_update2d; + struct hc_cmd_gather cmd_gather; + struct hc_cmd_save cmd_save; + struct hc_cmd_restore cmd_restore; +}; + +#endif /* CMH_HC_ABI_H */ diff --git a/drivers/crypto/cmh/include/cmh_hcq_abi.h b/drivers/crypto/cmh/= include/cmh_hcq_abi.h new file mode 100644 index 000000000000..b9fc2a80a408 --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_hcq_abi.h @@ -0,0 +1,221 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- HCQ Core ABI Definitions + * + * Kernel-side definitions for the CMH HCQ ABI. + * All constants and layouts derived from the CMH eSW ABI. + */ + +#ifndef CMH_HCQ_ABI_H +#define CMH_HCQ_ABI_H + +#include +#include + +/* VCQ layout: header + [SYS cmds] + HCQ_CMD + [sys_read] + flush */ +#define HCQ_VCQ_CMDS_MIN 3 /* header + cmd + flush */ +#define HCQ_VCQ_CMDS_MAX 6 /* keygen: hdr+new+write+cmd+read+flush */ + +/* HCQ Command IDs */ +#define HCQ_CMD_XMSS_VERIFY 0x03U +#define HCQ_CMD_LMS_VERIFY 0x04U +#define HCQ_CMD_SLHDSA_VERIFY_INTERNAL 0x05U +#define HCQ_CMD_SLHDSA_VERIFY 0x06U +#define HCQ_CMD_SLHDSA_VERIFY_PREHASH 0x07U +#define HCQ_CMD_SLHDSA_VERIFY_PREHASH_DIGEST 0x08U +#define HCQ_CMD_SLHDSA_KEYGEN 0x09U +#define HCQ_CMD_SLHDSA_SIGN_INTERNAL 0x10U +#define HCQ_CMD_SLHDSA_SIGN 0x11U +#define HCQ_CMD_SLHDSA_SIGN_PREHASH 0x12U +#define HCQ_CMD_SLHDSA_SIGN_PREHASH_DIGEST 0x13U +#define HCQ_CMD_SLHDSA_PUBGEN 0x14U + +/* SLH-DSA Parameter Set IDs */ +#define HCQ_SLHDSA_SHAKE_128S 1U +#define HCQ_SLHDSA_SHAKE_128F 2U +#define HCQ_SLHDSA_SHAKE_192S 3U +#define HCQ_SLHDSA_SHAKE_192F 4U +#define HCQ_SLHDSA_SHAKE_256S 5U +#define HCQ_SLHDSA_SHAKE_256F 6U +#define HCQ_SLHDSA_SHA2_128S 7U +#define HCQ_SLHDSA_SHA2_128F 8U +#define HCQ_SLHDSA_SHA2_192S 9U +#define HCQ_SLHDSA_SHA2_192F 10U +#define HCQ_SLHDSA_SHA2_256S 11U +#define HCQ_SLHDSA_SHA2_256F 12U +#define HCQ_SLHDSA_PARAM_MAX 12U + +/* SLH-DSA Prehash Algorithm IDs */ +#define HCQ_SLHDSA_PREHASH_SHA256 1U +#define HCQ_SLHDSA_PREHASH_SHA512 2U +#define HCQ_SLHDSA_PREHASH_SHAKE128 3U +#define HCQ_SLHDSA_PREHASH_SHAKE256 4U + +/* SLH-DSA size limits */ +#define SLHDSA_MAX_PK_SIZE 64U /* 2*n, n=3D32 */ +#define SLHDSA_MAX_SK_SIZE 128U /* 4*n, n=3D32 */ +#define SLHDSA_MAX_SEED_SIZE 96U /* 3*n, n=3D32 */ +#define SLHDSA_MAX_SIG_SIZE 49856U /* SHAKE-256f / SHA2-256f */ +#define SLHDSA_MAX_MSG_LEN 128U +#define SLHDSA_MAX_CTX_LEN 255U + +/* LMS/HSS size limits -- derived from eSW HCQ ABI constraints */ +#define LMS_MAX_PK_LEN 60U /* eSW public-key buffer */ +#define LMS_MAX_MSG_LEN 256U /* SHS_LMS_MESSAGE_LEN_MAX */ +#define LMS_MAX_SIG_LEN 13364U /* eSW signature buffer */ + +/* XMSS/XMSS-MT size limits -- derived from eSW HCQ ABI constraints */ +#define XMSS_MAX_PK_LEN 136U /* eSW public-key buffer */ +#define XMSS_MAX_MSG_LEN 64U /* SHS_XMSS_MESSAGE_LEN_MAX */ +#define XMSS_MAX_SIG_LEN 27688U /* eSW signature buffer */ + +/* SLH-DSA n-value for each parameter set (index =3D param_set - 1) */ +extern const u32 slhdsa_n[]; + +/* SLH-DSA signature sizes (index =3D param_set - 1) */ +extern const u32 slhdsa_sig_size[]; + +/* Derive PK/SK/seed sizes from n */ +static inline u32 slhdsa_pk_size(u32 param_set) +{ + if (param_set < 1U || param_set > HCQ_SLHDSA_PARAM_MAX) + return 0; + return 2U * slhdsa_n[param_set - 1U]; +} + +static inline u32 slhdsa_sk_size(u32 param_set) +{ + if (param_set < 1U || param_set > HCQ_SLHDSA_PARAM_MAX) + return 0; + return 4U * slhdsa_n[param_set - 1U]; +} + +static inline u32 slhdsa_seed_size(u32 param_set) +{ + if (param_set < 1U || param_set > HCQ_SLHDSA_PARAM_MAX) + return 0; + return 3U * slhdsa_n[param_set - 1U]; +} + +static inline u32 slhdsa_get_sig_size(u32 param_set) +{ + if (param_set < 1U || param_set > HCQ_SLHDSA_PARAM_MAX) + return 0; + return slhdsa_sig_size[param_set - 1U]; +} + +/* HCQ Command Structures -- match CMH eSW ABI exactly */ + +struct hcq_cmd_xmss_verify { + u32 xmss_mt; /* 0 =3D XMSS, 1 =3D XMSS-MT */ + u32 pk_len; + u32 sig_len; + u32 dig_len; + u64 pk; + u64 sig; + u64 dig; +}; + +struct hcq_cmd_lms_verify { + u32 lms_hss; /* 0 =3D LMS, 1 =3D LMS-HSS */ + u32 pk_len; + u32 sig_len; + u32 dig_len; + u64 pk; + u64 sig; + u64 dig; +}; + +struct hcq_cmd_slhdsa_verify_internal { + u32 parameter_set; + u32 message_len; + u64 message; + u64 pk; + u64 sig; +}; + +struct hcq_cmd_slhdsa_verify { + u32 parameter_set; + u32 message_len; + u64 message; + u64 context; + u64 pk; + u64 sig; + u32 context_len; +}; + +struct hcq_cmd_slhdsa_verify_prehash { + u32 parameter_set; + u32 prehash_algo; + u32 message_len; + u32 context_len; + u64 message; + u64 context; + u64 pk; + u64 sig; +}; + +struct hcq_cmd_slhdsa_keygen { + u32 parameter_set; + u32 seed_len; + u32 pk_len; + u32 sk_len; + u64 seed; /* DS reference */ + u64 pk; /* extmem addr */ + u64 sk; /* DS reference */ +}; + +struct hcq_cmd_slhdsa_sign_internal { + u32 parameter_set; + u32 message_len; + u64 add_random; /* extmem addr, 0 =3D none */ + u64 message; + u64 sk; /* DS reference */ + u64 sig; /* extmem addr */ +}; + +struct hcq_cmd_slhdsa_sign { + u32 parameter_set; + u32 message_len; + u64 add_random; + u64 message; + u64 context; + u64 sk; /* DS reference */ + u64 sig; /* extmem addr */ + u32 context_len; +}; + +struct hcq_cmd_slhdsa_sign_prehash { + u32 parameter_set; + u32 prehash_algo; + u32 message_len; + u32 context_len; + u64 add_random; + u64 message; + u64 context; + u64 sk; /* DS reference */ + u64 sig; /* extmem addr */ +}; + +struct hcq_cmd_slhdsa_pubgen { + u32 parameter_set; + u32 sk_len; + u64 sk; /* DS reference */ + u64 pk; /* extmem addr */ +}; + +union hcq_cmd { + struct hcq_cmd_xmss_verify cmd_xmss_verify; + struct hcq_cmd_lms_verify cmd_lms_verify; + struct hcq_cmd_slhdsa_verify_internal cmd_slhdsa_verify_internal; + struct hcq_cmd_slhdsa_verify cmd_slhdsa_verify; + struct hcq_cmd_slhdsa_verify_prehash cmd_slhdsa_verify_prehash; + struct hcq_cmd_slhdsa_keygen cmd_slhdsa_keygen; + struct hcq_cmd_slhdsa_sign_internal cmd_slhdsa_sign_internal; + struct hcq_cmd_slhdsa_sign cmd_slhdsa_sign; + struct hcq_cmd_slhdsa_sign_prehash cmd_slhdsa_sign_prehash; + struct hcq_cmd_slhdsa_pubgen cmd_slhdsa_pubgen; +}; + +#endif /* CMH_HCQ_ABI_H */ diff --git a/drivers/crypto/cmh/include/cmh_kic_abi.h b/drivers/crypto/cmh/= include/cmh_kic_abi.h new file mode 100644 index 000000000000..7f4fe3b9fd89 --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_kic_abi.h @@ -0,0 +1,77 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- KIC Core ABI Definitions + * + * Kernel-side definitions for the CMH KIC ABI (KIC commands only). + * Derived from the CMH eSW ABI. + */ + +#ifndef CMH_KIC_ABI_H +#define CMH_KIC_ABI_H + +#include + +/* KIC Commands */ + +#define KIC_CMD_HKDF1 0x06U +#define KIC_CMD_HKDF2 0x07U +#define KIC_CMD_AES_CMAC_KDF 0x08U +#define KIC_CMD_DKEK_DERIVE 0x09U + +/* Maximum key size for KIC operations (bytes) */ +#define KIC_KEY_SIZE 32U + +/* + * KIC Command Structures + * + * Field names (llen, len) mirror the CMH eSW ABI register layout. + * llen =3D label length, len =3D output key length. + */ + +struct kic_cmd_hkdf1 { + u64 dst; /* DS ref for derived key (SYS_REF_LAST) */ + u64 base; /* base key reference (e.g., KIC_KEY1) */ + u64 label; /* label pointer (0 for inline-next-slot) */ + u32 llen; /* label length */ + u32 len; /* output key length */ + u32 type; /* SYS_TYPE_SET(flags, core_id) */ +}; + +struct kic_cmd_hkdf2 { + u64 dst; /* DS ref for derived key */ + u64 base; /* base key reference */ + u64 salt; /* salt key reference (SYS_REF_NONE =3D no salt) */ + u64 label; /* label pointer */ + u32 llen; /* label length */ + u32 len; /* output key length */ + u32 type; /* SYS_TYPE_SET(flags, core_id) */ +}; + +struct kic_cmd_aes_cmac_kdf { + u64 base_key; /* KIC/DS reference for base key */ + u64 out_key; /* DS reference for derived key */ + u64 label; /* label DMA address */ + u32 key_len; /* base & output key length (must be 32) */ + u32 label_len; /* label length */ + u32 type; /* SYS_TYPE_SET(flags, core_id) for output */ +}; + +struct kic_cmd_dkek_derive { + u64 base_key; /* KIC base key reference */ + u64 out_key; /* DS reference for the derived KEK */ + u32 host_id; /* host ID (0 =3D caller's own) */ + u32 metadata_len; /* metadata length */ + u64 metadata; /* metadata DMA address */ +}; + +/* KIC Command Union */ + +union kic_cmd { + struct kic_cmd_hkdf1 cmd_hkdf1; + struct kic_cmd_hkdf2 cmd_hkdf2; + struct kic_cmd_aes_cmac_kdf cmd_aes_cmac_kdf; + struct kic_cmd_dkek_derive cmd_dkek_derive; +}; + +#endif /* CMH_KIC_ABI_H */ diff --git a/drivers/crypto/cmh/include/cmh_mqi.h b/drivers/crypto/cmh/incl= ude/cmh_mqi.h new file mode 100644 index 000000000000..202d52b86fb8 --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_mqi.h @@ -0,0 +1,35 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Mailbox Queue Initializer + * + * Allocates DMA-capable queue buffers and programs MBX registers + * via the MBX lock/setup/enable/unlock register sequence. + */ + +#ifndef CMH_MQI_H +#define CMH_MQI_H + +#include "cmh_config.h" + +#define MBX_LOCK_TIMEOUT_MS 1000 +#define MBX_LOCK_POLL_MIN_US 10 +#define MBX_LOCK_POLL_MAX_US 50 + +/** + * cmh_mqi_init() - Allocate MBX queue buffers and program registers + * @cfg: Global device configuration + * + * Performs the lock/setup/enable/unlock sequence for each configured MBX. + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mqi_init(struct cmh_config *cfg); + +/** + * cmh_mqi_cleanup() - Free MBX queue buffers and release locks + * @cfg: Global device configuration + */ +void cmh_mqi_cleanup(struct cmh_config *cfg); + +#endif /* CMH_MQI_H */ diff --git a/drivers/crypto/cmh/include/cmh_pke_abi.h b/drivers/crypto/cmh/= include/cmh_pke_abi.h new file mode 100644 index 000000000000..e0e7b946b4e3 --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_pke_abi.h @@ -0,0 +1,272 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- PKE Core ABI Definitions + * + * Kernel-side definitions for the CMH PKE ABI. + * All constants and layouts derived from the CMH eSW ABI. + */ + +#ifndef CMH_PKE_ABI_H +#define CMH_PKE_ABI_H + +#include + +/* PKE Command IDs */ + +#define PKE_CMD_ECDSA_VERIFY 0x03U +#define PKE_CMD_ECDSA_SIGN 0x04U +#define PKE_CMD_ECDSA_PUBGEN 0x05U +#define PKE_CMD_ECDSA_KEYGEN 0x06U +#define PKE_CMD_EDDSA_VERIFY 0x07U +#define PKE_CMD_EDDSA_SIGN 0x08U +#define PKE_CMD_EDDSA_PUBGEN 0x09U +#define PKE_CMD_ECDH_KEYGEN 0x0AU +#define PKE_CMD_ECDH 0x0BU +#define PKE_CMD_RSA_ENC 0x0CU +#define PKE_CMD_RSA_DEC 0x0DU +#define PKE_CMD_RSA_KEYGEN 0x0EU +#define PKE_CMD_RSA_CRT_DEC 0x0FU +#define PKE_CMD_SM2_ECDH_KEYGEN 0x16U +#define PKE_CMD_SM2_ECDH 0x17U +#define PKE_CMD_SM2_DEC_POINT 0x18U +#define PKE_CMD_SM2_ENC_POINT 0x19U +#define PKE_CMD_SM2_ID_DIGEST 0x1AU +#define PKE_CMD_SM2_ECDH_HASH 0x1BU +#define PKE_CMD_SM2_DEC_HASH 0x1CU +#define PKE_CMD_SM2_ENC_HASH 0x1DU +#define PKE_CMD_EDDSA_PRIV_KEYGEN_SCA 0x21U +#define PKE_CMD_FLUSH 0xFFU + +/* EC Curve IDs (per CMH PKE ABI) */ + +#define PKE_CURVE_P192 0x01U +#define PKE_CURVE_P224 0x02U +#define PKE_CURVE_P256 0x03U +#define PKE_CURVE_P384 0x04U +#define PKE_CURVE_P521 0x05U +#define PKE_CURVE_SECP256K1 0x07U +#define PKE_CURVE_BP192R1 0x11U +#define PKE_CURVE_BP224R1 0x12U +#define PKE_CURVE_BP256R1 0x13U +#define PKE_CURVE_BP320R1 0x14U +#define PKE_CURVE_BP384R1 0x15U +#define PKE_CURVE_BP512R1 0x16U +#define PKE_CURVE_ANSSI_FRP256V1 0x17U +#define PKE_CURVE_SM2 0x18U +#define PKE_CURVE_25519 0x21U +#define PKE_CURVE_448 0x22U + +/* PKE Command Structures -- match CMH eSW ABI exactly */ + +struct pke_cmd_ecdsa_verify { + u32 curve; + u32 digest_len; + u64 public_key; + u64 digest; + u64 signature; + u64 rprime; +}; + +struct pke_cmd_ecdsa_sign { + u32 curve; + u32 secret_key_len; + u64 digest; + u64 signature; + u64 secret_key; /* DS reference */ + u32 digest_len; +}; + +struct pke_cmd_ecdsa_pubgen { + u32 curve; + u32 secret_key_len; + u64 public_key; + u64 secret_key; /* DS reference */ +}; + +struct pke_cmd_ecdsa_keygen { + u32 curve; + u32 secret_key_len; + u64 secret_key; /* DS reference */ + u32 secret_key_type; +}; + +struct pke_cmd_eddsa_verify { + u32 curve; + u32 digest_len; + u64 public_key_y; + u64 digest; + u64 signature; + u64 rprime; +}; + +struct pke_cmd_eddsa_sign { + u32 curve; + u32 secret_key_len; + u64 digest; + u64 signature; + u64 secret_key; /* DS reference */ + u32 digest_len; +}; + +struct pke_cmd_eddsa_pubgen { + u32 curve; + u32 secret_key_len; + u64 public_key_y; + u64 secret_key; /* DS reference */ +}; + +struct pke_cmd_ecdh_keygen { + u32 curve; + u32 secret_key_len; + u64 public_key_x; + u64 secret_key; /* DS reference */ +}; + +struct pke_cmd_ecdh { + u32 curve; + u32 secret_key_len; + u32 shared_secret_len; + u32 shared_secret_type; + u64 peer_key_x; + u64 secret_key; /* DS reference */ + u64 shared_secret; /* DS reference for result */ +}; + +struct pke_cmd_rsa_enc { + u32 bits; + u32 e_len; + u64 e; + u64 n; + u64 m; + u64 c; +}; + +struct pke_cmd_rsa_dec { + u32 bits; + u32 e_len; + u64 e; + u64 n; + u64 c; + u64 m; + u64 d; /* DS reference */ +}; + +struct pke_cmd_rsa_crt_dec { + u32 bits; + u32 e_len; + u64 e; + u64 n; + u64 c; + u64 m; + u64 crt; /* DS reference */ +}; + +struct pke_cmd_rsa_keygen { + u32 bits; + u32 d_type; + u64 e; + u64 n; + u64 d; /* DS reference */ + u64 crt; /* DS reference */ + u32 crt_type; +}; + +struct pke_cmd_eddsa_keygen_sca { + u32 curve; + u64 secret_key; /* DS reference: input normal SK */ + u64 sca_secret_key; /* DS reference: output blinded SK */ +}; + +/* SM2 Command Structures */ + +struct pke_cmd_sm2_ecdh_keygen { + u64 nonce; /* DMA addr (32B input or output) */ + u64 session_key; /* DMA addr output (64B) */ + u32 nonce_len; /* 0 =3D HW generates, 32 =3D caller provides */ +}; + +struct pke_cmd_sm2_ecdh { + u32 nonce_len; /* 0 or 32 */ + u32 private_key_len; /* must be 32 */ + u64 nonce; /* DMA addr (32B) */ + u64 peer_public_key; /* DMA addr (64B) */ + u64 peer_session_key; /* DMA addr (64B) */ + u64 private_key; /* DS reference */ + u64 shared_point; /* DS reference (output, 64B) */ + u32 shared_point_type; /* SYS_TYPE_SET(flags, CORE_ID_PKE) */ +}; + +struct pke_cmd_sm2_dec_point { + u32 ciphertext_len; /* total CT length (97..128) */ + u32 private_key_len; /* must be 32 */ + u64 ciphertext; /* DMA addr (64B: C1 point) */ + u64 dec_point; /* DMA addr output (64B) */ + u64 private_key; /* DS reference */ +}; + +struct pke_cmd_sm2_enc_point { + u64 nonce; /* DMA addr (32B, optional) */ + u64 public_key; /* DMA addr (64B) */ + u64 ciphertext; /* DMA addr output (64B: C1) */ + u64 enc_point; /* DMA addr output (64B) */ + u32 nonce_len; /* 0 or 32 */ +}; + +struct pke_cmd_sm2_id_digest { + u64 id; /* DMA addr (identity, <=3D32B) */ + u64 public_key; /* DMA addr (64B) */ + u64 digest; /* DMA addr output (32B) */ + u32 id_len; /* identity length in bytes */ +}; + +struct pke_cmd_sm2_ecdh_hash { + u64 peer_id_digest; /* DMA addr (32B) */ + u64 id_digest; /* DMA addr (32B) */ + u64 shared_point; /* DS reference (64B input) */ + u64 shared_key; /* DS reference (16B output) */ + u32 shared_key_type; /* SYS_TYPE_SET(flags, CORE_ID_PKE) */ +}; + +struct pke_cmd_sm2_dec_hash { + u64 ciphertext; /* DMA addr (full ciphertext) */ + u64 dec_point; /* DMA addr (64B) */ + u64 plaintext; /* DMA addr output (ct_len - 96 bytes) */ + u32 ciphertext_len; /* 97..128 */ +}; + +struct pke_cmd_sm2_enc_hash { + u64 message; /* DMA addr (plaintext) */ + u64 enc_point; /* DMA addr (64B) */ + u64 ciphertext; /* DMA addr output (96 + msg_len) */ + u32 message_len; /* 1..32 */ +}; + +/* PKE Command Union */ + +union pke_cmd { + struct pke_cmd_ecdsa_verify cmd_ecdsa_verify; + struct pke_cmd_ecdsa_sign cmd_ecdsa_sign; + struct pke_cmd_ecdsa_pubgen cmd_ecdsa_pubgen; + struct pke_cmd_ecdsa_keygen cmd_ecdsa_keygen; + struct pke_cmd_eddsa_verify cmd_eddsa_verify; + struct pke_cmd_eddsa_sign cmd_eddsa_sign; + struct pke_cmd_eddsa_pubgen cmd_eddsa_pubgen; + struct pke_cmd_ecdh_keygen cmd_ecdh_keygen; + struct pke_cmd_ecdh cmd_ecdh; + struct pke_cmd_rsa_enc cmd_rsa_enc; + struct pke_cmd_rsa_dec cmd_rsa_dec; + struct pke_cmd_rsa_crt_dec cmd_rsa_crt_dec; + struct pke_cmd_rsa_keygen cmd_rsa_keygen; + struct pke_cmd_eddsa_keygen_sca cmd_eddsa_keygen_sca; + struct pke_cmd_sm2_ecdh_keygen cmd_sm2_ecdh_keygen; + struct pke_cmd_sm2_ecdh cmd_sm2_ecdh; + struct pke_cmd_sm2_dec_point cmd_sm2_dec_point; + struct pke_cmd_sm2_enc_point cmd_sm2_enc_point; + struct pke_cmd_sm2_id_digest cmd_sm2_id_digest; + struct pke_cmd_sm2_ecdh_hash cmd_sm2_ecdh_hash; + struct pke_cmd_sm2_dec_hash cmd_sm2_dec_hash; + struct pke_cmd_sm2_enc_hash cmd_sm2_enc_hash; +}; + +#endif /* CMH_PKE_ABI_H */ diff --git a/drivers/crypto/cmh/include/cmh_qse_abi.h b/drivers/crypto/cmh/= include/cmh_qse_abi.h new file mode 100644 index 000000000000..9834620e21d7 --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_qse_abi.h @@ -0,0 +1,181 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- QSE Core ABI Definitions + * + * Kernel-side definitions for the CMH QSE ABI. + * All constants and layouts derived from the CMH eSW ABI. + */ + +#ifndef CMH_QSE_ABI_H +#define CMH_QSE_ABI_H + +#include +#include +#include + +/* VCQ layout: header + [SYS_NEW] + QSE_CMD + flush */ +#define QSE_VCQ_CMDS_MIN 3 /* header + cmd + flush */ +#define QSE_VCQ_CMDS_MAX 4 /* header + sys_new + cmd + flush */ + +/* QSE Flags */ +#define QSE_FLAG_USE_REF BIT(0) +#define QSE_FLAG_USE_RNG BIT(1) + +/* QSE Command IDs */ +#define QSE_CMD_ML_KEM_KEYGEN 0x01U +#define QSE_CMD_ML_KEM_ENC 0x02U +#define QSE_CMD_ML_KEM_DEC 0x03U +#define QSE_CMD_ML_DSA_KEYGEN 0x04U +#define QSE_CMD_ML_DSA_SIGN 0x05U +#define QSE_CMD_ML_DSA_VERIFY 0x06U +#define QSE_CMD_ML_KEM_KEYGEN_MASKED 0x07U +#define QSE_CMD_ML_KEM_ENC_MASKED 0x08U +#define QSE_CMD_ML_KEM_DEC_MASKED 0x09U +#define QSE_CMD_ML_DSA_KEYGEN_MASKED 0x0AU +#define QSE_CMD_ML_DSA_SIGN_MASKED 0x0BU + +/* ML-KEM category values */ +#define ML_KEM_K_512 2U +#define ML_KEM_K_768 3U +#define ML_KEM_K_1024 4U + +/* ML-DSA mode values */ +#define ML_DSA_MODE_44 2U +#define ML_DSA_MODE_65 3U +#define ML_DSA_MODE_87 5U + +/* ML-DSA special message length for externalMu (pre-hashed 64-byte input)= */ +#define ML_DSA_MLEN_EXTERNAL_MU 0xFFFFFFFFU +#define ML_DSA_EXTMU_LEN 64U /* actual copy size for externalMu */ + +/* ML-DSA maximum message length */ +#define ML_DSA_MAX_MLEN 10240U + +/* Shared secret size */ +#define ML_KEM_SS_LEN 32U +#define ML_KEM_SS_LEN_MASKED 64U + +/* Seed sizes */ +#define QSE_SEED_LEN 32U +#define QSE_SEED_LEN_MASKED 64U + +/* + * ML-KEM size tables -- indexed by (k - 2). + * [0] =3D ML-KEM-512 (k=3D2) + * [1] =3D ML-KEM-768 (k=3D3) + * [2] =3D ML-KEM-1024 (k=3D4) + */ +#define ML_KEM_LEVELS 3U + +#define ML_KEM_EK_SIZE(k) (384U * (k) + 32U) +#define ML_KEM_DK_SIZE(k) (768U * (k) + 96U) +#define ML_KEM_DK_SIZE_MASKED(k) (1152U * (k) + 128U) + +static inline u32 ml_kem_ct_size(u32 k) +{ + u32 du =3D (k =3D=3D 4U) ? 11U : 10U; + u32 dv =3D (k =3D=3D 4U) ? 5U : 4U; + + return 32U * (k * du + dv); +} + +#define ML_KEM_CT_SIZE(k) ml_kem_ct_size(k) + +/* + * ML-DSA size tables -- indexed by mode. + * Mode values: 2 (ML-DSA-44), 3 (ML-DSA-65), 5 (ML-DSA-87). + */ +extern const u32 ml_dsa_pk_size[]; +extern const u32 ml_dsa_sk_size[]; +extern const u32 ml_dsa_sk_size_masked[]; +extern const u32 ml_dsa_sig_size[]; + +/* Map ML-DSA mode (2/3/5) -> table index (0/1/2) */ +static inline int ml_dsa_mode_idx(u32 mode) +{ + switch (mode) { + case 2: return 0; + case 3: return 1; + case 5: return 2; + default: return -1; + } +} + +/* Map ML-KEM k (2/3/4) -> table index (0/1/2), or -1 if invalid */ +static inline int ml_kem_k_idx(u32 k) +{ + if (k >=3D 2U && k <=3D 4U) + return (int)(k - 2U); + return -1; +} + +/* QSE Command Structures -- match CMH eSW ABI exactly */ + +struct qse_cmd_ml_kem_keygen { + u32 k; + u32 flags; + u64 seed; + u64 z; + u64 ek; + u64 dk; + u32 dk_type; +}; + +struct qse_cmd_ml_kem_enc { + u32 k; + u32 flags; + u64 coin; + u64 ek; + u64 ct; + u64 ss; + u32 ss_type; +}; + +struct qse_cmd_ml_kem_dec { + u32 k; + u32 flags; + u64 ct; + u64 dk; + u64 ss; + u32 ss_type; +}; + +struct qse_cmd_ml_dsa_keygen { + u32 mode; + u32 flags; + u64 seed; + u64 pk; + u64 sk; + u32 sk_type; +}; + +struct qse_cmd_ml_dsa_sign { + u32 mode; + u32 flags; + u64 rnd; + u64 m; + u64 sk; + u64 sig; + u32 mlen; +}; + +struct qse_cmd_ml_dsa_verify { + u32 mode; + u32 flags; + u64 m; + u64 pk; + u64 sig; + u32 mlen; +}; + +union qse_cmd { + struct qse_cmd_ml_kem_keygen cmd_ml_kem_keygen; + struct qse_cmd_ml_kem_enc cmd_ml_kem_enc; + struct qse_cmd_ml_kem_dec cmd_ml_kem_dec; + struct qse_cmd_ml_dsa_keygen cmd_ml_dsa_keygen; + struct qse_cmd_ml_dsa_sign cmd_ml_dsa_sign; + struct qse_cmd_ml_dsa_verify cmd_ml_dsa_verify; +}; + +#endif /* CMH_QSE_ABI_H */ diff --git a/drivers/crypto/cmh/include/cmh_registers.h b/drivers/crypto/cm= h/include/cmh_registers.h new file mode 100644 index 000000000000..668cf319cd70 --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_registers.h @@ -0,0 +1,161 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Hardware Register Definitions + * + * Derived from the CMH hardware register specification. + * All offsets are taken directly from the hardware documentation. + */ + +#ifndef CMH_REGISTERS_H +#define CMH_REGISTERS_H + +#include +#include + +/* MBX Instance Addressing */ + +#define CMH_MBX_INSTANCE_SHIFT 12 +#define CMH_MBX_INSTANCE_SIZE BIT(CMH_MBX_INSTANCE_SHIFT) /* 0x100= 0 */ +#define CMH_MAX_MBX_INSTANCES 64U + +/* MBX Per-Instance Register Offsets */ + +#define R_MBX_LOCK 0x000U +#define R_MBX_HOST_INFO 0x004U +#define R_MBX_QUEUE_LO 0x008U +#define R_MBX_QUEUE_HI 0x00CU +#define R_MBX_QUEUE_SLOTS 0x010U +#define R_MBX_QUEUE_STRIDE 0x014U +#define R_MBX_QUEUE_HEAD 0x018U +#define R_MBX_QUEUE_TAIL 0x01CU +#define R_MBX_INTERRUPT 0x020U +#define R_MBX_INTERRUPT_MASK 0x024U +#define R_MBX_COMMAND 0x028U +#define R_MBX_STATUS 0x02CU +#define R_MBX_CHILD 0x030U +#define R_MBX_ID 0x034U +#define R_MBX_HOST_CONFIG 0x038U +#define R_MBX_SCRATCH 0x03CU + +#define MBX_QUEUE_ALIGNMENT 0x4U + +/* MBX Interrupt Bits */ + +#define MBX_DONE_IRQ BIT(0) +#define MBX_ERROR_IRQ BIT(1) +#define MBX_IRQ_MASK (MBX_DONE_IRQ | MBX_ERROR_IRQ) + +/* MBX Command Values */ + +#define MBX_COMMAND_RUN 0x000U +#define MBX_COMMAND_PAUSE 0xC2FU +#define MBX_COMMAND_CONTINUE 0x5DBU +#define MBX_COMMAND_RESTART 0xB78U +#define MBX_COMMAND_ABORT 0x6F6U +#define MBX_COMMAND_FLUSH 0x3A5U + +/* MBX Status Values */ + +#define MBX_STATUS_IDLE 0x01U +#define MBX_STATUS_BUSY 0x10U +#define MBX_STATUS_HOLD 0x20U +#define MBX_STATUS_PAUSED 0x28U +#define MBX_STATUS_SUCCESS 0x40U +#define MBX_STATUS_ERROR 0x80U +#define MBX_STATUS_OFFLINE 0x88U /* ERROR | 0x08: offline/stop= ped */ + +#define MBX_MASK_DONE (MBX_STATUS_IDLE | MBX_STATUS_SUCCES= S) +#define MBX_MASK_RUNNING (MBX_STATUS_BUSY | MBX_STATUS_HOLD) +#define MBX_MASK_STOPPED MBX_STATUS_OFFLINE + +/* MBX Status Field Extraction */ + +#define MBX_STATUS_CODE(v) ((v) & 0xFFU) +#define MBX_STATUS_CORE_ID(v) (((v) >> 8) & 0xFFU) +#define MBX_STATUS_ERROR_CODE(v) (((v) >> 16) & 0xFFU) +#define MBX_STATUS_CMD_INDEX(v) (((v) >> 24) & 0xFFU) + +/* SIC Register Offsets (relative to SIC base / instance 0 base) */ + +#define R_SIC_BOOT_STATUS 0x100U +#define SIC_BOOT_STATUS_MASK 0x77U +#define SIC_BOOT_STATUS_PASS 0x66U + +#define R_SIC_MBX_AVAILABILITY 0x104U +#define R_SIC_MBX_AVAILABILITY2 0x108U + +#define R_SIC_SW_BOOT_STATUS 0x12CU +#define SIC_SW_BOOT_STATUS_STARTED BIT(0) +#define SIC_SW_BOOT_STATUS_READY BIT(1) +#define SIC_SW_BOOT_STATUS_MISSION BIT(6) +#define SIC_SW_BOOT_STATUS_MISSION2 BIT(7) + +#define R_SIC_SW_ERROR_INFO 0x130U +#define R_SIC_SW_HEARTBEAT 0x154U + +#define R_SIC_GPINTERRUPT 0x160U + +#define R_SIC_HW_VERSION0 0x200U +#define R_SIC_SW_VERSION 0x218U +#define R_SIC_CORE_ENABLE 0x22CU + +/* + * Per-core dual-rail enable fields within R_SIC_CORE_ENABLE. Each core + * occupies a 2-bit field; the value 0b01 (low bit set, high bit clear) + * means enabled. A core is present iff (val & (mask | mask << 1)) =3D=3D= mask. + */ +#define SIC_CORE_ENABLE_HC 0x00001U +#define SIC_CORE_ENABLE_AES 0x00004U +#define SIC_CORE_ENABLE_SM4 0x00010U +#define SIC_CORE_ENABLE_SM3 0x00040U +#define SIC_CORE_ENABLE_HCQ 0x00100U +#define SIC_CORE_ENABLE_QSE 0x00400U +#define SIC_CORE_ENABLE_PKE 0x01000U +#define SIC_CORE_ENABLE_DRBG 0x04000U +#define SIC_CORE_ENABLE_CCP 0x10000U + +/* Register Access Helpers */ + +static inline u32 cmh_reg_read32(void __iomem *base, u32 offset) +{ + return ioread32((u8 __iomem *)base + offset); +} + +static inline void cmh_reg_write32(u32 value, void __iomem *base, u32 offs= et) +{ + iowrite32(value, (u8 __iomem *)base + offset); +} + +/* + * 64-bit register access via two 32-bit reads/writes. Only correct for + * register pairs where split access is defined (e.g. QUEUE_LO/HI). + * Do not use for registers requiring atomic 64-bit access. + * + * No explicit barrier between the two halves is needed: ioread32/iowrite32 + * include implicit ordering guarantees on all supported architectures + * (MMIO accessors are strongly ordered with respect to each other). + */ +static inline u64 cmh_reg_read64(void __iomem *base, u32 offset) +{ + u32 lo =3D ioread32((u8 __iomem *)base + offset); + u32 hi =3D ioread32((u8 __iomem *)base + offset + 4); + + return ((u64)hi << 32) | lo; +} + +static inline void cmh_reg_write64(u64 value, void __iomem *base, u32 offs= et) +{ + iowrite32((u32)value, (u8 __iomem *)base + offset); + iowrite32((u32)(value >> 32), (u8 __iomem *)base + offset + 4); +} + +/* Return the ioremap'd base for MBX instance N within the SIC region */ +static inline void __iomem *cmh_mbx_instance_base(void __iomem *sic_mapped, + u32 instance) +{ + return (u8 __iomem *)sic_mapped + + ((unsigned long)instance << CMH_MBX_INSTANCE_SHIFT); +} + +#endif /* CMH_REGISTERS_H */ diff --git a/drivers/crypto/cmh/include/cmh_rh.h b/drivers/crypto/cmh/inclu= de/cmh_rh.h new file mode 100644 index 000000000000..b182c203a475 --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_rh.h @@ -0,0 +1,93 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Response Handler + * + * IRQ-driven completion processing. Uses request_threaded_irq(): + * - Hardirq: read+clear MBX interrupt registers, wake thread + * - Threaded handler: walk per-MBX transaction queues, + * fire completion callbacks, free transaction objects + * + * The Response Handler consumes transaction_obj entries enqueued + * by the Transaction Manager (cmh_txn.c) on each per-mailbox txq. + */ + +#ifndef CMH_RH_H +#define CMH_RH_H + +#include "cmh_config.h" + +/** + * cmh_rh_init() - Register IRQ handler and start response processing + * @cfg: Global device configuration + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_rh_init(struct cmh_config *cfg); + +/** + * cmh_rh_cleanup() - Free IRQ and stop response processing + * @cfg: Global device configuration + */ +void cmh_rh_cleanup(struct cmh_config *cfg); + +/** + * cmh_rh_suspend() - Quiesce RH for system suspend + * @cfg: Global device configuration + * + * Cancels the watchdog timer and masks MBX interrupts at the hardware + * level. IRQ handlers remain registered (standard PM pattern). + * The threaded IRQ handler stays active so that cmh_tm_quiesce() + * (called after this) can still drain in-flight transactions via + * IRQ-driven completions. + */ +void cmh_rh_suspend(struct cmh_config *cfg); + +/** + * cmh_rh_resume() - Restart RH after system resume + * @cfg: Global device configuration + * + * Re-synchronises per-MBX head tracking with hardware, clears stale + * interrupt bits, re-enables MBX interrupt masks, and re-arms the + * watchdog timer. Must be called before cmh_tm_resume(). + */ +void cmh_rh_resume(struct cmh_config *cfg); + +/* debugfs timeout accessor (debug builds only) */ +#ifdef CONFIG_CRYPTO_DEV_CMH_DEBUG +unsigned int *cmh_rh_timeout_watchdog_ptr(void); +#endif + +/** + * cmh_rh_force_drain_mbx() - FLUSH + drain all pending transactions on a = MBX + * @mbx_idx: Mailbox index to drain + * + * Issues MBX_COMMAND_FLUSH, drains all pending transactions with + * -ECANCELED, and resets all recovery bookkeeping (including the + * wedged flag). Safe to call at any time; acquires rh_process_lock. + * Intended for debugfs last-resort recovery. + */ +void cmh_rh_force_drain_mbx(u32 mbx_idx); + +/** + * cmh_rh_mbx_is_wedged() - Check if a mailbox is permanently wedged + * @mbx_idx: Mailbox index to check + * + * Returns true if the mailbox has failed RESTART+FLUSH recovery and + * is offline. Used by the TM to avoid submitting new work to a dead + * mailbox. + * + * Return: true if wedged, false otherwise (including out-of-range idx). + */ +bool cmh_rh_mbx_is_wedged(u32 mbx_idx); + +/** + * cmh_rh_abort_mbx() - Issue MBX_COMMAND_ABORT under rh_process_lock + * @mbx_idx: Mailbox index to abort + * + * Serialises the ABORT write with RESTART/FLUSH commands issued by the + * watchdog, preventing command-register clobber races. + */ +void cmh_rh_abort_mbx(u32 mbx_idx); + +#endif /* CMH_RH_H */ diff --git a/drivers/crypto/cmh/include/cmh_rng.h b/drivers/crypto/cmh/incl= ude/cmh_rng.h new file mode 100644 index 000000000000..7d24682890c7 --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_rng.h @@ -0,0 +1,31 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Hardware RNG (DRBG) Driver + * + * Registers a struct hwrng backed by the CMH DRBG core. + * Each .read() builds a VCQ with DRBG_CMD_GENERATE and submits it + * through the Transaction Manager for synchronous completion. + * + * The DRBG must be configured (CONFIG command) by the management host + * before the LKM is loaded -- the LKM only issues GENERATE requests. + * + * CRNG seeding: hwrng .quality is left at 0, which the hwrng core + * elevates to full trust (1024) for a hardware RNG. Set a fixed + * quality in the hwrng initializer to lower the entropy estimate. + */ + +#ifndef CMH_RNG_H +#define CMH_RNG_H + +struct platform_device; + +int cmh_rng_register(struct platform_device *pdev); +void cmh_rng_unregister(void); + +/* debugfs timeout accessor (debug builds only) */ +#ifdef CONFIG_CRYPTO_DEV_CMH_DEBUG +unsigned int *cmh_rng_timeout_drbg_ptr(void); +#endif + +#endif /* CMH_RNG_H */ diff --git a/drivers/crypto/cmh/include/cmh_sm3_abi.h b/drivers/crypto/cmh/= include/cmh_sm3_abi.h new file mode 100644 index 000000000000..cbbe80fe18d6 --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_sm3_abi.h @@ -0,0 +1,79 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- SM3 Hash Core ABI Definitions + * + * Kernel-side definitions for the CMH SM3 ABI. + * All constants and layouts derived from the CMH eSW ABI. + */ + +#ifndef CMH_SM3_ABI_H +#define CMH_SM3_ABI_H + +#include + +/* SM3 Commands */ + +#define SM3_CMD_INIT 0x01U +#define SM3_CMD_UPDATE 0x02U +#define SM3_CMD_FINAL 0x03U +#define SM3_CMD_UPDATE2D 0x04U +#define SM3_CMD_GATHER 0x06U +#define SM3_CMD_SAVE 0x07U +#define SM3_CMD_RESTORE 0x08U + +/* SM3 Digest / Block Sizes */ + +#define CMH_SM3_DIGEST_SIZE 32U +#define CMH_SM3_BLOCK_SIZE 64U + +/* SM3 Context (for SAVE/RESTORE) */ + +#define SM3_CONTEXT_WORDS 29U +#define SM3_CONTEXT_SIZE (SM3_CONTEXT_WORDS * 4 + 4) /* ctx[29] + = crc */ + +/* SM3 Command Structures */ + +struct sm3_cmd_update { + u64 input; /* DMA physical address of input data */ + u32 inlen; /* input data length in bytes */ +}; + +struct sm3_cmd_final { + u64 digest; /* DMA physical address for output digest */ + u32 outlen; /* digest length in bytes */ +}; + +struct sm3_cmd_update2d { + u64 input; /* DMA source address for input data */ + u64 output; /* DMA destination address for pass-through data */ + u32 iolen; /* input/pass-through data length in bytes */ +}; + +struct sm3_cmd_gather { + u64 lista; /* DMA address of dma_scattergather_item chain */ + u32 sgcmd; /* SM3 sub-command: SM3_CMD_UPDATE or SM3_CMD_UPDATE2D */ +}; + +struct sm3_cmd_save { + u64 output; /* DMA physical address for saved context */ + u32 outlen; /* must be SM3_CONTEXT_SIZE */ +}; + +struct sm3_cmd_restore { + u64 input; /* DMA physical address of saved context */ + u32 inlen; /* must be SM3_CONTEXT_SIZE */ +}; + +/* SM3 Command Union */ + +union sm3_cmd { + struct sm3_cmd_update cmd_update; + struct sm3_cmd_final cmd_final; + struct sm3_cmd_update2d cmd_update2d; + struct sm3_cmd_gather cmd_gather; + struct sm3_cmd_save cmd_save; + struct sm3_cmd_restore cmd_restore; +}; + +#endif /* CMH_SM3_ABI_H */ diff --git a/drivers/crypto/cmh/include/cmh_sm4_abi.h b/drivers/crypto/cmh/= include/cmh_sm4_abi.h new file mode 100644 index 000000000000..a34faea613dc --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_sm4_abi.h @@ -0,0 +1,101 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- SM4 Core ABI Definitions + * + * Kernel-side definitions for the CMH SM4 ABI. + * All constants and layouts derived from the CMH eSW ABI. + */ + +#ifndef CMH_SM4_ABI_H +#define CMH_SM4_ABI_H + +#include + +/* SM4 Block Size */ + +#define CMH_SM4_BLOCK_SIZE 16U +#define CMH_SM4_IV_SIZE 16U +#define CMH_SM4_KEY_SIZE 16U /* SM4 always uses 128-bit keys */ + +/* SM4 Modes (per CMH SM4 ABI) */ + +#define SM4_MODE_ECB 1U +#define SM4_MODE_CBC 2U +#define SM4_MODE_CTR 3U +#define SM4_MODE_CFB 5U +#define SM4_MODE_GCM 6U +#define SM4_MODE_CMAC 7U +#define SM4_MODE_CCM 8U +#define SM4_MODE_XTS 9U +#define SM4_MODE_XCBC 10U + +/* SM4 Operations (per CMH SM4 ABI) */ + +#define SM4_OP_DECRYPT 1U +#define SM4_OP_ENCRYPT 2U + +/* SM4 Command IDs */ + +#define SM4_CMD_INIT 0x01U +#define SM4_CMD_AAD_UPDATE 0x02U +#define SM4_CMD_AAD_FINAL 0x03U +#define SM4_CMD_UPDATE 0x04U +#define SM4_CMD_FINAL 0x05U +#define SM4_CMD_SCATTERGATHER 0x06U +#define SM4_CMD_CCM_INIT 0x09U + +/* SM4 Command Structures */ + +struct sm4_cmd_init { + u64 key; /* datastore reference for the key */ + u64 iv; /* DMA address of the IV */ + u32 keylen; /* key length in bytes (16, or 32 for XTS) */ + u32 ivlen; /* IV length in bytes (0..16) */ + u32 mode; /* SM4 mode (SM4_MODE_*) */ + u32 op; /* SM4 operation (SM4_OP_*) */ + u32 aadlen; /* AAD length or 0 */ + u32 iolen; /* plaintext/ciphertext length */ +}; + +struct sm4_cmd_update { + u64 input; /* DMA address of input data */ + u64 output; /* DMA address of output data */ + u32 iolen; /* input/output data length */ +}; + +struct sm4_cmd_final { + u64 input; /* DMA address of last input data */ + u64 output; /* DMA address of last output data */ + u64 tag; /* DMA address of tag (AEAD only) */ + u32 iolen; /* last input/output data length */ + u32 taglen; /* tag length (AEAD only) */ +}; + +struct sm4_cmd_aad_final { + u64 data; /* DMA address of AAD data */ + u32 datalen; /* AAD data length */ +}; + +struct sm4_cmd_ccm_init { + u64 key; /* datastore reference for the key */ + u64 nonce; /* DMA address of the nonce */ + u32 keylen; /* key length in bytes (always 16) */ + u32 noncelen; /* nonce length (15 - L) */ + u32 op; /* SM4 operation (SM4_OP_*) */ + u32 aadlen; /* AAD length */ + u32 iolen; /* plaintext/ciphertext length */ + u32 taglen; /* tag length */ +}; + +/* SM4 Command Union */ + +union sm4_cmd { + struct sm4_cmd_init cmd_init; + struct sm4_cmd_update cmd_update; + struct sm4_cmd_final cmd_final; + struct sm4_cmd_aad_final cmd_aad_final; + struct sm4_cmd_ccm_init cmd_ccm_init; +}; + +#endif /* CMH_SM4_ABI_H */ diff --git a/drivers/crypto/cmh/include/cmh_sys_abi.h b/drivers/crypto/cmh/= include/cmh_sys_abi.h new file mode 100644 index 000000000000..64110311e552 --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_sys_abi.h @@ -0,0 +1,148 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- SYS Core ABI Definitions + * + * Kernel-side definitions for the CMH SYS ABI. + * All constants and layouts derived from the CMH eSW ABI. + */ + +#ifndef CMH_SYS_ABI_H +#define CMH_SYS_ABI_H + +#include +#include + +/* SYS Commands (per CMH SYS ABI) */ + +#define SYS_CMD_RUN 0x01U +#define SYS_CMD_NOP 0x02U +#define SYS_CMD_IMPORT 0x07U +#define SYS_CMD_EXPORT 0x08U +#define SYS_CMD_NEW 0x0AU +#define SYS_CMD_READ 0x0BU +#define SYS_CMD_WRITE 0x0CU +#define SYS_CMD_GRANT 0x0DU +#define SYS_CMD_LIST 0x0EU +#define SYS_CMD_FIND 0x0FU +#define SYS_CMD_DATA 0x11U + +/* SYS Reference Constants */ + +#define SYS_REF_NONE 0x0000000000000000ULL +#define SYS_REF_TEMP 0x1111111111111111ULL +#define SYS_REF_LAST 0xFFFFFFFFFFFFFFFFULL + +typedef u64 sys_ref_t; + +/* SYS CID */ + +#define SYS_CID_NONE 0x0000000000000000ULL + +/* SYS Type Encoding -- bits [7:0] =3D core_id, bits [23:16] =3D flags */ + +#define SYS_TYPE_FLAG_PT BIT(16) /* can be read as plaintext */ +#define SYS_TYPE_FLAG_XC BIT(17) /* can be exported over XC bus */ +#define SYS_TYPE_FLAG_SCA BIT(18) /* SCA key in 2 shares */ + +#define SYS_TYPE_SET(flags, core) \ + (((flags) & 0xFF0000UL) | ((core) & 0xFFUL)) +#define SYS_TYPE_CORE(type) ((type) & 0xFFU) +#define SYS_TYPE_FLAGS(type) ((type) & 0xFF0000U) +#define SYS_TYPE_NONE 0U /* DMA output, no DS storage */ + +#define SYS_WRAP_HDR_SIZE 16 /* sys_read plaintext header */ + +/* SYS Command Structures */ + +struct sys_cmd_new { + u64 cid; /* caller id (name) for the object */ + u64 ref; /* DMA address -- CMH eSW writes back reference here */ + u32 len; /* size of the new object in bytes */ +}; + +struct sys_cmd_write { + u64 ref; /* object datastore reference */ + u64 src; /* DMA source address of key data */ + u64 key; /* wrapping key reference (SYS_REF_NONE =3D plaintext) */ + u32 len; /* source buffer length */ + u32 type; /* SYS_TYPE_SET(flags, core_id) */ +}; + +struct sys_cmd_read { + u64 ref; /* object datastore reference */ + u64 dst; /* DMA destination for key data */ + u64 key; /* wrapping key reference (SYS_REF_NONE =3D plaintext) */ + u32 len; /* destination buffer length */ +}; + +struct sys_cmd_data { + u64 ref; /* object datastore reference */ + u64 dst; /* DMA destination for object data */ + u32 len; /* destination buffer length */ +}; + +struct sys_cmd_find { + u64 cid; /* caller id to search for */ + u64 dst; /* DMA destination for struct sys_list_item */ + u32 len; /* destination buffer length */ +}; + +struct sys_cmd_list { + u64 ref; /* starting DS reference (SYS_REF_NONE =3D first) */ + u64 dst; /* DMA destination for struct sys_list_item */ + u32 len; /* destination buffer length */ +}; + +struct sys_cmd_grant { + u64 ref; /* object datastore reference */ + u64 read; /* bitfield: allow read for mailboxes */ + u64 write; /* bitfield: allow write for mailboxes */ + u64 execute; /* bitfield: allow use for mailboxes */ +}; + +struct sys_cmd_export { + u64 cid; /* caller id for the response */ + u64 dst; /* DMA destination for the export blob */ + u64 key; /* wrapping key datastore reference */ + u32 len; /* destination buffer length */ +}; + +struct sys_cmd_import { + u64 src; /* DMA source address of import blob */ + u64 key; /* wrapping key datastore reference */ + u32 len; /* source buffer length */ +}; + +/* SYS List/Find Response Item */ + +struct sys_list_item { + u64 ref; /* object datastore reference */ + u64 cid; /* caller id */ + u32 len; /* object length */ + u32 type; /* object type (SYS_TYPE_SET packed) */ +}; + +/* Wrapped-read header (prepended to SYS_CMD_READ responses) */ + +struct sys_wrap_hdr { + u64 cid; /* caller id */ + u32 wrap; /* wrap data length following this header */ + u32 len; /* object data length following wrap data */ +}; + +/* SYS Command Union */ + +union sys_cmd { + struct sys_cmd_new cmd_new; + struct sys_cmd_write cmd_write; + struct sys_cmd_read cmd_read; + struct sys_cmd_data cmd_data; + struct sys_cmd_find cmd_find; + struct sys_cmd_list cmd_list; + struct sys_cmd_grant cmd_grant; + struct sys_cmd_export cmd_export; + struct sys_cmd_import cmd_import; +}; + +#endif /* CMH_SYS_ABI_H */ diff --git a/drivers/crypto/cmh/include/cmh_sysfs.h b/drivers/crypto/cmh/in= clude/cmh_sysfs.h new file mode 100644 index 000000000000..864cf1c8fa00 --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_sysfs.h @@ -0,0 +1,14 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- sysfs Device Attributes + */ + +#ifndef CMH_SYSFS_H +#define CMH_SYSFS_H + +struct attribute_group; + +extern const struct attribute_group *cmh_sysfs_groups[]; + +#endif /* CMH_SYSFS_H */ diff --git a/drivers/crypto/cmh/include/cmh_txn.h b/drivers/crypto/cmh/incl= ude/cmh_txn.h new file mode 100644 index 000000000000..cdfae716adea --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_txn.h @@ -0,0 +1,491 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Transaction Manager + * + * Dedicated kthread managing concurrent VCQ submissions. + * + * Callers post command_msg objects into the Command Message Queue (CMQ). + * The TM thread dequeues them, selects a mailbox, builds VCQ(s) in the + * DMA queue slot, creates a transaction_obj, and rings the doorbell. + * + * The Response Handler (cmh_rh.c) walks per-mailbox transaction queues + * when an IRQ fires and fires completion callbacks. + */ + +#ifndef CMH_TXN_H +#define CMH_TXN_H + +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "cmh_config.h" +#include "cmh_vcq.h" + +/* Command Message (caller -> TM) */ + +typedef void (*cmh_completion_fn)(void *data, int error); + +struct command_msg { + struct list_head list; /* CMQ linked list node */ + u32 command_id; /* VCQ_CMD_ID(core, flags, span, cmd)= */ + void *vcq_data; /* heap-owned copy of VCQ entries */ + u32 vcq_count; /* total vcq_cmd entries across all V= CQs */ + u32 num_vcqs; /* how many VCQs in vcq_data (0 or 1 = =3D single) */ + s32 target_mbx; /* MBX index from core affinity, or -= 1 fallback */ + s32 actual_mbx; /* MBX selected by TM thread, -1 unti= l dispatched */ + u32 first_vcq_id; /* HW VCQ id of first slot, U32_MAX u= ntil submitted */ + cmh_completion_fn complete; /* completion callback (may be NULL) = */ + void *completion_data; + refcount_t refs; /* submit_sync: 2 =3D waiter + TM */ + bool backlog_ok; /* accept into backlog when CMQ is fu= ll */ + unsigned long timeout_jiffies;/* per-txn async timeout (0 =3D none)= */ +}; + +/* Transaction Object (TM -> RH) */ + +/* Per-transaction FSM states for async timeout resolution */ +#define TXN_INFLIGHT 0 +#define TXN_COMPLETE 1 +#define TXN_TIMED_OUT 2 + +struct transaction_obj { + struct list_head list; /* per-mailbox txn queue node */ + u32 first_vcq_id; + u32 last_vcq_id; + u32 mailbox_idx; /* index into cfg->mailboxes[] */ + u32 command_id; /* VCQ_CMD_ID from first payload cmd = */ + int error_code; + cmh_completion_fn complete; + void *completion_data; + atomic_t state; /* TXN_INFLIGHT / COMPLETE / TIMED_OU= T */ + struct timer_list timeout_timer; /* per-request async timeout */ + refcount_t refs; /* owner + timer (if armed) */ +}; + +/* Per-Mailbox Transaction Queue */ + +struct cmh_mbx_txq { + struct list_head head; + spinlock_t lock; /* protects head list + depth */ + u32 depth; /* number of in-flight transactions */ + struct mutex dispatch_lock; /* serialises VCQ dispatch + MBX flus= h */ +}; + +/* Public Interface */ + +/** + * cmh_tm_init() - Initialise the Transaction Manager + * @cfg: Global device configuration (mailbox layout, IRQ, etc.) + * + * Starts the TM kthread and initialises per-mailbox transaction queues. + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_tm_init(struct cmh_config *cfg); + +/** + * cmh_tm_cleanup() - Stop the TM kthread and drain all queues + */ +void cmh_tm_cleanup(void); + +/** + * cmh_tm_quiesce() - Stop TM kthread and drain in-flight transactions + * + * Stops the TM kthread, rejects new posts, then waits (with a + * configurable timeout) for all per-MBX transaction queues to drain. + * If the timeout fires, remaining transactions are cancelled with + * -ECANCELED. + */ +void cmh_tm_quiesce(void); + +/** + * cmh_tm_resume() - Restart the TM kthread after resume + * + * Return: 0 on success, negative errno if the kthread fails to start. + */ +int cmh_tm_resume(void); + +/** + * cmh_tm_post_command() - Post a command to the TM for submission + * @msg: Command message with pre-built VCQ data and completion callback + * + * Round-robin selects the next MBX with enough free slots for + * msg->num_vcqs VCQs. All VCQs in a message are written to + * consecutive slots on the same MBX (back-to-back). + * The caller retains ownership of @msg until the completion callback fire= s. + * + * Return: 0 on success, -EAGAIN if queue full, -ENODEV if TM stopped. + */ +int cmh_tm_post_command(struct command_msg *msg); + +/* + * Synchronous submit -- post one or more VCQs and wait for completion. + * + * Combines post_command + refcounted wait + timeout + cancel into one + * call. This is the standard pattern for all synchronous crypto ops. + * + * Context: must be called from a sleepable (task) context. + * Performs GFP_KERNEL allocations and sleeps on + * wait_for_completion_timeout(). A WARN_ON_ONCE fires + * if called from atomic / IRQ / softirq context. + * + * vcq_cmds: pre-built VCQ array (headers + commands, contiguous) + * vcq_count: total number of vcq_cmd entries across all VCQs + * num_vcqs: number of VCQs in the array (0 or 1 =3D single VCQ) + * + * For multi-VCQ submissions, the array contains multiple VCQs laid + * out contiguously, each starting with its own header. All VCQs are + * written to consecutive MBX slots and share one transaction object. + * + * Returns 0 on success, -ETIMEDOUT, or CMH eSW error code. + */ +int cmh_tm_submit_sync(struct vcq_cmd *vcq_cmds, u32 vcq_count, + u32 num_vcqs); + +/* + * Synchronous submit pinned to a specific mailbox. + * target_mbx: -1 =3D round-robin, >=3D 0 =3D pin to that MBX index. + */ +int cmh_tm_submit_sync_mbx(struct vcq_cmd *vcq_cmds, u32 vcq_count, + u32 num_vcqs, s32 target_mbx); + +/* + * Synchronous submit with explicit timeout. + * timeout_hz: completion timeout in jiffies (use msecs_to_jiffies()). + */ + +/* + * Extended timeout for slow crypto operations: RSA keygen, PQC + * keygen/sign/verify. Controlled by the slow_op_timeout_ms module + * parameter. + */ +unsigned long cmh_tm_slow_op_timeout_jiffies(void); + +int cmh_tm_submit_sync_tmo(struct vcq_cmd *vcq_cmds, u32 vcq_count, + u32 num_vcqs, s32 target_mbx, + unsigned long timeout_hz); + +/* + * Synchronous submit that never issues MBX_COMMAND_ABORT on timeout. + * Returns -EAGAIN if cancelled from queue, -EINPROGRESS if the VCQ is + * left in-flight. On -EINPROGRESS, @orphan_cb(@orphan_data) will be + * called when the VCQ eventually completes (RH callback fires and the + * last sync_ctx ref drops). Use this to defer DMA cleanup. + * Safe for background/kthread callers that must not disrupt other MBX wor= k. + */ +int cmh_tm_submit_sync_noabort(struct vcq_cmd *vcq_cmds, u32 vcq_count, + u32 num_vcqs, unsigned long timeout_hz, + void (*orphan_cb)(void *), + void *orphan_data); + +/* + * Asynchronous submit -- post VCQs and return immediately. + * + * On successful return (0), the provided @callback may be invoked from + * either the RH threaded IRQ context (normal completion path) or the TM + * kthread (if VCQ dispatch to the HW ring fails after the message was + * posted to the CMQ). The caller must not assume a specific callback + * context. + * + * After a successful post, the caller must NOT touch VCQ buffers -- + * ownership transfers to the TM. If this function returns non-zero, + * the message was not posted, the callback will NOT fire, and the caller + * must perform cleanup. + * + * Uses GFP_ATOMIC internally -- the crypto API may invoke driver ops + * from softirq context (e.g. IPsec), so GFP_KERNEL would deadlock. + * + * If @backlog_ok is true and the CMQ is full, the message is placed on + * an overflow backlog queue and -EBUSY is returned. The caller must + * treat -EBUSY as "accepted" (like -EINPROGRESS): the callback WILL + * fire once the request is promoted from backlog and completes. When + * @backlog_ok is false, CMQ-full returns -EAGAIN (caller must clean up). + * + * Returns: 0 on successful post, -EBUSY (backlogged -- callback will + * fire), -ENOMEM, -EINVAL (bad vcq_count), -EAGAIN (CMQ full, + * no backlog), -ENODEV. + */ +int cmh_tm_submit_async(struct vcq_cmd *vcq_cmds, u32 vcq_count, + u32 num_vcqs, s32 target_mbx, + cmh_completion_fn callback, void *callback_data, + bool backlog_ok, unsigned long timeout_jiffies); + +/** + * cmh_tm_async_timeout_jiffies() - Default per-request async timeout + * + * Returns the debugfs-configurable timeout for symmetric data-path + * ops (async_timeout_ms converted to jiffies). Akcipher/kpp callers + * should pass 0 instead (no per-request timeout; vcq_timeout_ms is the + * safety net). + */ +unsigned long cmh_tm_async_timeout_jiffies(void); + +/** + * cmh_tm_flush_mbx() - Issue MBX_COMMAND_FLUSH and wait for completion + * @mbx_idx: Mailbox index + * + * Resets the eSW child mailbox state including the temp stack. + * Must be called when no VCQ submission is in progress on @mbx_idx. + * + * Return: 0 on success, -ETIMEDOUT if eSW does not clear the command, + * -EBUSY if a command is already pending. + */ +int cmh_tm_flush_mbx(s32 mbx_idx); + +/** + * cmh_tm_poke_tail() - Re-write R_MBX_QUEUE_TAIL to wake the eSW + * @base: Mailbox register base + * + * Serialised against the submit_vcq() doorbell so a re-poke cannot roll + * back a concurrent submission. Safe from softirq/timer context. + */ +void cmh_tm_poke_tail(void __iomem *base); + +/** + * cmh_tm_try_cancel_command() - Try to cancel a queued command + * @msg: Command message to cancel + * + * Return: true if removed from CMQ, false if already consumed by the TM t= hread. + */ +bool cmh_tm_try_cancel_command(struct command_msg *msg); + +/** + * cmh_tm_peek_transaction() - Peek at the oldest transaction on a mailbox + * @mbx_idx: Mailbox index + * + * For use by the Response Handler. Caller must hold txq->lock or call + * from a context where no concurrent pop is possible (e.g. threaded IRQ). + * + * Return: Pointer to the oldest transaction_obj, or NULL if empty. + */ +struct transaction_obj *cmh_tm_peek_transaction(u32 mbx_idx); + +/** + * cmh_tm_pop_transaction() - Remove and return the oldest transaction + * @mbx_idx: Mailbox index + * + * Return: Pointer to the removed transaction_obj, or NULL if empty. + */ +struct transaction_obj *cmh_tm_pop_transaction(u32 mbx_idx); + +/** + * cmh_txn_finish() - Complete a transaction with FSM + timer handling + * @txn: Transaction popped from the TXQ + * @error: Error code (0 for success, negative errno) + * + * Resolves the timer-vs-completion race via atomic cmpxchg, cancels + * the per-txn timeout timer if still pending, fires the completion + * callback (if this path wins the race), and drops the owner reference. + * The transaction is freed when the last reference is dropped. + * + * Called by the Response Handler after popping a completed transaction. + */ +void cmh_txn_finish(struct transaction_obj *txn, int error); + +/** + * cmh_tm_max_cmds_per_vcq() - Max vcq_cmd entries per MBX slot + * + * Returns the minimum across all configured MBXes so callers can pack + * VCQs without knowing which MBX will be selected. + * + * Return: At least MIN_VCQ_CMDS (2). + */ +u32 cmh_tm_max_cmds_per_vcq(void); + +/** + * cmh_tm_mbx_count() - Return the number of configured mailboxes + * + * Return: cfg->mbx_count. + */ +u32 cmh_tm_mbx_count(void); + +/** + * cmh_core_default_id() - Return the default core_id for a core type + * @type: Logical core type enum + * + * Returns the core_id of the first (index-0) instance without advancing + * the round-robin counter. Intended for callers pinned to a fixed MBX + * (e.g. mgmt ioctls on MGMT_MBX) that only need the VCQ core_id field. + * + * In multi-instance configurations the returned core_id is always that + * of instance[0], regardless of which MBX instance[0] is assigned to. + * Mgmt callers submit on MGMT_MBX (0) -- the eSW accepts any valid + * core_id on any MBX for command dispatch. + * + * Return: u32 core_id. + */ +u32 cmh_core_default_id(enum cmh_core_type type); + +/** + * cmh_core_select_instance() - Multi-instance core dispatch selection + * @type: Logical core type enum + * + * Returns the next (core_id, mbx_idx) pair for @type using round-robin + * across configured instances. On first use for an instance whose MBX + * is not pre-assigned, atomically assigns the next available MBX. + * + * With single-instance defaults, this degenerates to the same behaviour + * as the old single-entry core_to_mbx[] table -- one core type, one MBX. + * + * Return: struct core_dispatch with core_id and mbx_idx. + */ +struct core_dispatch cmh_core_select_instance(enum cmh_core_type type); + +/** + * cmh_core_num_instances() - Return count of configured instances + * @type: Logical core type enum + * + * Return: Number of instances for @type, or 0 if the core is absent from + * the CORE_ENABLE register (not present in this hardware build). + */ +u32 cmh_core_num_instances(enum cmh_core_type type); + +/** + * cmh_core_present() - Test whether a core type was discovered + * @type: Logical core type enum + * + * Return: true if at least one instance of @type is present in hardware, + * false if the core is absent. Used to gate crypto-API + * registration so an absent core advertises no algorithms. + */ +static inline bool cmh_core_present(enum cmh_core_type type) +{ + return cmh_core_num_instances(type) > 0; +} + +/** + * cmh_core_get_instance() - Get a specific instance by index + * @type: Logical core type enum + * @idx: Instance index (0-based, must be < cmh_core_num_instances()) + * + * Returns (core_id, mbx_idx) for the given instance without advancing + * the round-robin counter. Triggers auto-assign if the instance has + * no MBX yet. + * + * Return: struct core_dispatch with core_id and mbx_idx. + */ +struct core_dispatch cmh_core_get_instance(enum cmh_core_type type, u32 id= x); + +/** + * cmh_tm_affinity_reset() - Reset all core-to-MBX assignments + * + * Called during init and cleanup. + */ +void cmh_tm_affinity_reset(void); + +/** + * cmh_tm_txq_completion_notify() - Wake TM thread after TXQ completion + * + * Called by the Response Handler after completing a transaction to + * unblock the TM thread if it is waiting for a free MBX slot. + */ +void cmh_tm_txq_completion_notify(void); + +/* + * Pack @count payload commands (no headers) into one or more VCQs + * respecting the per-slot size limit, then submit synchronously. + * + * @payload: flat array of vcq_cmd entries (no headers) + * @count: number of entries in @payload + * @packed: caller-provided scratch buffer for the packed output + * @max_packed: size of @packed in vcq_cmd entries + * @target_mbx: -1 =3D round-robin, >=3D 0 =3D pin to this MBX index + * + * Each VCQ gets its own header. All VCQs are submitted as a single + * back-to-back transaction on the same MBX. + */ +int cmh_vcq_pack_and_submit(const struct vcq_cmd *payload, u32 count, + struct vcq_cmd *packed, u32 max_packed, + s32 target_mbx); + +/** + * cmh_vcq_pack_and_submit_async() - Pack payload commands and submit async + * @payload: Flat array of VCQ command entries (no headers) + * @count: Number of entries in @payload + * @packed: Caller-provided scratch buffer for packed output + * @max_packed: Size of @packed in vcq_cmd entries + * @target_mbx: Mailbox index (-1 for round-robin) + * @callback: Completion callback + * @callback_data: Opaque data passed to @callback + * @backlog_ok: If true, accept into backlog when CMQ is full + * @timeout_jiffies: Per-request timeout (0 to disable) + * + * Async variant of cmh_vcq_pack_and_submit(). Returns 0 on successful + * post; after a successful post, @callback may run from RH threaded IRQ + * context on normal completion, from the TM kthread if VCQ dispatch + * fails after posting, or from TM teardown paths such as + * cmh_tm_cleanup() / cmh_tm_quiesce() when queued or in-flight work is + * cancelled. Callers must not assume a single callback context. On + * non-zero return, the callback will NOT fire. + * + * @payload: flat array of vcq_cmd entries (no headers) + * @count: number of entries in @payload + * @packed: caller-provided scratch buffer for the packed output + * @max_packed: size of @packed in vcq_cmd entries + * @target_mbx: -1 =3D round-robin, >=3D 0 =3D pin to this MBX index + * @callback: completion callback (may run from IRQ or TM context) + * @callback_data: opaque pointer passed to @callback + * @backlog_ok: if true, queue the request when all MBXs are busy + * @timeout_jiffies: maximum wait time for MBX slot (0 =3D no wait) + * + * Return: 0 on successful post, -EBUSY (backlogged), negative errno on fa= ilure. + */ +int cmh_vcq_pack_and_submit_async(const struct vcq_cmd *payload, u32 coun= t, + struct vcq_cmd *packed, u32 max_packed, + s32 target_mbx, + cmh_completion_fn callback, + void *callback_data, + bool backlog_ok, + unsigned long timeout_jiffies); + +/* debugfs timeout accessors (debug builds only) */ +#ifdef CONFIG_CRYPTO_DEV_CMH_DEBUG +unsigned int *cmh_tm_timeout_async_ptr(void); +unsigned int *cmh_tm_timeout_vcq_ptr(void); +unsigned int *cmh_tm_timeout_slow_op_ptr(void); +unsigned int *cmh_tm_timeout_drain_ptr(void); +unsigned int *cmh_tm_cmq_max_depth_ptr(void); +unsigned int *cmh_tm_backlog_max_depth_ptr(void); +#endif + +/* -- Crypto request completion helper ---------------------------------- = */ + +struct device *cmh_dev(void); + +/** + * cmh_complete() - Complete a crypto request with optional error logging + * @req: The async crypto request to complete + * @err: Completion code: 0 =3D success, < 0 =3D errno, -EINPROGRESS =3D b= acklog + * promotion signal, > 0 =3D byte count (ahash BLOCK_ONLY leftover the + * crypto API must re-buffer) + * + * Logs a rate-limited diagnostic on genuine errors (err < 0), then hands + * the request back to the crypto framework. -EINPROGRESS and non-negative + * completions (success, or a positive ahash remainder) are not errors and + * are not logged. Centralizes error reporting so individual algorithm + * drivers do not need per-callback logging. + */ +static inline void cmh_complete(struct crypto_async_request *req, int err) +{ + if (err < 0 && err !=3D -EINPROGRESS) { + /* + * For template instances (e.g. hmac(sha3-512-cmh)) the + * driver name will be the outer template's, not ours. + * Still useful for triage -- identifies the failing tfm. + */ + dev_dbg_ratelimited(cmh_dev(), "op error: alg=3D%s err=3D%d\n", + crypto_tfm_alg_driver_name(req->tfm), + err); + } + crypto_request_complete(req, err); +} + +#endif /* CMH_TXN_H */ diff --git a/drivers/crypto/cmh/include/cmh_vcq.h b/drivers/crypto/cmh/incl= ude/cmh_vcq.h new file mode 100644 index 000000000000..8ebcbccd2aca --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_vcq.h @@ -0,0 +1,288 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- VCQ (Virtual Command Queue) Definitions + * + * Kernel-side definitions for the CMH VCQ and DMA scatter-gather ABI, + * so the LKM can build VCQs without depending on CMH eSW headers. + * + * All constants and layouts are derived from the CMH eSW ABI. + * + * This ABI is little-endian only: the VCQ command and descriptor fields + * use native integer types and the CMH block shares the little-endian SoC + * with its eSW. The driver depends on !CPU_BIG_ENDIAN (see the CMH + * Kconfig). + * + * Per-core command definitions live in their own ABI headers (cmh_hc_abi.= h, + * cmh_aes_abi.h, etc.) and are included here to form the hwc_cmd union. + */ + +#ifndef CMH_VCQ_H +#define CMH_VCQ_H + +#include +#include +#include +#include + +#include "cmh_hc_abi.h" +#include "cmh_sm3_abi.h" +#include "cmh_drbg_abi.h" +#include "cmh_sys_abi.h" +#include "cmh_kic_abi.h" +#include "cmh_aes_abi.h" +#include "cmh_sm4_abi.h" +#include "cmh_ccp_abi.h" +#include "cmh_pke_abi.h" +#include "cmh_qse_abi.h" +#include "cmh_hcq_abi.h" +#include "cmh_eac_abi.h" + +/* VCQ Magic Numbers */ + +#define VCQ_HDR_MAGIC 0x01514356U /* 'V' 'C' 'Q' 0x01 */ +#define VCQ_CMD_MAGIC 0x01444D43U /* 'C' 'M' 'D' 0x01 */ + +/* VCQ Command ID Encoding */ + +#define VCQ_CMD_MASK 0x000000FFU +#define VCQ_SPAN_MASK 0x0000FF00U +#define VCQ_FLAG_MASK 0x00FF0000U +#define VCQ_CORE_MASK 0xFF000000U + +#define VCQ_CMD_ID(core, flags, span, cmd) \ + (((u32)(core) << 24) | ((flags) & VCQ_FLAG_MASK) | \ + (((u32)(span) << 8) & VCQ_SPAN_MASK) | ((cmd) & VCQ_CMD_MASK)) + +/* Core IDs (per CMH hardware specification) */ + +#define CORE_ID_SYS 0x00U +#define CORE_ID_DMA 0x01U +#define CORE_ID_HC 0x02U +#define CORE_ID_AES 0x03U +#define CORE_ID_SM4 0x04U +#define CORE_ID_SM3 0x05U +#define CORE_ID_XC 0x07U +#define CORE_ID_HCQ 0x08U +#define CORE_ID_QSE 0x09U +#define CORE_ID_PKE 0x0AU +#define CORE_ID_TIC 0x0BU +#define CORE_ID_KIC 0x0CU +#define CORE_ID_MPU 0x0EU +#define CORE_ID_DRBG 0x0FU +#define CORE_ID_EMC 0x11U +#define CORE_ID_CCP 0x18U +#define CORE_ID_EAC 0x1EU +#define CORE_ID_NUM 0x1FU /* eSW g_drvs[] array size sentinel= */ +#define CORE_ID_MAX 0xFFU /* VCQ encoding limit (8-bit field)= */ + +/** + * enum cmh_core_type - Logical core type for multi-instance dispatch + * @CMH_CORE_HC: Hash / HMAC / CSHAKE / KMAC (CORE_ID_HC) + * @CMH_CORE_AES: AES (CORE_ID_AES) + * @CMH_CORE_SM4: SM4 (CORE_ID_SM4) + * @CMH_CORE_SM3: SM3 (CORE_ID_SM3) + * @CMH_CORE_CCP: ChaCha20 / Poly1305 (CORE_ID_CCP) + * @CMH_CORE_PKE: RSA / ECDSA / ECDH / EdDSA / SM2 (CORE_ID_PKE) + * @CMH_CORE_QSE: ML-KEM / ML-DSA (CORE_ID_QSE) + * @CMH_CORE_HCQ: SLH-DSA / LMS / XMSS (CORE_ID_HCQ) + * @CMH_NUM_CORE_TYPES: Number of core types (array sizing sentinel) + * + * Algorithm drivers use this enum (not raw CORE_ID_* constants) for + * MBX selection and VCQ dispatch. Each value indexes into a config + * table that maps to one or more (core_id, mbx) pairs. + * + * Raw CORE_ID_* defines remain for: + * - SYS_TYPE_SET() key-type tags in datastore operations + * - DT child node ``reg`` values (hardware core identity for config loo= kup) + * - Singleton system cores (SYS, KIC, DRBG, EAC) not in this enum + */ +enum cmh_core_type { + CMH_CORE_HC =3D 0, + CMH_CORE_AES, + CMH_CORE_SM4, + CMH_CORE_SM3, + CMH_CORE_CCP, + CMH_CORE_PKE, + CMH_CORE_QSE, + CMH_CORE_HCQ, + CMH_NUM_CORE_TYPES +}; + +/** + * struct core_dispatch - VCQ dispatch target returned by core selection + * @core_id: Hardware core ID to encode in VCQ_CMD_ID() + * @mbx_idx: Mailbox index to submit the VCQ to + */ +struct core_dispatch { + u32 core_id; + s32 mbx_idx; +}; + +/* Common VCQ Command (per CMH VCQ ABI) */ + +#define VCQ_CMD_FLUSH 0xFFU + +/** + * struct vcq_hdr - VCQ header occupying the first slot of every VCQ + * @cmds: Total number of commands including the header itself + * @rsvd: Reserved -- used internally by CMH eSW firmware + */ +struct vcq_hdr { + u32 cmds; + u32 rsvd[13]; +}; + +/* DMA Scatter-Gather Item (per CMH DMAC hardware specification) */ + +/** + * struct dma_scattergather_item - DMA scatter-gather descriptor node + * @lli: Next descriptor address (0 =3D end of list) + * @src: Source address for input particle + * @dst: Destination address for output particle + * @len: Particle length (low 32 bits used by hardware) + * + * Linked-list node walked by the DMAC hardware. @lli chains to the + * next item or is zero for end-of-list. + */ +struct dma_scattergather_item { + u64 lli; + u64 src; + u64 dst; + u64 len; +}; + +/* Unified HWC Command Union */ +/* + * Each per-core ABI header defines a union _cmd. + * Add new cores here as they are implemented. + */ + +union hwc_cmd { + struct vcq_hdr hdr; + union hc_cmd hc; + union sm3_cmd sm3; + union drbg_cmd drbg; + union sys_cmd sys; + union kic_cmd kic; + union aes_cmd aes; + union sm4_cmd sm4; + union ccp_cmd ccp; + union pke_cmd pke; + union qse_cmd qse; + union hcq_cmd hcq; + union eac_cmd eac; +}; + +/** + * struct vcq_cmd - Single VCQ command entry (always 64 bytes) + * @magic: VCQ_HDR_MAGIC for the header slot, VCQ_CMD_MAGIC for commands + * @id: Encoded command ID built via VCQ_CMD_ID(core, flags, span, cmd) + * @hwc: Per-core command payload union + */ +struct vcq_cmd { + u32 magic; + u32 id; + union hwc_cmd hwc; +}; + +static_assert(sizeof(struct vcq_cmd) =3D=3D 64, + "struct vcq_cmd must be exactly 64 bytes (one VCQ slot)"); + +/** + * vcq_set_header() - Write the standard VCQ header at slot[0] + * @slot: Pointer to the first VCQ slot + * @total_cmds: Total number of commands including the header + */ +static inline void vcq_set_header(struct vcq_cmd *slot, u32 total_cmds) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_HDR_MAGIC; + slot->id =3D VCQ_CMD_ID(CORE_ID_SYS, 0, 1, SYS_CMD_RUN); + slot->hwc.hdr.cmds =3D total_cmds; +} + +/* VCQ Command Limits */ + +#define MIN_VCQ_CMDS 2U /* header + at least one command */ +#define MAX_VCQ_CMDS 15U /* including the header */ +#define MAX_VCQ_SIZE (MAX_VCQ_CMDS * sizeof(struct vcq_cmd)) + +/** + * vcq_add_inline_data() - Pack inline data into consecutive VCQ slots + * @slot: Pointer to the command slot preceding the inline data + * @data: Source data to copy into subsequent slots + * @data_len: Length of @data in bytes + * + * Appends data starting at slot+1 and updates the span field in + * slot->id. The caller must ensure enough slots are reserved. + * + * Return: Total number of slots consumed (1 + inline slots). + */ +static inline u32 vcq_add_inline_data(struct vcq_cmd *slot, + const void *data, u32 data_len) +{ + u32 inline_slots, total_span; + + if (!data_len) + return 1; + + inline_slots =3D (data_len + sizeof(struct vcq_cmd) - 1) / + sizeof(struct vcq_cmd); + total_span =3D 1 + inline_slots; + + /* Zero the inline slots, then copy data */ + memset(slot + 1, 0, inline_slots * sizeof(struct vcq_cmd)); + memcpy(slot + 1, data, data_len); + + /* Update span in the command's id field */ + slot->id =3D (slot->id & ~VCQ_SPAN_MASK) | + (((u32)total_span << 8) & VCQ_SPAN_MASK); + + return total_span; +} + +/** + * vcq_add_flush() - Build a generic VCQ_CMD_FLUSH command + * @slot: Pointer to the VCQ slot to populate + * @core_id: Hardware core ID for the flush command + */ +static inline void vcq_add_flush(struct vcq_cmd *slot, u32 core_id) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, VCQ_CMD_FLUSH); +} + +/* Shared HC VCQ Builders -- used by hash, hmac, cshake, kmac drivers */ + +static inline void vcq_add_hc_init(struct vcq_cmd *slot, u32 core_id, + u32 algo) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, HC_CMD_INIT); + slot->hwc.hc.cmd_init.algo =3D algo; +} + +static inline void vcq_add_hc_final(struct vcq_cmd *slot, u32 core_id, + u64 digest_phys, u32 outlen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, HC_CMD_FINAL); + slot->hwc.hc.cmd_final.digest =3D digest_phys; + slot->hwc.hc.cmd_final.outlen =3D outlen; +} + +static inline void vcq_add_hc_gather(struct vcq_cmd *slot, u32 core_id, + u64 lista_phys, u32 sgcmd) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, HC_CMD_GATHER); + slot->hwc.hc.cmd_gather.lista =3D lista_phys; + slot->hwc.hc.cmd_gather.sgcmd =3D sgcmd; +} + +#endif /* CMH_VCQ_H */ --=20 2.43.7 From nobody Fri Sep 25 01:20:33 2026 Received: from BL2PR02CU003.outbound.protection.outlook.com (mail-eastusazon11021074.outbound.protection.outlook.com [52.101.52.74]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 67DBD4B1B38; Thu, 17 Sep 2026 22:59:42 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.52.74 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789686008; cv=fail; b=okbJezYTmZhmOzq7wE1DDJ0cvZsjOwGP8uop7Y75CwQQ3ETB/cLRlgIpNcqjpgKAogkZug8QrJntnBAbSMoYbkQuAP1m9qYqY+HAYYbiwdIJmefqjRrXiAWa86pk815OchokKgtYBmWpk1UJ2JkWgAwjF3kkREhJv/2tRZeczIA= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789686008; c=relaxed/simple; bh=CBqAEkKry/n5P+FEmtQe+yJPFy30XqUu0/orPfMW+44=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=Vw0bDc8uRqY+QZBmwLA9BQJE/7xQpbZyVe23AcX0rmxKGDaJnTMEn5A8IC28JLIzaIs0Jm/mkLLepZstCd6f4rhRDiLuSSgxa0kn+elZlrqpvWsc/9Paw5PSRDcupoWJjm8sjMig2MBYKdNdtr+np0a4XQfXYf5wuzAFiyyByhA= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com; spf=fail smtp.mailfrom=rambus.com; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b=uazxL8EZ; arc=fail smtp.client-ip=52.101.52.74 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=rambus.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b="uazxL8EZ" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=JZOZXd4ederkaM3olIk3Faev94H5o/h/e4OrBwBAmXw1Tld0KihpZPJl5MEDCJ2qBt1hFZhDevvWhC0gYXHR5OFKQSjT5DCHEqJv8fbmITp5fPOmD2WZyvqJ/8Q1Nc2BLtlT0ynJi1YUqg4Xw8vH3ZYxnjP+aqJxsSQu/YWe571If8q4/iOJjORnPYsMwVIDixy4zWPLRW00qDIBYswPnmkCE7LM6g8cZz73Ohk/BFTFqRE4bGGWkgTPJisTU1DagVzr9dKadZ5wjX2Stba1FRWI/4xPGicaqa5O0qT1rC8fTn5A5YoqqXFJTCxuhJqyZwNnLlGWu9uWJEdj+PQVSQ== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=F65vaCF2DFEcnnF9zk3JH6ZUM/HQ6L+KaUgA+98KEvY=; b=dOeAEAiZeRT6eaj3ZQfkQD3Za8Wo21hBdcPgATtJI8ddvfjTPLYFCYJ1/N5giv/O/liUPcat1az/99ZAdR7R5XId0jN6UW33Z+Ff9J35VR4sz9K75CySqPJzxQJBP7TwS0i9VGMzjQl+/yTsfN8/Lv7bDXHKqQQuLfNYd+qeADE+Tn2Y6aQYlRLPUBr82A9nGnj79UZBSWTJSDTkjcrfGQZteMTbD5NnAusDbdntw+lMZV8ioIeAO0INlXt3TxpA0zM/ajOAij3Hf5CZbvSL1S/WISIK+U0evjzMf7DuwE4q2mpJGQoPHZhagDVfOLgr1e83/qUan72fGqJ4pao98w== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 192.86.86.210) smtp.rcpttodomain=cryptography.com smtp.mailfrom=rambus.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=rambus.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=rambus.com; s=selector1; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=F65vaCF2DFEcnnF9zk3JH6ZUM/HQ6L+KaUgA+98KEvY=; b=uazxL8EZHisHzpzomMuF9bt6F6TDsrdPstjd0jJwK6T9lt85Cj5ukfxmk6B1J9lwfAKWEAmsP/PLstSMPIJ5ssUszD1rAEih1swrs+JoGQdNtA6nKJPwdCNZsR0j4VsWv4gyx+u4VeiT2178cLQnt1Kl8r8i4v4j5e3x6smO+dTl9TMwxUoqbcouJCg1tGJCWSS29CGhM/JrAW3aLhKiBHsClPfyApW78k5WsTAH3R3CREHuU2G/LVs9pSiRsJ0rhtZq/T8+Cb2lnbe12JNYJZjCcNmW+wLw+xzGfZVjma8BvYMyhZx/BJWU/GdvFs8ocHFIC4DX5q/7osS1re9hHQ== Received: from BN0PR10CA0010.namprd10.prod.outlook.com (2603:10b6:408:143::14) by SJ0PR04MB7791.namprd04.prod.outlook.com (2603:10b6:a03:3ac::11) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.406.12; Thu, 17 Sep 2026 22:59:33 +0000 Received: from BN2PEPF0000A800.namprd02.prod.outlook.com (2603:10b6:408:143:cafe::21) by BN0PR10CA0010.outlook.office365.com (2603:10b6:408:143::14) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.428.12 via Frontend Transport; Thu, 17 Sep 2026 22:59:33 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 192.86.86.210) smtp.mailfrom=rambus.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=rambus.com; Received-SPF: Pass (protection.outlook.com: domain of rambus.com designates 192.86.86.210 as permitted sender) receiver=protection.outlook.com; client-ip=192.86.86.210; helo=hqxsv-psmtppxy02.rambus.com; pr=C Received: from hqxsv-psmtppxy02.rambus.com (192.86.86.210) by BN2PEPF0000A800.mail.protection.outlook.com (10.167.245.167) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.451.8 via Frontend Transport; Thu, 17 Sep 2026 22:59:32 +0000 Received: from hqxsv-cmdev3-aousherovitch.rambus.com (hqn-lb-int-float.rambus.com [10.12.20.20]) by hqxsv-psmtppxy02.rambus.com (Postfix) with ESMTP id D46B9180175D; Thu, 17 Sep 2026 22:59:31 +0000 (UTC) From: Alex Ousherovitch To: Albert Ou , Alex Ousherovitch , Conor Dooley , "David S. Miller" , Herbert Xu , Jonathan Corbet , Krzysztof Kozlowski , Palmer Dabbelt , Paul Walmsley , Rob Herring , Saravanakrishnan Krishnamoorthy , Shuah Khan Cc: Alexandre Ghiti , devicetree@vger.kernel.org, Joel Wittenauer , linux-api@vger.kernel.org, linux-crypto@vger.kernel.org, linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-riscv@lists.infradead.org, Shuah Khan , Thi Nguyen Subject: [PATCH v5 03/19] crypto: cmh - add key provisioning and management Date: Thu, 17 Sep 2026 15:59:12 -0700 Message-ID: <20260917225929.2494111-4-aousherovitch@rambus.com> X-Mailer: git-send-email 2.43.7 In-Reply-To: <20260917225929.2494111-1-aousherovitch@rambus.com> References: <20260917225929.2494111-1-aousherovitch@rambus.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: BN2PEPF0000A800:EE_|SJ0PR04MB7791:EE_ X-MS-Office365-Filtering-Correlation-Id: f2868f2c-f489-4a45-7184-08df150f565b X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|23010399003|82310400026|1800799024|376014|7416014|36860700016|10067099003|3023799007|22082099003|18002099003|6133799003|5023799004|56012099006|11063799006|921020; X-Microsoft-Antispam-Message-Info: RFCyuuonRSYSKLimzxDr54MhGR2BUKpsHtyRQ2SklHklgibOjCINpvJimkqVG56N8Xxbi4YrTPpNaRFwA1lkHa+CAC2Y/onVk8r2clZ7+LBlBhjMg4xpSSa9BdnPT28WNXwCIGR5JCd1m1RNYOhtjqDl7P88+tQKa0kylL4bmNHZ087+GevhH6Yu6aEobBHgPxDXj4/A1NldJo9QzcHRZ6utOitWPZaxFOakvJFvNpn+cH7K8DM2wsTfn0KF6USf/sxP6vWjEFjSIkjgln4mK+MV3wEoweS4sSvK+HF5BOnQE12EWACPPirmVe/EHNRLFNI/AL06GTBD/qK/bV08F9sJaBg6cxasiml8JT21KccxBDM3eiU7VUmqj+vo3d/3pBjRYB+N1YX1eYeMV7CRVwGf5ca5Cki/aKI5d0C+iWq0SJo8PomnVaxEw6an66zgbMdff2gmEf5lv0P1zG7wmGCGqNu2+rWwyyXJHwbC+JHjcUpKV9lNurNckzjWDPSHBJkgUoMIHibaiFxNqHlpxtqvbh1WUR5Kewq8vNxjAi0ivCHp4rBMKGBbHdCOg09htWIybc7KyCSdsQMeKm2kH9EIRD/+GflEs0WkUmhlYFKdiDKerzzF/fA4ifo5mlIZXFHyqckwBz1oXaF7r40sVpArvRtjcYXsn3zS/CcEHasRaA79OXle9fUzpYa0f7du4/VmQUNWsqZZAJnPlL4eczc10C/6tJPaUGGPyHbBRzocCsM+wLVtApWCkblAy/xB X-Forefront-Antispam-Report: CIP:192.86.86.210;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:hqxsv-psmtppxy02.rambus.com;PTR:InfoDomainNonexistent;CAT:NONE;SFS:(13230040)(23010399003)(82310400026)(1800799024)(376014)(7416014)(36860700016)(10067099003)(3023799007)(22082099003)(18002099003)(6133799003)(5023799004)(56012099006)(11063799006)(921020);DIR:OUT;SFP:1102; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: On9cYtXri5Nq7v9pjZ9Ucsw0AsO0CCiCE2SXj5cueWx3TNSiIN69zWcgCyu1kiX8tZ+EkuRRi7nVn/ffJCyssm/LwkFTaXpMei+syNGzq6O+b5eRDatDD0yeXd4nTK/XPbVegqllf4CCNOcWYtLr3TD3sTclS35CzqDK7DdX7kDWHV9jrGVjYD4TvGJCMVoO9hrfMiaoZHaWk0anWLLIq+duk24SIFnHqdKGhOkBgGPwXrNt1eWEhbPpnNcI294L0qwASLGjzHCAdJJxwSc5Vz+ljrNDq5+8C+yOP6eiEHYmnKAguc0krqv/Od8kLkcJo+2KAW/coTngNAHIAV6XLXfE80tcRT6tC86esh3BKkXP7mPkKD6YTcdBW6rFiayg8nAk3YhshQZWhUmbCQTVdL7QRNTmz0AXGsTYDL3lBlQiVmkVioF1M0ZrjZFLX+G/ X-OriginatorOrg: rambus.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 17 Sep 2026 22:59:32.5116 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: f2868f2c-f489-4a45-7184-08df150f565b X-MS-Exchange-CrossTenant-Id: bd0ba799-c2b9-413c-9c56-5d1731c4827c X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=bd0ba799-c2b9-413c-9c56-5d1731c4827c;Ip=[192.86.86.210];Helo=[hqxsv-psmtppxy02.rambus.com] X-MS-Exchange-CrossTenant-AuthSource: BN2PEPF0000A800.namprd02.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: SJ0PR04MB7791 Content-Type: text/plain; charset="utf-8" The in-kernel crypto API cannot express the CMH management plane, so this patch adds a character device (/dev/cmh_mgmt) for it. The ioctl commands fall into two groups; the overlap with the transforms the driver also registers is deliberate and bounded: (1) Operations with no crypto API representation - the majority. The API has no transform type or verb for these, so a character device is the only available UAPI: - hardware key lifecycle: create, import, export, derive, destroy, enumerate (keystore CRUD) - no keystore verb - KIC key derivation (HKDF, AES-CMAC-KDF, DKEK) - asymmetric key generation (RSA, EC, EdDSA, ML-DSA, SLH-DSA) and public-key derivation - the API has no keygen verb - ML-KEM encapsulate/decapsulate - no kernel KEM API exists - SM2 encrypt/decrypt and key exchange (multi-step GM/T 0003) - EdDSA sign/verify - not registered with the crypto API - EAC Chip Authentication and DRBG (re)configuration (2) Hardware-held-key operations on algorithms that ARE also registered (RSA decrypt, ECDSA/ML-DSA/SLH-DSA sign, ECDH). These name the same primitives as the registered akcipher/sig/kpp transforms, but the crypto API set_priv_key()/set_secret() accept only raw key bytes supplied by the caller; they cannot reference a private key that is generated inside, and never leaves, the hardware datastore - the central security property of this device. The ioctl path keeps the private key hardware-resident, while the registered transforms serve raw-key in-kernel users. The two paths are complementary, not redundant. The subsystem provides: - Key provisioning: create, import, derive, and destroy hardware keys stored in the CMH datastore - System object management: allocate and free CMH system objects - Management ioctl interface (/dev/cmh_mgmt): key lifecycle, KIC key derivation, PKE (RSA, ECDSA, ECDH, EdDSA), PQC (ML-KEM, ML-DSA, SLH-DSA), SM2, EAC, and DRBG reseeding - SM2 ioctl handlers: encrypt, decrypt, sign, and key exchange -- multi-step protocol flows not expressible through the crypto API sig interface - UAPI header: cmh_mgmt_ioctl.h (ioctl definitions and structures) The device requires CAP_SYS_ADMIN for open() and is built conditionally on CONFIG_CRYPTO_DEV_CMH_MGMT (default n); when disabled the ioctl interface is absent while all kernel crypto API algorithms remain registered. Signed-off-by: Alex Ousherovitch Co-developed-by: Saravanakrishnan Krishnamoorthy Signed-off-by: Saravanakrishnan Krishnamoorthy --- Documentation/ABI/testing/cmh-mgmt | 136 ++ drivers/crypto/cmh/Kconfig | 19 + drivers/crypto/cmh/Makefile | 11 +- drivers/crypto/cmh/cmh_key.c | 164 ++ drivers/crypto/cmh/cmh_main.c | 9 + drivers/crypto/cmh/cmh_mgmt.c | 1770 ++++++++++++++++++++++ drivers/crypto/cmh/cmh_mgmt_pke.c | 1143 ++++++++++++++ drivers/crypto/cmh/cmh_mgmt_pqc.c | 1365 +++++++++++++++++ drivers/crypto/cmh/cmh_pke_sm2.c | 862 +++++++++++ drivers/crypto/cmh/cmh_sys.c | 376 +++++ drivers/crypto/cmh/include/cmh_key.h | 82 + drivers/crypto/cmh/include/cmh_mgmt.h | 69 + drivers/crypto/cmh/include/cmh_pke.h | 245 +++ drivers/crypto/cmh/include/cmh_pke_sm2.h | 30 + drivers/crypto/cmh/include/cmh_pqc.h | 25 + drivers/crypto/cmh/include/cmh_sys.h | 111 ++ include/uapi/linux/cmh_mgmt_ioctl.h | 902 +++++++++++ 17 files changed, 7318 insertions(+), 1 deletion(-) create mode 100644 Documentation/ABI/testing/cmh-mgmt create mode 100644 drivers/crypto/cmh/cmh_key.c create mode 100644 drivers/crypto/cmh/cmh_mgmt.c create mode 100644 drivers/crypto/cmh/cmh_mgmt_pke.c create mode 100644 drivers/crypto/cmh/cmh_mgmt_pqc.c create mode 100644 drivers/crypto/cmh/cmh_pke_sm2.c create mode 100644 drivers/crypto/cmh/cmh_sys.c create mode 100644 drivers/crypto/cmh/include/cmh_key.h create mode 100644 drivers/crypto/cmh/include/cmh_mgmt.h create mode 100644 drivers/crypto/cmh/include/cmh_pke.h create mode 100644 drivers/crypto/cmh/include/cmh_pke_sm2.h create mode 100644 drivers/crypto/cmh/include/cmh_pqc.h create mode 100644 drivers/crypto/cmh/include/cmh_sys.h create mode 100644 include/uapi/linux/cmh_mgmt_ioctl.h diff --git a/Documentation/ABI/testing/cmh-mgmt b/Documentation/ABI/testing= /cmh-mgmt new file mode 100644 index 000000000000..7da2630edb20 --- /dev/null +++ b/Documentation/ABI/testing/cmh-mgmt @@ -0,0 +1,136 @@ +What: /dev/cmh_mgmt +Date: June 2026 +KernelVersion: 7.1 +Contact: linux-crypto@vger.kernel.org +Description: + Character device (misc) providing a management and + key-operations ioctl interface to the Rambus CryptoManager Hub + hardware crypto accelerator. Used for operations that + cannot be represented through the standard in-kernel + crypto API: datastore key CRUD, key derivation, + asymmetric crypto (EdDSA, SM2), and post-quantum crypto + (ML-KEM, ML-DSA, SLH-DSA). + + The ioctl magic is 'J'. All struct arguments are + versioned via a leading __u32 version field set to + CMH_MGMT_V1 (1). + + Ioctl commands are grouped by function: + + **Key Management (0x01-0x0E):** + + - CMH_IOCTL_KEY_NEW (0x01): Allocate a new datastore slot. + Accepts ds_type (CMH_DS_* constant), key length, flags, + and caller ID. Returns a 64-bit key reference. + + - CMH_IOCTL_KEY_WRITE (0x02): Write key material into + a previously allocated datastore slot. Supports + plaintext or wrapped key import via a wrapping key ref. + + - CMH_IOCTL_KEY_READ (0x03): Read key material from + a datastore slot, optionally wrapped. Returns data + plus a 16-byte SYS header for wrapped reads. + + - CMH_IOCTL_KEY_FIND (0x04): Look up a key reference + by caller ID (CID). + + - CMH_IOCTL_KEY_GRANT (0x05): Grant access to a key. + + - CMH_IOCTL_KEY_DELETE (0x06): Delete a datastore slot. + + - CMH_IOCTL_DS_EXPORT (0x07): Export the entire datastore + as an encrypted blob. + + - CMH_IOCTL_DS_IMPORT (0x08): Import a previously exported + datastore blob. + + - CMH_IOCTL_KEY_NEW_RANDOM (0x0B): Allocate a datastore + slot and fill it with hardware-generated random data. + + - CMH_IOCTL_KEY_LIST (0x0E): List active datastore entries, + returning CIDs, types, lengths, and flags. + + **Key Derivation -- KIC (0x09-0x0D):** + + - CMH_IOCTL_KIC_HKDF1 (0x09): HKDF-Extract step. + - CMH_IOCTL_KIC_HKDF2 (0x0A): HKDF-Expand step. + - CMH_IOCTL_KIC_AES_CMAC_KDF (0x0C): AES-CMAC KDF. + - CMH_IOCTL_KIC_DKEK_DERIVE (0x0D): DKEK derivation. + + **EAC -- Error and Alarm (0x0F):** + + - CMH_IOCTL_EAC_READ (0x0F): Read and clear hardware + error, alarm, and safety notification registers. + + **PKE -- Public Key Engine (0x10-0x1C):** + + - CMH_IOCTL_PKE_RSA_ENC (0x10): RSA public-key encrypt. + - CMH_IOCTL_PKE_RSA_DEC (0x11): RSA private-key decrypt. + - CMH_IOCTL_PKE_RSA_CRT_DEC (0x12): RSA-CRT decrypt. + - CMH_IOCTL_PKE_RSA_KEYGEN (0x13): RSA key pair generation. + - CMH_IOCTL_PKE_ECDSA_SIGN (0x14): ECDSA sign. + - CMH_IOCTL_PKE_ECDH (0x16): ECDH shared secret. + - CMH_IOCTL_PKE_ECDH_KEYGEN (0x17): ECDH key pair generation. + - CMH_IOCTL_PKE_EDDSA_SIGN (0x18): EdDSA sign (Ed25519/Ed448). + - CMH_IOCTL_PKE_EDDSA_VERIFY (0x19): EdDSA verify. + - CMH_IOCTL_PKE_EC_KEYGEN (0x1A): EC key pair generation. + - CMH_IOCTL_PKE_EC_PUBGEN (0x1B): EC public key derivation. + - CMH_IOCTL_PKE_EDDSA_KEYGEN_SCA (0x1C): EdDSA SCA-protected + key generation. + + **PQC -- Post-Quantum Crypto (0x20-0x2D):** + + - CMH_IOCTL_ML_KEM_KEYGEN (0x20): ML-KEM key pair generation + (modes 512/768/1024). + - CMH_IOCTL_ML_KEM_ENC (0x21): ML-KEM encapsulation. + - CMH_IOCTL_ML_KEM_DEC (0x22): ML-KEM decapsulation. + - CMH_IOCTL_ML_DSA_KEYGEN (0x23): ML-DSA key pair generation + (modes 44/65/87). + - CMH_IOCTL_ML_DSA_SIGN (0x24): ML-DSA sign. + - CMH_IOCTL_SLHDSA_KEYGEN (0x28): SLH-DSA key pair generation + (12 parameter sets). + - CMH_IOCTL_SLHDSA_SIGN (0x29): SLH-DSA sign. + - CMH_IOCTL_SLHDSA_SIGN_PREHASH (0x2D): SLH-DSA prehash sign. + + **SM2 Operations (0x30-0x37):** + + - CMH_IOCTL_SM2_ECDH_KEYGEN (0x30): SM2 ephemeral key gen. + - CMH_IOCTL_SM2_ECDH (0x31): SM2 key exchange. + - CMH_IOCTL_SM2_DEC_POINT (0x32): SM2 decrypt (point step). + - CMH_IOCTL_SM2_ENC_POINT (0x33): SM2 encrypt (point step). + - CMH_IOCTL_SM2_ID_DIGEST (0x34): SM2 ID digest (ZA). + - CMH_IOCTL_SM2_ECDH_HASH (0x35): SM2 key exchange hash step. + - CMH_IOCTL_SM2_DEC_HASH (0x36): SM2 decrypt (hash step). + - CMH_IOCTL_SM2_ENC_HASH (0x37): SM2 encrypt (hash step). + + The SM2 encrypt/decrypt hash-step ioctls accept payloads + of at most 32 bytes. The underlying hardware KDF emits a + single 32-byte SM3 block, so longer messages cannot be + processed in a single command and are rejected with + -EINVAL. + + **DRBG Management (0x40):** + + - CMH_IOCTL_DRBG_CONFIG (0x40): Configure the hardware + DRBG entropy ratio and security strength. Normally + called once at system start-up before hwrng reads. + + All structs contain ``__reserved`` fields that must be + zero; the driver returns ``-EINVAL`` if any reserved field + is non-zero. This ensures forward compatibility when + reserved fields gain meaning in future versions. + + All ioctls return 0 on success or a negative errno on + failure. Common errors: + + - EINVAL: Invalid version, parameter, key type, or + non-zero reserved field. + - ENOENT: Key reference not found in datastore. + - ENOMEM: DMA allocation failure. + - EBUSY: Hardware mailbox full. + - ETIMEDOUT: VCQ operation timed out. + - EFAULT: Bad user-space pointer. + + The ioctl UAPI header is . + All structures, constants, and type definitions are + documented in that header file. diff --git a/drivers/crypto/cmh/Kconfig b/drivers/crypto/cmh/Kconfig index fca66d5e2f89..512e7b3bdc59 100644 --- a/drivers/crypto/cmh/Kconfig +++ b/drivers/crypto/cmh/Kconfig @@ -63,3 +63,22 @@ config CRYPTO_DEV_CMH_DEBUG Useful for bringup, validation, and performance analysis. Not recommended for production. =20 + +config CRYPTO_DEV_CMH_MGMT + bool "CMH management ioctl device (/dev/cmh_mgmt)" + depends on CRYPTO_DEV_CMH + default n + help + Expose /dev/cmh_mgmt, a misc device providing ioctl commands + for operations that have no kernel crypto API binding: hardware + key lifecycle (create, import, derive, destroy), KIC key + derivation, PQC keygen/encaps/decaps (ML-KEM, ML-DSA, SLH-DSA), + EdDSA sign/verify, SM2 key exchange, and DRBG + configuration. + + The device requires CAP_SYS_ADMIN. Disabling this option + removes the ioctl interface but all kernel crypto API + algorithms (consumed by in-kernel users and validated by the + crypto test manager) remain fully functional. + + If unsure, say N. diff --git a/drivers/crypto/cmh/Makefile b/drivers/crypto/cmh/Makefile index 742e65e3917e..bb7772555d6a 100644 --- a/drivers/crypto/cmh/Makefile +++ b/drivers/crypto/cmh/Makefile @@ -12,7 +12,16 @@ cmh-y :=3D \ cmh_txn.o \ cmh_rh.o \ cmh_dma.o \ - cmh_sysfs.o + cmh_sysfs.o \ + cmh_key.o \ + cmh_sys.o + +# Management ioctl device (/dev/cmh_mgmt): key lifecycle, PKE, PQC ioctls. +cmh-$(CONFIG_CRYPTO_DEV_CMH_MGMT) +=3D \ + cmh_mgmt.o \ + cmh_mgmt_pke.o \ + cmh_mgmt_pqc.o \ + cmh_pke_sm2.o =20 ccflags-y +=3D -I$(src)/include =20 diff --git a/drivers/crypto/cmh/cmh_key.c b/drivers/crypto/cmh/cmh_key.c new file mode 100644 index 000000000000..fde8be50b25c --- /dev/null +++ b/drivers/crypto/cmh/cmh_key.c @@ -0,0 +1,164 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Dual Key Path Implementation + * + * Two key provisioning paths are supported: + * + * Raw key: key bytes -> stored in tfm context -> + * SYS_CMD_WRITE(SYS_REF_TEMP) packed into every crypto VCQ. + * The raw key buffer is DMA-mapped once at setkey time and remains + * mapped for the lifetime of the transform (unmapped in destroy). + * + * Raw key DMA lifetime rationale + * ------------------------------ + * Raw keys are DMA-mapped at setkey time and the mapping persists + * until the transform is destroyed (cmh_key_destroy). This is a + * deliberate design choice, consistent with upstream HW crypto + * drivers (CAAM, ccree, CCP) that also map keys at setkey for + * transform-lifetime reuse: + * + * - The Linux crypto framework expects setkey to prepare the + * transform for repeated encrypt/decrypt calls. Remapping the + * same key on every request would add DMA API overhead per crypto + * operation with no security benefit. + * - On destroy, kfree_sensitive() scrubs the key buffer and the + * DMA mapping is released. For key-by-ID (persistent), the + * per-MBX ref cache is zeroed with memzero_explicit(). + * - No key material is ever logged; dev_dbg() messages only show + * CIDs (content identifiers), not key bytes. + * + * Hardware-required behaviors (not driver policy) + * ------------------------------------------------ + * - SYS_REF_TEMP lifetime: the eSW firmware reclaims temporary + * datastore objects when the mailbox slot is reused. This is a + * hardware constraint; the driver packs SYS_CMD_WRITE into every + * VCQ to re-provision the raw key for each operation. + * - Mailbox flush (SYS_CMD_FLUSH): reclaims temp-stack space on the + * target MBX. Required by HW to prevent temp-stack exhaustion + * across multi-VCQ operations. + */ + +#include +#include +#include + +#include "cmh_key.h" +#include "cmh_sys.h" +#include "cmh_txn.h" +#include "cmh_dma.h" +#include "cmh_sys_abi.h" +#include + +/** + * cmh_ds_type_to_core_id() - Map a datastore type to a logical core ID + * @ds_type: Datastore type constant (e.g. CMH_DS_AES_KEY, CMH_DS_SM4_KEY) + * + * Returns the algorithm-family identity (e.g. CORE_ID_AES =3D 0x03), NOT = the + * VCQ dispatch core_id. With multi-instance, a second AES engine dispatc= hes + * at CORE_ID_AES2 (0x06) but keys are still tagged with CORE_ID_AES (0x03) + * -- the eSW validates against the logical identity, not the dispatch ID. + * + * Return: Logical core ID on success, CORE_ID_NUM for unknown @ds_type. + */ +u32 cmh_ds_type_to_core_id(u32 ds_type) +{ + switch (ds_type) { + case CMH_DS_AES_KEY: + case CMH_DS_AES_XTS_KEY: + return CORE_ID_AES; + case CMH_DS_SM4_KEY: + return CORE_ID_SM4; + case CMH_DS_HMAC_KEY: + case CMH_DS_KMAC_KEY: + return CORE_ID_HC; + case CMH_DS_CHACHA20_KEY: + return CORE_ID_CCP; + case CMH_DS_RSA_PRIV_KEY: + case CMH_DS_RSA_PUB_KEY: + case CMH_DS_RSA_CRT_KEY: + case CMH_DS_ECDSA_PRIV_KEY: + case CMH_DS_ECDSA_PUB_KEY: + case CMH_DS_ECDH_PRIV_KEY: + case CMH_DS_EDDSA_PRIV_KEY: + case CMH_DS_SHARED_SECRET: + case CMH_DS_SM2_PRIV_KEY: + return CORE_ID_PKE; + case CMH_DS_ML_KEM_DK: + case CMH_DS_ML_DSA_SK: + return CORE_ID_QSE; + case CMH_DS_SLHDSA_SK: + return CORE_ID_HCQ; + default: + return CORE_ID_NUM; + } +} + +/** + * cmh_key_setkey_raw() - Store a raw key in the key context + * @ctx: Key context to populate + * @key: Pointer to the raw key bytes + * @keylen: Length of @key in bytes + * @core_id: Logical core ID for SYS_TYPE tagging + * + * Duplicates the raw key, DMA-maps the copy for the lifetime of the + * transform, and stores the mapping in @ctx. Any previously held key + * is destroyed first. + * + * The DMA mapping persists until cmh_key_destroy() is called (typically + * from the algorithm .exit_tfm callback). This avoids per-request DMA + * mapping overhead and matches the setkey-to-destroy lifetime model used + * by other upstream HW crypto drivers (CAAM, ccree, CCP). The key + * buffer is freed via kfree_sensitive() on destroy. + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_key_setkey_raw(struct cmh_key_ctx *ctx, const u8 *key, + u32 keylen, u32 core_id) +{ + dma_addr_t dma; + u8 *copy; + + if (!keylen || !key) + return -EINVAL; + + copy =3D kmemdup(key, keylen, GFP_KERNEL); + if (!copy) + return -ENOMEM; + + /* Pre-map for the lifetime of the transform */ + dma =3D cmh_dma_map_single(copy, keylen, DMA_TO_DEVICE); + if (cmh_dma_map_error(dma)) { + kfree_sensitive(copy); + return -ENOMEM; + } + + /* Clean up any previous key */ + cmh_key_destroy(ctx); + + ctx->mode =3D CMH_KEY_RAW; + ctx->raw.data =3D copy; + ctx->raw.len =3D keylen; + ctx->raw.dma =3D dma; + ctx->raw.sys_type =3D SYS_TYPE_SET(SYS_TYPE_FLAG_PT, core_id); + + return 0; +} + +/** + * cmh_key_destroy() - Destroy and zero-fill a key context + * @ctx: Key context to destroy + * + * For raw keys, unmaps the DMA buffer and securely frees the key material. + * Resets the key mode to CMH_KEY_NONE. + */ +void cmh_key_destroy(struct cmh_key_ctx *ctx) +{ + if (ctx->mode =3D=3D CMH_KEY_RAW && ctx->raw.data) { + cmh_dma_unmap_single(ctx->raw.dma, ctx->raw.len, + DMA_TO_DEVICE); + kfree_sensitive(ctx->raw.data); + memzero_explicit(&ctx->raw, sizeof(ctx->raw)); + } + ctx->mode =3D CMH_KEY_NONE; +} diff --git a/drivers/crypto/cmh/cmh_main.c b/drivers/crypto/cmh/cmh_main.c index c037e8daa877..d7753c4630ba 100644 --- a/drivers/crypto/cmh/cmh_main.c +++ b/drivers/crypto/cmh/cmh_main.c @@ -31,6 +31,7 @@ #include "cmh_mqi.h" #include "cmh_txn.h" #include "cmh_rh.h" +#include "cmh_mgmt.h" #include "cmh_registers.h" #include "cmh_debugfs.h" #include "cmh_sysfs.h" @@ -201,10 +202,17 @@ static int cmh_probe(struct platform_device *pdev) if (ret) goto err_rh_init; =20 + /* Register key management device (/dev/cmh_mgmt) */ + ret =3D cmh_mgmt_register(); + if (ret) + goto err_mgmt_register; + platform_set_drvdata(pdev, dev); =20 return 0; =20 +err_mgmt_register: + cmh_rh_cleanup(cfg); err_rh_init: cmh_tm_cleanup(); err_tm_init: @@ -229,6 +237,7 @@ static void cmh_remove(struct platform_device *pdev) =20 cfg =3D &dev->config; =20 + cmh_mgmt_unregister(); cmh_rh_cleanup(cfg); cmh_tm_cleanup(); cmh_mqi_cleanup(cfg); diff --git a/drivers/crypto/cmh/cmh_mgmt.c b/drivers/crypto/cmh/cmh_mgmt.c new file mode 100644 index 000000000000..32719b9f1f6e --- /dev/null +++ b/drivers/crypto/cmh/cmh_mgmt.c @@ -0,0 +1,1770 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Key Management misc_device (/dev/cmh_mgmt) + * + * Provides ioctl interface for key provisioning (NEW, NEW_RANDOM, WRITE, = READ, + * FIND, GRANT, DELETE) and datastore lifecycle (EXPORT, IMPORT). + * + * Each ioctl handler: copy_from_user -> validate -> DMA alloc -> + * build VCQ -> cmh_tm_submit_sync -> copy_to_user -> DMA free. + * + * Access requires CAP_SYS_ADMIN (checked in open()). The device node + * is mode 0660; DAC further limits access to owner/group. + * CMH eSW enforces per-MBX access control on top of this. + */ + +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "cmh_mgmt.h" +#include "cmh_sys.h" +#include "cmh_txn.h" +#include "cmh_key.h" +#include "cmh_dma.h" +#include "cmh_config.h" +#include "cmh_sys_abi.h" +#include "cmh_pke.h" +#include "cmh_pke_sm2.h" +#include "cmh_qse_abi.h" +#include "cmh_hcq_abi.h" +#include + +#include + +/* + * Pin all mgmt ioctls to a single management mailbox (MBX 0). + * + * This is a deliberate, structural choice -- not a performance default. + * The /dev/cmh_mgmt path is *stateful* with respect to the eSW datastore, + * and that state is per-mailbox, so every step of a key's lifecycle must + * land on the same mailbox: + * + * 1. Datastore access control is per-mailbox AND opaque to the driver. + * SYS_CMD_NEW grants the creating mailbox a (1 << mbx_id) access mask + * (read/write/execute). Crucially, the returned 64-bit ref encodes a + * randomised offset -- NOT the owning mailbox -- so given only a ref + * (as KEY_GRANT/READ/DELETE/DS_EXPORT receive), the driver cannot + * recover which mailbox owns the object. A fixed management mailbox + * is therefore the only way to guarantee that NEW, WRITE, GRANT, READ + * and the subsequent hardware-held-key compute ops all share the + * mailbox that holds the access rights, without exposing mailbox + * identity in the UABI. (User space may still widen access to other + * mailboxes explicitly via KEY_GRANT.) + * + * 2. The eSW SYS_REF_TEMP scratch store is per-mailbox and persists + * across ioctl calls. A derivation that writes SYS_REF_TEMP (e.g. a + * KIC_* derive) must be consumed by a later ioctl on the *same* + * mailbox (e.g. DS_EXPORT with wrap_key=3DSYS_REF_TEMP). + * + * Device-tree per-mailbox ``rambus,cores`` affinity applies to the *state= less* + * registered crypto API path (cmh_core_select_instance()), which carries + * no datastore state across calls and is free to balance across mailboxes. + * + * Note: MBX 0 is NOT reserved exclusively for mgmt -- registered crypto + * operations may also land here via TM round-robin (target_mbx =3D -1). + * This is safe because those ops do not allocate from the temp store. + */ + +/* VCQ layout: header + command + flush =3D 3 entries */ +#define MGMT_VCQ_CMDS 3 + +/* + * Tracks whether any operation has left residual state in the device's + * per-mailbox temporary key store since the last flush. The device + * reclaims temp storage only on a full mailbox flush (MBX_COMMAND_FLUSH), + * which also terminates any executing command queue with -EPIPE. + * + * To avoid killing concurrent in-flight operations, the flush in + * cmh_mgmt_ioctl() is conditional: it fires only when this flag is set. + * Operations that allocate temp storage (currently: KIC derivations + * targeting SYS_REF_TEMP) set this flag on success. + */ +static atomic_t mgmt_temp_dirty =3D ATOMIC_INIT(0); + +/* + * Serialise the entire management plane. SYS_REF_TEMP, the temp-dirty + * flush bookkeeping and the datastore are a single shared hardware + * resource, so ioctls -- even from different opens of /dev/cmh_mgmt -- + * must not interleave. + */ +static DEFINE_MUTEX(cmh_mgmt_lock); + +/** + * cmh_mgmt_ds_scrub() - Scrub orphaned datastore objects from a failed io= ctl + * @ref0: first datastore object ref to scrub, or 0 for none + * @ref1: second datastore object ref to scrub, or 0 for none + * + * Grant-with-no-access wipes the key material and clears the CID (the + * datastore stack space is only reclaimed by a full reset). Used on error + * paths where a SYS_CMD_NEW object's ref never reached user space, or a l= ater + * command in the same submission failed. Best-effort: a ref that was nev= er + * created returns -ENOENT. + */ +void cmh_mgmt_ds_scrub(u64 ref0, u64 ref1) +{ + struct vcq_cmd vcq[4]; + u32 idx =3D 0; + + if (!ref0 && !ref1) + return; + + vcq_set_header(&vcq[idx++], 2 + !!ref0 + !!ref1); + if (ref0) + vcq_add_sys_grant(&vcq[idx++], ref0, 0, 0, 0); + if (ref1) + vcq_add_sys_grant(&vcq[idx++], ref1, 0, 0, 0); + vcq_add_sys_flush(&vcq[idx++]); + cmh_tm_submit_sync_mbx(vcq, idx, 1, MGMT_MBX); +} + +/* -- KEY_NEW -------------------------- */ + +static int cmh_mgmt_key_new(void __user *argp) +{ + struct cmh_ioctl_key_new req; + struct vcq_cmd vcq[MGMT_VCQ_CMDS]; + u64 *ref_buf; + dma_addr_t ref_dma; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (!req.len || req.len > CMH_MGMT_MAX_DATA_LEN) + return -EINVAL; + + /* DMA buffer for CMH eSW to write back the ref */ + ref_buf =3D kmalloc_obj(*ref_buf, GFP_KERNEL); + if (!ref_buf) + return -ENOMEM; + + *ref_buf =3D 0; + ref_dma =3D cmh_dma_map_single(ref_buf, sizeof(*ref_buf), + DMA_FROM_DEVICE); + if (cmh_dma_map_error(ref_dma)) { + kfree(ref_buf); + return -ENOMEM; + } + + vcq_set_header(&vcq[0], MGMT_VCQ_CMDS); + vcq_add_sys_new(&vcq[1], req.cid, ref_dma, req.len); + vcq_add_sys_flush(&vcq[2]); + + ret =3D cmh_tm_submit_sync_mbx(vcq, MGMT_VCQ_CMDS, 1, MGMT_MBX); + + /* + * Unmap before CPU read: single-phase operation (no re-use of + * the DMA mapping), so unmap transfers ownership back to the + * CPU. On SWIOTLB systems the unmap copies the bounce buffer + * to the original allocation. This is the correct pattern for + * single-shot sync submits where the buffer is not re-mapped. + */ + cmh_dma_unmap_single(ref_dma, sizeof(*ref_buf), DMA_FROM_DEVICE); + + if (ret) { + /* Scrub a slot sys_new may have created before the failure. */ + if (*ref_buf) + cmh_mgmt_ds_scrub(*ref_buf, 0); + kfree(ref_buf); + return ret; + } + + req.ref =3D *ref_buf; + kfree(ref_buf); + + if (copy_to_user(argp, &req, sizeof(req))) { + /* + * The DS slot was created but its ref never reached + * userspace, so the caller cannot free it. Logically + * delete the orphaned slot before returning. + */ + cmh_mgmt_ds_scrub(req.ref, 0); + dev_warn(cmh_dev(), + "mgmt: KEY_NEW copy_to_user failed, DS slot cleaned up\n"); + return -EFAULT; + } + + dev_dbg(cmh_dev(), "mgmt: KEY_NEW cid=3D0x%llx len=3D%u -> ref=3D0x%llx\n= ", + req.cid, req.len, req.ref); + return 0; +} + +/* -- KEY_WRITE ------------------------- */ + +static int cmh_mgmt_key_write(void __user *argp) +{ + struct cmh_ioctl_key_write req; + struct vcq_cmd vcq[MGMT_VCQ_CMDS]; + void *dmabuf; + dma_addr_t dma_addr; + u32 core_id, sys_type; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (!req.len || req.len > CMH_MGMT_MAX_DATA_LEN) + return -EINVAL; + + core_id =3D cmh_ds_type_to_core_id(req.ds_type); + if (core_id =3D=3D CORE_ID_NUM) + return -EINVAL; + sys_type =3D SYS_TYPE_SET(req.flags, core_id); + + dmabuf =3D kmalloc(req.len, GFP_KERNEL); + if (!dmabuf) + return -ENOMEM; + + if (copy_from_user(dmabuf, u64_to_user_ptr(req.data), + req.len)) { + kfree_sensitive(dmabuf); + return -EFAULT; + } + + dma_addr =3D cmh_dma_map_single(dmabuf, req.len, DMA_TO_DEVICE); + if (cmh_dma_map_error(dma_addr)) { + kfree_sensitive(dmabuf); + return -ENOMEM; + } + + vcq_set_header(&vcq[0], MGMT_VCQ_CMDS); + vcq_add_sys_write(&vcq[1], req.ref, dma_addr, req.wrap_key, + req.len, sys_type); + /* + * PKE keys on Weierstrass curves and RSA keys must be byte-swapped + * when stored in the DS so they match the internal big-endian + * representation used by the PKE sidecar. Edwards curve keys + * (EdDSA) use native byte order and must NOT be swapped. + */ + switch (req.ds_type) { + case CMH_DS_RSA_PRIV_KEY: + case CMH_DS_RSA_PUB_KEY: + case CMH_DS_RSA_CRT_KEY: + case CMH_DS_ECDSA_PRIV_KEY: + case CMH_DS_ECDSA_PUB_KEY: + case CMH_DS_ECDH_PRIV_KEY: + case CMH_DS_SHARED_SECRET: + case CMH_DS_SM2_PRIV_KEY: + vcq[1].id |=3D PKE_SWAP_FLAGS; + break; + default: + /* EdDSA, symmetric keys -- no swap */ + break; + } + vcq_add_sys_flush(&vcq[2]); + + ret =3D cmh_tm_submit_sync_mbx(vcq, MGMT_VCQ_CMDS, 1, MGMT_MBX); + + cmh_dma_unmap_single(dma_addr, req.len, DMA_TO_DEVICE); + kfree_sensitive(dmabuf); + + if (ret) + return ret; + + dev_dbg(cmh_dev(), "mgmt: KEY_WRITE ref=3D0x%llx len=3D%u type=3D0x%x\n", + req.ref, req.len, sys_type); + return 0; +} + +/* -- KEY_READ -------------------------- */ + +static int cmh_mgmt_key_read(void __user *argp) +{ + struct cmh_ioctl_key_read req; + struct vcq_cmd vcq[MGMT_VCQ_CMDS]; + void *dmabuf; + dma_addr_t dma_addr; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.__reserved) + return -EINVAL; + if (!req.len || req.len > CMH_MGMT_MAX_DATA_LEN) + return -EINVAL; + + dmabuf =3D kzalloc(req.len, GFP_KERNEL); + if (!dmabuf) + return -ENOMEM; + + dma_addr =3D cmh_dma_map_single(dmabuf, req.len, DMA_FROM_DEVICE); + if (cmh_dma_map_error(dma_addr)) { + kfree(dmabuf); + return -ENOMEM; + } + + vcq_set_header(&vcq[0], MGMT_VCQ_CMDS); + vcq_add_sys_read(&vcq[1], req.ref, dma_addr, req.wrap_key, req.len); + vcq_add_sys_flush(&vcq[2]); + + ret =3D cmh_tm_submit_sync_mbx(vcq, MGMT_VCQ_CMDS, 1, MGMT_MBX); + + cmh_dma_unmap_single(dma_addr, req.len, DMA_FROM_DEVICE); + + if (ret) { + kfree_sensitive(dmabuf); + return ret; + } + + if (copy_to_user(u64_to_user_ptr(req.data), + dmabuf, req.len)) { + kfree_sensitive(dmabuf); + return -EFAULT; + } + + req.out_len =3D req.len; + kfree_sensitive(dmabuf); + + if (copy_to_user(argp, &req, sizeof(req))) + return -EFAULT; + + dev_dbg(cmh_dev(), "mgmt: KEY_READ ref=3D0x%llx len=3D%u\n", + req.ref, req.out_len); + return 0; +} + +/* -- KEY_FIND -------------------------- */ + +static int cmh_mgmt_key_find(void __user *argp) +{ + struct cmh_ioctl_key_find req; + struct vcq_cmd vcq[MGMT_VCQ_CMDS]; + struct sys_list_item *item; + dma_addr_t item_dma; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.__reserved) + return -EINVAL; + + item =3D kzalloc_obj(*item, GFP_KERNEL); + if (!item) + return -ENOMEM; + + item_dma =3D cmh_dma_map_single(item, sizeof(*item), DMA_FROM_DEVICE); + if (cmh_dma_map_error(item_dma)) { + kfree(item); + return -ENOMEM; + } + + vcq_set_header(&vcq[0], MGMT_VCQ_CMDS); + vcq_add_sys_find(&vcq[1], req.cid, item_dma, sizeof(*item)); + vcq_add_sys_flush(&vcq[2]); + + ret =3D cmh_tm_submit_sync_mbx(vcq, MGMT_VCQ_CMDS, 1, MGMT_MBX); + + cmh_dma_unmap_single(item_dma, sizeof(*item), DMA_FROM_DEVICE); + + if (ret) { + kfree(item); + return ret; + } + + req.ref =3D item->ref; + req.len =3D item->len; + req.type =3D item->type; + kfree(item); + + if (copy_to_user(argp, &req, sizeof(req))) + return -EFAULT; + + dev_dbg(cmh_dev(), "mgmt: KEY_FIND cid=3D0x%llx -> ref=3D0x%llx\n", + req.cid, req.ref); + return 0; +} + +/* -- KEY_LIST ------------------------- */ + +static int cmh_mgmt_key_list(void __user *argp) +{ + struct cmh_ioctl_key_list req; + struct vcq_cmd vcq[MGMT_VCQ_CMDS]; + struct sys_list_item *item; + dma_addr_t item_dma; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + + if (req.__reserved) + return -EINVAL; + + item =3D kzalloc_obj(*item, GFP_KERNEL); + if (!item) + return -ENOMEM; + + item_dma =3D cmh_dma_map_single(item, sizeof(*item), DMA_FROM_DEVICE); + if (cmh_dma_map_error(item_dma)) { + kfree(item); + return -ENOMEM; + } + + vcq_set_header(&vcq[0], MGMT_VCQ_CMDS); + vcq_add_sys_list(&vcq[1], req.start_ref, item_dma, sizeof(*item)); + vcq_add_sys_flush(&vcq[2]); + + ret =3D cmh_tm_submit_sync_mbx(vcq, MGMT_VCQ_CMDS, 1, MGMT_MBX); + + cmh_dma_unmap_single(item_dma, sizeof(*item), DMA_FROM_DEVICE); + + if (ret) { + kfree(item); + return ret; + } + + req.ref =3D item->ref; + req.cid =3D item->cid; + req.len =3D item->len; + req.type =3D item->type; + kfree(item); + + if (copy_to_user(argp, &req, sizeof(req))) + return -EFAULT; + + return 0; +} + +/* -- KEY_GRANT / KEY_DELETE --------------------- */ + +static int cmh_mgmt_key_grant(void __user *argp, bool is_delete) +{ + struct cmh_ioctl_key_grant req; + struct vcq_cmd vcq[MGMT_VCQ_CMDS]; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.__reserved) + return -EINVAL; + + /* DELETE =3D GRANT with all permissions zeroed */ + if (is_delete) { + req.read =3D 0; + req.write =3D 0; + req.execute =3D 0; + } + + vcq_set_header(&vcq[0], MGMT_VCQ_CMDS); + vcq_add_sys_grant(&vcq[1], req.ref, req.read, req.write, req.execute); + vcq_add_sys_flush(&vcq[2]); + + ret =3D cmh_tm_submit_sync_mbx(vcq, MGMT_VCQ_CMDS, 1, MGMT_MBX); + if (ret) + return ret; + + dev_dbg(cmh_dev(), "mgmt: KEY_%s ref=3D0x%llx r=3D0x%llx w=3D0x%llx x=3D0= x%llx\n", + is_delete ? "DELETE" : "GRANT", + req.ref, req.read, req.write, req.execute); + return 0; +} + +/* -- DS_EXPORT ------------------------- */ + +static int cmh_mgmt_ds_export(void __user *argp) +{ + struct cmh_ioctl_ds_export req; + struct vcq_cmd vcq[MGMT_VCQ_CMDS]; + void *dmabuf; + dma_addr_t dma_addr; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.__reserved) + return -EINVAL; + if (!req.len || req.len > CMH_MGMT_MAX_DATA_LEN) + return -EINVAL; + /* + * The eSW writes at least a sys_wrap_hdr into the buffer; reject a + * request too small to hold it up front. This also prevents the + * out-of-bounds read of hdr->wrap/hdr->len below on a short buffer. + */ + if (req.len < sizeof(struct sys_wrap_hdr)) + return -EINVAL; + + /* + * req.len is the exact DMA buffer size given to the eSW. + * Userspace must size it to at least the export blob: + * + * wrapped: sizeof(sys_wrap_hdr) + 2*AES_BLOCK_SIZE + obj_len + * =3D 16 + 32 + obj_len =3D 48 + obj_len + * plaintext: sizeof(sys_wrap_hdr) + obj_len + * =3D 16 + obj_len + * + * obj_len is known from KEY_NEW or KEY_FIND. If req.len is + * too small, the eSW rejects the command and we return -EIO. + */ + dmabuf =3D kzalloc(req.len, GFP_KERNEL); + if (!dmabuf) + return -ENOMEM; + + dma_addr =3D cmh_dma_map_single(dmabuf, req.len, DMA_FROM_DEVICE); + if (cmh_dma_map_error(dma_addr)) { + kfree(dmabuf); + return -ENOMEM; + } + + vcq_set_header(&vcq[0], MGMT_VCQ_CMDS); + vcq_add_sys_export(&vcq[1], req.cid, dma_addr, req.wrap_key, req.len); + vcq_add_sys_flush(&vcq[2]); + + ret =3D cmh_tm_submit_sync_mbx(vcq, MGMT_VCQ_CMDS, 1, MGMT_MBX); + + cmh_dma_unmap_single(dma_addr, req.len, DMA_FROM_DEVICE); + + if (ret) { + kfree_sensitive(dmabuf); + return ret; + } + + /* Parse actual blob size from the eSW-written header */ + { + struct sys_wrap_hdr *hdr =3D (struct sys_wrap_hdr *)dmabuf; + u64 actual; + + if (check_add_overflow((u64)sizeof(*hdr), (u64)hdr->wrap, + &actual) || + check_add_overflow(actual, (u64)hdr->len, &actual) || + actual > req.len) { + kfree_sensitive(dmabuf); + return -EIO; + } + req.out_len =3D (u32)actual; + } + + if (copy_to_user(u64_to_user_ptr(req.data), + dmabuf, req.out_len)) { + kfree_sensitive(dmabuf); + return -EFAULT; + } + + kfree_sensitive(dmabuf); + + if (copy_to_user(argp, &req, sizeof(req))) + return -EFAULT; + + dev_dbg(cmh_dev(), "mgmt: DS_EXPORT wrap_key=3D0x%llx len=3D%u\n", + req.wrap_key, req.out_len); + return 0; +} + +/* -- DS_IMPORT ------------------------- */ + +static int cmh_mgmt_ds_import(void __user *argp) +{ + struct cmh_ioctl_ds_import req; + struct vcq_cmd vcq[MGMT_VCQ_CMDS]; + void *dmabuf; + dma_addr_t dma_addr; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (!req.len || req.len > CMH_MGMT_MAX_DATA_LEN) + return -EINVAL; + + dmabuf =3D kmalloc(req.len, GFP_KERNEL); + if (!dmabuf) + return -ENOMEM; + + if (copy_from_user(dmabuf, u64_to_user_ptr(req.data), + req.len)) { + kfree_sensitive(dmabuf); + return -EFAULT; + } + + dma_addr =3D cmh_dma_map_single(dmabuf, req.len, DMA_TO_DEVICE); + if (cmh_dma_map_error(dma_addr)) { + kfree_sensitive(dmabuf); + return -ENOMEM; + } + + vcq_set_header(&vcq[0], MGMT_VCQ_CMDS); + vcq_add_sys_import(&vcq[1], dma_addr, req.wrap_key, req.len); + vcq_add_sys_flush(&vcq[2]); + + ret =3D cmh_tm_submit_sync_mbx(vcq, MGMT_VCQ_CMDS, 1, MGMT_MBX); + + cmh_dma_unmap_single(dma_addr, req.len, DMA_TO_DEVICE); + kfree_sensitive(dmabuf); + + if (ret) + return ret; + + dev_dbg(cmh_dev(), "mgmt: DS_IMPORT wrap_key=3D0x%llx len=3D%u\n", + req.wrap_key, req.len); + return 0; +} + +/* -- KIC key derivation ioctls -------- + * + * All four KIC derivation handlers (HKDF1, HKDF2, AES-CMAC-KDF, + * DKEK-derive) share the same two-mode structure and temp-flush pattern. + * + * Temp-storage flush rationale: + * + * The device maintains a small per-mailbox temporary key store + * (~960 bytes, LIFO). A derivation targeting SYS_REF_TEMP allocates + * from this store; the allocation persists across command-queue + * boundaries until either (a) a subsequent command consumes it or + * (b) a mailbox flush resets the store. + * + * Our single-derivation ioctls produce a temp key with no consumer + * in the same queue -- the key is consumed by a *later* ioctl + * (e.g. DS_EXPORT with wrap_key=3DSYS_REF_TEMP). If no consumer + * follows, the allocation persists. Sequential temp derivations + * accumulate allocations until the store is exhausted (3--8 calls + * depending on key size), after which the device returns ENOMEM. + * + * A mailbox flush (cmh_tm_flush_mbx / MBX_COMMAND_FLUSH) resets the + * temp store. It does NOT destroy persistent keys, datastore + * objects, or DRBG state -- only the command queue and temp store. + * + * Safe for cross-ioctl temp flows (e.g. export-to-file: + * HKDF1->TEMP in ioctl 1, then DS_EXPORT with wrap_key=3DTEMP in + * ioctl 2): the flush only happens in derivation handlers and in + * the pre-PKE dispatch path, not in DS_EXPORT/DS_IMPORT, so the + * temp key survives until consumed. + * + * The ioctl dispatch also flushes before PKE/SM2/PQC ioctls to + * protect them from temp residue left by earlier derivations on the + * same mailbox. The per-handler flushes here remain necessary + * because sequential temp derivations (without an intervening + * PKE/SM2/PQC ioctl) would still exhaust the store. + */ + +/* -- KIC_HKDF1 ------------------------- */ + +/* + * Derive a key from a KIC base key via one-step HKDF. + * + * Two modes controlled by CMH_KIC_FLAG_TEMP: + * + * TEMP (flag set) -- 3-command VCQ: + * [0] SYS header + * [1] KIC_CMD_HKDF1 (dst=3DSYS_REF_TEMP) + * [2] flush + * Returns SYS_REF_TEMP as ref. No DS entry created. + * + * Persistent (flag clear) -- 4-command VCQ: + * [0] SYS header + * [1] SYS_CMD_NEW (allocate DS slot, CMH eSW writes ref) + * [2] KIC_CMD_HKDF1 (dst=3DSYS_REF_LAST =3D just-allocated slot) + * [3] flush + * Returns the new DS reference. + */ +#define KDF_VCQ_MAX 4 +#define KDF_MAX_KEY_LEN 64 +#define KDF_MAX_LABEL_LEN 56 + +static int cmh_mgmt_kic_hkdf1(void __user *argp) +{ + struct cmh_ioctl_kic_hkdf1 req; + struct vcq_cmd vcq[KDF_VCQ_MAX]; + bool temp; + u64 *ref_buf =3D NULL; + void *label_buf =3D NULL; + dma_addr_t ref_dma =3D DMA_MAPPING_ERROR, label_dma =3D DMA_MAPPING_ERROR; + unsigned int n_cmds; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (!req.key_len || req.key_len > KDF_MAX_KEY_LEN) + return -EINVAL; + if (req.label_len > KDF_MAX_LABEL_LEN) + return -EINVAL; + + temp =3D !!(req.flags & CMH_KIC_FLAG_TEMP); + + /* + * Persistent path: need DMA buffer for CMH eSW to write the + * newly-allocated DS reference. + */ + if (!temp) { + ref_buf =3D kmalloc_obj(*ref_buf, GFP_KERNEL); + if (!ref_buf) + return -ENOMEM; + *ref_buf =3D 0; + ref_dma =3D cmh_dma_map_single(ref_buf, sizeof(*ref_buf), + DMA_FROM_DEVICE); + if (cmh_dma_map_error(ref_dma)) { + kfree(ref_buf); + return -ENOMEM; + } + } + + /* DMA buffer for label data (CMH eSW DMA-reads it) */ + if (req.label_len > 0) { + label_buf =3D kzalloc(req.label_len, GFP_KERNEL); + if (!label_buf) { + ret =3D -ENOMEM; + goto out_ref; + } + if (copy_from_user(label_buf, + u64_to_user_ptr(req.label), + req.label_len)) { + ret =3D -EFAULT; + goto out_label; + } + label_dma =3D cmh_dma_map_single(label_buf, req.label_len, + DMA_TO_DEVICE); + if (cmh_dma_map_error(label_dma)) { + ret =3D -ENOMEM; + goto out_label; + } + } + + /* Build VCQ */ + memset(vcq, 0, sizeof(vcq)); + + if (temp) { + /* Flush MBX to reset temp stack -- see KIC section comment */ + ret =3D cmh_tm_flush_mbx(MGMT_MBX); + if (ret) + goto out_unmap_label; + + n_cmds =3D 3; + vcq_set_header(&vcq[0], n_cmds); + vcq_add_kic_hkdf1(&vcq[1], SYS_REF_TEMP, req.base_key, + label_dma, req.key_len, req.label_len, + SYS_TYPE_SET(0, CORE_ID_AES)); + vcq_add_sys_flush(&vcq[2]); + } else { + n_cmds =3D 4; + vcq_set_header(&vcq[0], n_cmds); + vcq_add_sys_new(&vcq[1], req.cid, ref_dma, req.key_len); + vcq_add_kic_hkdf1(&vcq[2], SYS_REF_LAST, req.base_key, + label_dma, req.key_len, req.label_len, + SYS_TYPE_SET(0, CORE_ID_AES)); + vcq_add_sys_flush(&vcq[3]); + } + + ret =3D cmh_tm_submit_sync_mbx(vcq, n_cmds, 1, MGMT_MBX); + + /* Cleanup label DMA */ + if (label_buf) { + cmh_dma_unmap_single(label_dma, req.label_len, DMA_TO_DEVICE); + kfree(label_buf); + label_buf =3D NULL; + } + + if (ret) + goto out_ref; + + if (temp) { + req.ref =3D SYS_REF_TEMP; + atomic_set(&mgmt_temp_dirty, 1); + } else { + cmh_dma_unmap_single(ref_dma, sizeof(*ref_buf), + DMA_FROM_DEVICE); + req.ref =3D *ref_buf; + kfree(ref_buf); + ref_buf =3D NULL; + } + + if (copy_to_user(argp, &req, sizeof(req))) { + /* + * The derived-key DS slot was created but its ref never + * reached userspace, so the caller cannot free it. Logically + * delete the orphaned slot: grant-with-no-access wipes the key + * material and clears the CID (temp results use SYS_REF_TEMP + * and need no cleanup). + */ + if (!temp) { + cmh_mgmt_ds_scrub(req.ref, 0); + dev_warn(cmh_dev(), + "mgmt: KIC_HKDF1 copy_to_user failed, DS slot cleaned up\n"); + } + return -EFAULT; + } + + dev_dbg(cmh_dev(), + "mgmt: KIC_HKDF1 base=3D0x%llx len=3D%u flags=3D0x%x -> ref=3D0x%llx\n", + req.base_key, req.key_len, req.flags, req.ref); + return 0; + +out_unmap_label: + if (label_buf && !cmh_dma_map_error(label_dma)) + cmh_dma_unmap_single(label_dma, req.label_len, DMA_TO_DEVICE); +out_label: + kfree(label_buf); +out_ref: + if (ref_buf) { + cmh_dma_unmap_single(ref_dma, sizeof(*ref_buf), + DMA_FROM_DEVICE); + /* + * If sys_new created the slot before a later command in the + * VCQ failed, the eSW wrote its ref here (0 otherwise); scrub + * it so a partial failure does not leak a datastore slot. + */ + if (*ref_buf) + cmh_mgmt_ds_scrub(*ref_buf, 0); + kfree(ref_buf); + } + return ret; +} + +/* -- KIC_HKDF2 ------------------------- */ + +/* + * Two-step HKDF key derivation. Same as HKDF1 but adds a salt key + * reference: Step 1: HMAC(salt, base) -> PRK; Step 2: HMAC(PRK, label) ->= key. + */ + +static int cmh_mgmt_kic_hkdf2(void __user *argp) +{ + struct cmh_ioctl_kic_hkdf2 req; + struct vcq_cmd vcq[KDF_VCQ_MAX]; + bool temp; + u64 *ref_buf =3D NULL; + void *label_buf =3D NULL; + dma_addr_t ref_dma =3D DMA_MAPPING_ERROR, label_dma =3D DMA_MAPPING_ERROR; + unsigned int n_cmds; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (!req.key_len || req.key_len > KDF_MAX_KEY_LEN) + return -EINVAL; + if (req.label_len > KDF_MAX_LABEL_LEN) + return -EINVAL; + + temp =3D !!(req.flags & CMH_KIC_FLAG_TEMP); + + if (!temp) { + ref_buf =3D kmalloc_obj(*ref_buf, GFP_KERNEL); + if (!ref_buf) + return -ENOMEM; + *ref_buf =3D 0; + ref_dma =3D cmh_dma_map_single(ref_buf, sizeof(*ref_buf), + DMA_FROM_DEVICE); + if (cmh_dma_map_error(ref_dma)) { + kfree(ref_buf); + return -ENOMEM; + } + } + + if (req.label_len > 0) { + label_buf =3D kzalloc(req.label_len, GFP_KERNEL); + if (!label_buf) { + ret =3D -ENOMEM; + goto out_ref2; + } + if (copy_from_user(label_buf, + u64_to_user_ptr(req.label), + req.label_len)) { + ret =3D -EFAULT; + goto out_label2; + } + label_dma =3D cmh_dma_map_single(label_buf, req.label_len, + DMA_TO_DEVICE); + if (cmh_dma_map_error(label_dma)) { + ret =3D -ENOMEM; + goto out_label2; + } + } + + memset(vcq, 0, sizeof(vcq)); + + if (temp) { + /* Flush MBX to reset temp stack -- see KIC section comment */ + ret =3D cmh_tm_flush_mbx(MGMT_MBX); + if (ret) + goto out_unmap_label2; + + n_cmds =3D 3; + vcq_set_header(&vcq[0], n_cmds); + vcq_add_kic_hkdf2(&vcq[1], SYS_REF_TEMP, req.base_key, + req.salt_key, label_dma, + req.key_len, req.label_len, + SYS_TYPE_SET(0, CORE_ID_AES)); + vcq_add_sys_flush(&vcq[2]); + } else { + n_cmds =3D 4; + vcq_set_header(&vcq[0], n_cmds); + vcq_add_sys_new(&vcq[1], req.cid, ref_dma, req.key_len); + vcq_add_kic_hkdf2(&vcq[2], SYS_REF_LAST, req.base_key, + req.salt_key, label_dma, + req.key_len, req.label_len, + SYS_TYPE_SET(0, CORE_ID_AES)); + vcq_add_sys_flush(&vcq[3]); + } + + ret =3D cmh_tm_submit_sync_mbx(vcq, n_cmds, 1, MGMT_MBX); + + if (label_buf) { + cmh_dma_unmap_single(label_dma, req.label_len, DMA_TO_DEVICE); + kfree(label_buf); + label_buf =3D NULL; + } + + if (ret) + goto out_ref2; + + if (temp) { + req.ref =3D SYS_REF_TEMP; + atomic_set(&mgmt_temp_dirty, 1); + } else { + cmh_dma_unmap_single(ref_dma, sizeof(*ref_buf), + DMA_FROM_DEVICE); + req.ref =3D *ref_buf; + kfree(ref_buf); + ref_buf =3D NULL; + } + + if (copy_to_user(argp, &req, sizeof(req))) { + /* + * The derived-key DS slot was created but its ref never + * reached userspace, so the caller cannot free it. Logically + * delete the orphaned slot: grant-with-no-access wipes the key + * material and clears the CID (temp results use SYS_REF_TEMP + * and need no cleanup). + */ + if (!temp) { + cmh_mgmt_ds_scrub(req.ref, 0); + dev_warn(cmh_dev(), + "mgmt: KIC_HKDF2 copy_to_user failed, DS slot cleaned up\n"); + } + return -EFAULT; + } + + dev_dbg(cmh_dev(), + "mgmt: KIC_HKDF2 base=3D0x%llx salt=3D0x%llx len=3D%u flags=3D0x%x -> re= f=3D0x%llx\n", + req.base_key, req.salt_key, req.key_len, req.flags, req.ref); + return 0; + +out_unmap_label2: + if (label_buf && !cmh_dma_map_error(label_dma)) + cmh_dma_unmap_single(label_dma, req.label_len, DMA_TO_DEVICE); +out_label2: + kfree(label_buf); +out_ref2: + if (ref_buf) { + cmh_dma_unmap_single(ref_dma, sizeof(*ref_buf), + DMA_FROM_DEVICE); + /* + * If sys_new created the slot before a later command in the + * VCQ failed, the eSW wrote its ref here (0 otherwise); scrub + * it so a partial failure does not leak a datastore slot. + */ + if (*ref_buf) + cmh_mgmt_ds_scrub(*ref_buf, 0); + kfree(ref_buf); + } + return ret; +} + +/* -- KIC_AES_CMAC_KDF ------------------ */ + +/* + * Derive a key using AES-CMAC-based KDF (NIST SP800-108 style). + * Base key must be 32 bytes. Output is always non-PT (the hub driver + * rejects SYS_TYPE_FLAG_PT). + * + * VCQ layout matches HKDF: TEMP mode uses 3 commands, persistent uses 4. + */ +#define CMAC_KDF_KEY_LEN 32 + +static int cmh_mgmt_kic_aes_cmac_kdf(void __user *argp) +{ + struct cmh_ioctl_kic_aes_cmac_kdf req; + struct vcq_cmd vcq[KDF_VCQ_MAX]; + bool temp; + u64 *ref_buf =3D NULL; + void *label_buf =3D NULL; + dma_addr_t ref_dma =3D DMA_MAPPING_ERROR, label_dma =3D DMA_MAPPING_ERROR; + unsigned int n_cmds; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.key_len !=3D CMAC_KDF_KEY_LEN) + return -EINVAL; + if (req.label_len > KDF_MAX_LABEL_LEN) + return -EINVAL; + + temp =3D !!(req.flags & CMH_KIC_FLAG_TEMP); + + if (!temp) { + ref_buf =3D kmalloc_obj(*ref_buf, GFP_KERNEL); + if (!ref_buf) + return -ENOMEM; + *ref_buf =3D 0; + ref_dma =3D cmh_dma_map_single(ref_buf, sizeof(*ref_buf), + DMA_FROM_DEVICE); + if (cmh_dma_map_error(ref_dma)) { + kfree(ref_buf); + return -ENOMEM; + } + } + + if (req.label_len > 0) { + label_buf =3D kzalloc(req.label_len, GFP_KERNEL); + if (!label_buf) { + ret =3D -ENOMEM; + goto out_ref_cmac; + } + if (copy_from_user(label_buf, + u64_to_user_ptr(req.label), + req.label_len)) { + ret =3D -EFAULT; + goto out_label_cmac; + } + label_dma =3D cmh_dma_map_single(label_buf, req.label_len, + DMA_TO_DEVICE); + if (cmh_dma_map_error(label_dma)) { + ret =3D -ENOMEM; + goto out_label_cmac; + } + } + + memset(vcq, 0, sizeof(vcq)); + + if (temp) { + /* Flush MBX to reset temp stack -- see KIC section comment */ + ret =3D cmh_tm_flush_mbx(MGMT_MBX); + if (ret) + goto out_unmap_label_cmac; + + n_cmds =3D 3; + vcq_set_header(&vcq[0], n_cmds); + vcq_add_kic_aes_cmac_kdf(&vcq[1], SYS_REF_TEMP, + req.base_key, label_dma, + req.key_len, req.label_len, + SYS_TYPE_SET(0, CORE_ID_AES)); + vcq_add_sys_flush(&vcq[2]); + } else { + n_cmds =3D 4; + vcq_set_header(&vcq[0], n_cmds); + vcq_add_sys_new(&vcq[1], req.cid, ref_dma, req.key_len); + vcq_add_kic_aes_cmac_kdf(&vcq[2], SYS_REF_LAST, + req.base_key, label_dma, + req.key_len, req.label_len, + SYS_TYPE_SET(0, CORE_ID_AES)); + vcq_add_sys_flush(&vcq[3]); + } + + ret =3D cmh_tm_submit_sync_mbx(vcq, n_cmds, 1, MGMT_MBX); + + if (label_buf) { + cmh_dma_unmap_single(label_dma, req.label_len, DMA_TO_DEVICE); + kfree(label_buf); + label_buf =3D NULL; + } + + if (ret) + goto out_ref_cmac; + + if (temp) { + req.ref =3D SYS_REF_TEMP; + atomic_set(&mgmt_temp_dirty, 1); + } else { + cmh_dma_unmap_single(ref_dma, sizeof(*ref_buf), + DMA_FROM_DEVICE); + req.ref =3D *ref_buf; + kfree(ref_buf); + ref_buf =3D NULL; + } + + if (copy_to_user(argp, &req, sizeof(req))) { + /* + * The derived-key DS slot was created but its ref never + * reached userspace, so the caller cannot free it. Logically + * delete the orphaned slot: grant-with-no-access wipes the key + * material and clears the CID (temp results use SYS_REF_TEMP + * and need no cleanup). + */ + if (!temp) { + cmh_mgmt_ds_scrub(req.ref, 0); + dev_warn(cmh_dev(), + "mgmt: KIC_AES_CMAC_KDF copy_to_user failed, DS slot cleaned up\n"); + } + return -EFAULT; + } + + dev_dbg(cmh_dev(), + "mgmt: KIC_AES_CMAC_KDF base=3D0x%llx len=3D%u flags=3D0x%x -> ref=3D0x%= llx\n", + req.base_key, req.key_len, req.flags, req.ref); + return 0; + +out_unmap_label_cmac: + if (label_buf && !cmh_dma_map_error(label_dma) && label_dma) + cmh_dma_unmap_single(label_dma, req.label_len, DMA_TO_DEVICE); +out_label_cmac: + kfree(label_buf); +out_ref_cmac: + if (ref_buf) { + cmh_dma_unmap_single(ref_dma, sizeof(*ref_buf), + DMA_FROM_DEVICE); + /* Scrub a slot sys_new may have created before the failure. */ + if (*ref_buf) + cmh_mgmt_ds_scrub(*ref_buf, 0); + kfree(ref_buf); + } + return ret; +} + +/* -- KIC_DKEK_DERIVE ------------------- */ + +/* + * Derive a Key Encryption Key (KEK) from a KIC base key. + * Output is tagged CORE_ID_KIC (usable for further derivation only). + * host_id=3D0 means the caller's own host; non-zero requires management + * host privilege (eSW enforces this). + */ +#define DKEK_VCQ_MAX 4 + +static int cmh_mgmt_kic_dkek_derive(void __user *argp) +{ + struct cmh_ioctl_kic_dkek_derive req; + struct vcq_cmd vcq[DKEK_VCQ_MAX]; + bool temp; + u64 *ref_buf =3D NULL; + void *meta_buf =3D NULL; + dma_addr_t ref_dma =3D DMA_MAPPING_ERROR, meta_dma =3D DMA_MAPPING_ERROR; + unsigned int n_cmds; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.metadata_len > KIC_DKEK_MAX_METADATA) + return -EINVAL; + + temp =3D !!(req.flags & CMH_KIC_FLAG_TEMP); + + if (!temp) { + ref_buf =3D kmalloc_obj(*ref_buf, GFP_KERNEL); + if (!ref_buf) + return -ENOMEM; + *ref_buf =3D 0; + ref_dma =3D cmh_dma_map_single(ref_buf, sizeof(*ref_buf), + DMA_FROM_DEVICE); + if (cmh_dma_map_error(ref_dma)) { + kfree(ref_buf); + return -ENOMEM; + } + } + + if (req.metadata_len > 0) { + meta_buf =3D kzalloc(req.metadata_len, GFP_KERNEL); + if (!meta_buf) { + ret =3D -ENOMEM; + goto out_ref_dkek; + } + if (copy_from_user(meta_buf, + u64_to_user_ptr(req.metadata), + req.metadata_len)) { + ret =3D -EFAULT; + goto out_meta; + } + meta_dma =3D cmh_dma_map_single(meta_buf, req.metadata_len, + DMA_TO_DEVICE); + if (cmh_dma_map_error(meta_dma)) { + ret =3D -ENOMEM; + goto out_meta; + } + } + + memset(vcq, 0, sizeof(vcq)); + + if (temp) { + /* Flush MBX to reset temp stack -- see KIC section comment */ + ret =3D cmh_tm_flush_mbx(MGMT_MBX); + if (ret) + goto out_unmap_meta; + + n_cmds =3D 3; + vcq_set_header(&vcq[0], n_cmds); + vcq_add_kic_dkek_derive(&vcq[1], SYS_REF_TEMP, + req.base_key, req.host_id, + meta_dma, req.metadata_len); + vcq_add_sys_flush(&vcq[2]); + } else { + n_cmds =3D 4; + vcq_set_header(&vcq[0], n_cmds); + vcq_add_sys_new(&vcq[1], req.cid, ref_dma, KIC_KEY_SIZE); + vcq_add_kic_dkek_derive(&vcq[2], SYS_REF_LAST, + req.base_key, req.host_id, + meta_dma, req.metadata_len); + vcq_add_sys_flush(&vcq[3]); + } + + ret =3D cmh_tm_submit_sync_mbx(vcq, n_cmds, 1, MGMT_MBX); + + if (meta_buf) { + cmh_dma_unmap_single(meta_dma, req.metadata_len, + DMA_TO_DEVICE); + kfree(meta_buf); + meta_buf =3D NULL; + } + + if (ret) + goto out_ref_dkek; + + if (temp) { + req.ref =3D SYS_REF_TEMP; + atomic_set(&mgmt_temp_dirty, 1); + } else { + cmh_dma_unmap_single(ref_dma, sizeof(*ref_buf), + DMA_FROM_DEVICE); + req.ref =3D *ref_buf; + kfree(ref_buf); + ref_buf =3D NULL; + } + + if (copy_to_user(argp, &req, sizeof(req))) { + /* + * The derived-key DS slot was created but its ref never + * reached userspace, so the caller cannot free it. Logically + * delete the orphaned slot: grant-with-no-access wipes the key + * material and clears the CID (temp results use SYS_REF_TEMP + * and need no cleanup). + */ + if (!temp) { + cmh_mgmt_ds_scrub(req.ref, 0); + dev_warn(cmh_dev(), + "mgmt: KIC_DKEK_DERIVE copy_to_user failed, DS slot cleaned up\n"); + } + return -EFAULT; + } + + dev_dbg(cmh_dev(), + "mgmt: KIC_DKEK_DERIVE base=3D0x%llx host=3D%u meta_len=3D%u flags=3D0x%= x -> ref=3D0x%llx\n", + req.base_key, req.host_id, req.metadata_len, req.flags, + req.ref); + return 0; + +out_unmap_meta: + if (meta_buf && !cmh_dma_map_error(meta_dma) && meta_dma) + cmh_dma_unmap_single(meta_dma, req.metadata_len, DMA_TO_DEVICE); +out_meta: + kfree(meta_buf); +out_ref_dkek: + if (ref_buf) { + cmh_dma_unmap_single(ref_dma, sizeof(*ref_buf), + DMA_FROM_DEVICE); + /* Scrub a slot sys_new may have created before the failure. */ + if (*ref_buf) + cmh_mgmt_ds_scrub(*ref_buf, 0); + kfree(ref_buf); + } + return ret; +} + +/* -- KEY_NEW_RANDOM -- DRBG-backed key generation --- */ + +/* + * Allocate a new datastore slot and fill it with DRBG-generated + * random key material in a single atomic VCQ submission: + * + * [0] SYS header(5) + * [1] SYS_CMD_NEW -- allocate DS slot (CMH eSW writes ref) + * [2] DRBG_CMD_DATASTORE(SYS_REF_LAST) -- fill with random data + * [3] DRBG flush -- release DRBG core ownership + * [4] SYS flush + * + * The DRBG must be configured before this ioctl is used. + * Reuses struct cmh_ioctl_key_new (ds_type, flags, cid, len, ref). + */ +#define DRBG_KEYGEN_VCQ_CMDS 5 + +static int cmh_mgmt_key_new_random(void __user *argp) +{ + struct cmh_ioctl_key_new req; + struct vcq_cmd vcq[DRBG_KEYGEN_VCQ_CMDS]; + u64 *ref_buf; + dma_addr_t ref_dma; + u32 core_id, sys_type; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (!req.len || req.len > CMH_MGMT_MAX_DATA_LEN) + return -EINVAL; + + core_id =3D cmh_ds_type_to_core_id(req.ds_type); + if (core_id =3D=3D CORE_ID_NUM) + return -EINVAL; + sys_type =3D SYS_TYPE_SET(req.flags, core_id); + + ref_buf =3D kmalloc_obj(*ref_buf, GFP_KERNEL); + if (!ref_buf) + return -ENOMEM; + + *ref_buf =3D 0; + ref_dma =3D cmh_dma_map_single(ref_buf, sizeof(*ref_buf), + DMA_FROM_DEVICE); + if (cmh_dma_map_error(ref_dma)) { + kfree(ref_buf); + return -ENOMEM; + } + + vcq_set_header(&vcq[0], DRBG_KEYGEN_VCQ_CMDS); + vcq_add_sys_new(&vcq[1], req.cid, ref_dma, req.len); + vcq_add_drbg_datastore(&vcq[2], SYS_REF_LAST, req.len, sys_type); + vcq_add_flush(&vcq[3], CORE_ID_DRBG); + vcq_add_sys_flush(&vcq[4]); + + ret =3D cmh_tm_submit_sync_mbx(vcq, DRBG_KEYGEN_VCQ_CMDS, 1, MGMT_MBX); + + cmh_dma_unmap_single(ref_dma, sizeof(*ref_buf), DMA_FROM_DEVICE); + + if (ret) { + /* Scrub a slot sys_new may have created before the failure. */ + if (*ref_buf) + cmh_mgmt_ds_scrub(*ref_buf, 0); + kfree(ref_buf); + return ret; + } + + req.ref =3D *ref_buf; + kfree(ref_buf); + + if (copy_to_user(argp, &req, sizeof(req))) { + /* + * The DS slot was created but its ref never reached + * userspace, so the caller cannot free it. Logically delete + * the orphaned slot: grant-with-no-access wipes the key + * material and clears the CID (it does not reclaim the stack + * space -- only a datastore reset does). + */ + cmh_mgmt_ds_scrub(req.ref, 0); + dev_warn(cmh_dev(), + "mgmt: KEY_NEW_RANDOM copy_to_user failed, DS slot cleaned up\n"); + return -EFAULT; + } + + dev_dbg(cmh_dev(), + "mgmt: KEY_NEW_RANDOM cid=3D0x%llx len=3D%u type=3D0x%x -> ref=3D0x%llx\= n", + req.cid, req.len, sys_type, req.ref); + return 0; +} + +#define EAC_VCQ_CMDS 3 /* header + EAC_READ + flush */ + +static long cmh_mgmt_eac_read(void __user *argp) +{ + struct cmh_ioctl_eac_read req; + struct eac_read_rsp *rsp; + struct vcq_cmd vcq[EAC_VCQ_CMDS]; + dma_addr_t rsp_dma; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.__reserved !=3D 0) + return -EINVAL; + if (req.__pad !=3D 0) + return -EINVAL; + + /* + * Zero the response: the eSW writes the fields it populates, but a + * partial/aborted read must not copy uninitialised heap back to + * user space. + */ + rsp =3D kzalloc_obj(*rsp, GFP_KERNEL); + if (!rsp) + return -ENOMEM; + + rsp_dma =3D cmh_dma_map_single(rsp, sizeof(*rsp), DMA_FROM_DEVICE); + if (cmh_dma_map_error(rsp_dma)) { + kfree(rsp); + return -ENOMEM; + } + + vcq_set_header(&vcq[0], EAC_VCQ_CMDS); + vcq_add_eac_read(&vcq[1], rsp_dma, sizeof(*rsp)); + vcq_add_flush(&vcq[2], CORE_ID_EAC); + + ret =3D cmh_tm_submit_sync_mbx(vcq, EAC_VCQ_CMDS, 1, MGMT_MBX); + + cmh_dma_unmap_single(rsp_dma, sizeof(*rsp), DMA_FROM_DEVICE); + + if (ret) { + kfree(rsp); + return ret; + } + + /* Copy response fields into ioctl struct */ + req.mailbox_notification =3D rsp->mailbox_notification; + req.hw_error =3D rsp->hw_error; + req.hw_nmi =3D rsp->hw_nmi; + req.hw_panic =3D rsp->hw_panic; + req.safety_fatal =3D rsp->safety_fatal; + req.safety_notification =3D rsp->safety_notification; + req.sw_info0 =3D rsp->sw_info0; + req.sw_info1 =3D rsp->sw_info1; + memcpy(req.sram_bank_errors, rsp->sram_bank_errors, + sizeof(req.sram_bank_errors)); + req.__pad =3D 0; + + kfree(rsp); + + if (copy_to_user(argp, &req, sizeof(req))) + return -EFAULT; + + return 0; +} + +/* -- DRBG CONFIG (management) ------------ */ + +#define DRBG_CONFIG_VCQ_CMDS 4 /* header + RESET + CONFIG + flush */ + +static long cmh_mgmt_drbg_config(void __user *argp) +{ + struct cmh_ioctl_drbg_config req; + struct vcq_cmd vcq[DRBG_CONFIG_VCQ_CMDS]; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.__reserved !=3D 0) + return -EINVAL; + if (req.entropy_ratio > 3) + return -EINVAL; + if (req.security_strength !=3D CMH_DRBG_STRENGTH_128 && + req.security_strength !=3D CMH_DRBG_STRENGTH_256) + return -EINVAL; + + vcq_set_header(&vcq[0], DRBG_CONFIG_VCQ_CMDS); + vcq_add_drbg_reset(&vcq[1]); + vcq_add_drbg_config(&vcq[2], req.entropy_ratio, + req.security_strength); + vcq_add_flush(&vcq[3], CORE_ID_DRBG); + + ret =3D cmh_tm_submit_sync_mbx(vcq, DRBG_CONFIG_VCQ_CMDS, 1, MGMT_MBX); + if (ret) + dev_warn(cmh_dev(), "mgmt: DRBG CONFIG failed (rc=3D%d)\n", ret); + + return ret; +} + +/* -- ioctl dispatch ------------------------ */ + +/* + * PKE, SM2, and PQC ioctls use device-internal temporary storage for + * intermediate results. Residual allocations in the per-mailbox temp + * store (left by prior operations that targeted SYS_REF_TEMP) reduce + * the space available and can cause the device to return ENOMEM. + * + * Flush the mailbox before these operations to reset the temp store, + * but ONLY when the store is actually dirty (mgmt_temp_dirty flag). + * Unconditional flushing would kill in-flight command queues from + * concurrent callers on the same mailbox -- MBX_COMMAND_FLUSH + * terminates any executing queue with -EPIPE and discards all queued + * submissions. + * + * The conditional flush is safe: PKE/SM2/PQC ioctls do not consume + * SYS_REF_TEMP from a prior ioctl (unlike DS_EXPORT/DS_IMPORT which + * may reference a temp key produced by a preceding derivation), so + * clearing the temp store before them loses no needed state. + */ +static inline bool cmh_mgmt_needs_temp_flush(unsigned int cmd) +{ + unsigned int nr =3D _IOC_NR(cmd); + + /* + * Range invariant: all PKE/SM2/PQC ioctls must have consecutive + * NR values between PKE_RSA_ENC (0x10) and SM2_ENC_HASH (0x37). + * If a new ioctl is added outside this range, update the bounds + * and adjust these assertions. + */ + BUILD_BUG_ON(_IOC_NR(CMH_IOCTL_PKE_RSA_ENC) !=3D 0x10); + BUILD_BUG_ON(_IOC_NR(CMH_IOCTL_SM2_ENC_HASH) !=3D 0x37); + + return nr >=3D _IOC_NR(CMH_IOCTL_PKE_RSA_ENC) && + nr <=3D _IOC_NR(CMH_IOCTL_SM2_ENC_HASH); +} + +static long cmh_mgmt_ioctl_locked(struct file *file, unsigned int cmd, + void __user *argp) +{ + int ret; + + if (cmh_mgmt_needs_temp_flush(cmd) && + atomic_xchg(&mgmt_temp_dirty, 0)) { + ret =3D cmh_tm_flush_mbx(MGMT_MBX); + if (ret) { + /* Flush failed -- the temp region is still dirty. */ + atomic_set(&mgmt_temp_dirty, 1); + return ret; + } + } + + switch (cmd) { + case CMH_IOCTL_KEY_NEW: + return cmh_mgmt_key_new(argp); + case CMH_IOCTL_KEY_WRITE: + return cmh_mgmt_key_write(argp); + case CMH_IOCTL_KEY_READ: + return cmh_mgmt_key_read(argp); + case CMH_IOCTL_KEY_FIND: + return cmh_mgmt_key_find(argp); + case CMH_IOCTL_KEY_GRANT: + return cmh_mgmt_key_grant(argp, false); + case CMH_IOCTL_KEY_DELETE: + return cmh_mgmt_key_grant(argp, true); + case CMH_IOCTL_DS_EXPORT: + return cmh_mgmt_ds_export(argp); + case CMH_IOCTL_DS_IMPORT: + return cmh_mgmt_ds_import(argp); + case CMH_IOCTL_KIC_HKDF1: + return cmh_mgmt_kic_hkdf1(argp); + case CMH_IOCTL_KIC_HKDF2: + return cmh_mgmt_kic_hkdf2(argp); + case CMH_IOCTL_KEY_NEW_RANDOM: + return cmh_mgmt_key_new_random(argp); + case CMH_IOCTL_KIC_AES_CMAC_KDF: + return cmh_mgmt_kic_aes_cmac_kdf(argp); + case CMH_IOCTL_KIC_DKEK_DERIVE: + return cmh_mgmt_kic_dkek_derive(argp); + case CMH_IOCTL_KEY_LIST: + return cmh_mgmt_key_list(argp); + case CMH_IOCTL_EAC_READ: + return cmh_mgmt_eac_read(argp); + /* PKE operations */ + case CMH_IOCTL_PKE_RSA_ENC: + return cmh_mgmt_pke_rsa_enc(argp); + case CMH_IOCTL_PKE_RSA_DEC: + return cmh_mgmt_pke_rsa_dec(argp); + case CMH_IOCTL_PKE_RSA_CRT_DEC: + return cmh_mgmt_pke_rsa_crt_dec(argp); + case CMH_IOCTL_PKE_RSA_KEYGEN: + return cmh_mgmt_pke_rsa_keygen(argp); + case CMH_IOCTL_PKE_ECDSA_SIGN: + return cmh_mgmt_pke_ecdsa_sign(argp); + case CMH_IOCTL_PKE_ECDH: + return cmh_mgmt_pke_ecdh(argp); + case CMH_IOCTL_PKE_ECDH_KEYGEN: + return cmh_mgmt_pke_ecdh_keygen(argp); + case CMH_IOCTL_PKE_EDDSA_SIGN: + return cmh_mgmt_pke_eddsa_sign(argp); + case CMH_IOCTL_PKE_EDDSA_VERIFY: + return cmh_mgmt_pke_eddsa_verify(argp); + case CMH_IOCTL_PKE_EC_KEYGEN: + return cmh_mgmt_pke_ec_keygen(argp); + case CMH_IOCTL_PKE_EC_PUBGEN: + return cmh_mgmt_pke_ec_pubgen(argp); + case CMH_IOCTL_PKE_EDDSA_KEYGEN_SCA: + return cmh_mgmt_pke_eddsa_keygen_sca(argp); + /* SM2 operations */ + case CMH_IOCTL_SM2_ECDH_KEYGEN: + return cmh_mgmt_sm2_ecdh_keygen(argp); + case CMH_IOCTL_SM2_ECDH: + return cmh_mgmt_sm2_ecdh(argp); + case CMH_IOCTL_SM2_DEC_POINT: + return cmh_mgmt_sm2_dec_point(argp); + case CMH_IOCTL_SM2_ENC_POINT: + return cmh_mgmt_sm2_enc_point(argp); + case CMH_IOCTL_SM2_ID_DIGEST: + return cmh_mgmt_sm2_id_digest(argp); + case CMH_IOCTL_SM2_ECDH_HASH: + return cmh_mgmt_sm2_ecdh_hash(argp); + case CMH_IOCTL_SM2_DEC_HASH: + return cmh_mgmt_sm2_dec_hash(argp); + case CMH_IOCTL_SM2_ENC_HASH: + return cmh_mgmt_sm2_enc_hash(argp); + /* PQC operations */ + case CMH_IOCTL_ML_KEM_KEYGEN: + return cmh_mgmt_ml_kem_keygen(argp); + case CMH_IOCTL_ML_KEM_ENC: + return cmh_mgmt_ml_kem_enc(argp); + case CMH_IOCTL_ML_KEM_DEC: + return cmh_mgmt_ml_kem_dec(argp); + case CMH_IOCTL_ML_DSA_KEYGEN: + return cmh_mgmt_ml_dsa_keygen(argp); + case CMH_IOCTL_ML_DSA_SIGN: + return cmh_mgmt_ml_dsa_sign(argp); + case CMH_IOCTL_SLHDSA_KEYGEN: + return cmh_mgmt_slhdsa_keygen(argp); + case CMH_IOCTL_SLHDSA_SIGN: + return cmh_mgmt_slhdsa_sign(argp); + case CMH_IOCTL_SLHDSA_SIGN_PREHASH: + return cmh_mgmt_slhdsa_sign_prehash(argp); + /* DRBG management */ + case CMH_IOCTL_DRBG_CONFIG: + return cmh_mgmt_drbg_config(argp); + default: + return -ENOTTY; + } +} + +/* + * cmh_mgmt_ioctl() - serialised entry point for /dev/cmh_mgmt + * + * Holds cmh_mgmt_lock across the whole operation so the shared + * management-plane state (SYS_REF_TEMP, datastore) cannot be clobbered + * by a concurrent ioctl on another open of the device. + */ +static long cmh_mgmt_ioctl(struct file *file, unsigned int cmd, + unsigned long arg) +{ + void __user *argp =3D (void __user *)arg; + long ret; + + mutex_lock(&cmh_mgmt_lock); + ret =3D cmh_mgmt_ioctl_locked(file, cmd, argp); + mutex_unlock(&cmh_mgmt_lock); + return ret; +} + +/* -- File operations ----------------------- */ + +/* + * Capability is checked once at open time. A privileged process may + * pass the resulting fd to an unprivileged helper -- this delegation + * model is intentional and mirrors /dev/kvm, /dev/loop-control, etc. + */ +static int cmh_mgmt_open(struct inode *inode, struct file *file) +{ + if (!capable(CAP_SYS_ADMIN)) + return -EPERM; + + return 0; +} + +static const struct file_operations cmh_mgmt_fops =3D { + .owner =3D THIS_MODULE, + .open =3D cmh_mgmt_open, + .unlocked_ioctl =3D cmh_mgmt_ioctl, + .compat_ioctl =3D compat_ptr_ioctl, +}; + +static struct miscdevice cmh_mgmt_dev =3D { + .minor =3D MISC_DYNAMIC_MINOR, + .name =3D "cmh_mgmt", + .fops =3D &cmh_mgmt_fops, + .mode =3D 0660, +}; + +static bool cmh_mgmt_registered; + +/** + * cmh_mgmt_register() - Register the /dev/cmh_mgmt misc device + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_register(void) +{ + int ret; + + /* + * ABI size guards -- catch silent layout changes at compile time. + * All ioctl structs use only __u32 and __u64 with explicit padding, + * guaranteeing identical layout on 32-bit and 64-bit (compat_ptr_ioctl). + */ + BUILD_BUG_ON(sizeof(struct cmh_ioctl_key_new) !=3D 32); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_key_write) !=3D 40); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_key_read) !=3D 40); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_key_find) !=3D 32); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_key_list) !=3D 40); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_key_grant) !=3D 40); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_ds_export) !=3D 40); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_ds_import) !=3D 24); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_kic_hkdf1) !=3D 48); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_kic_hkdf2) !=3D 56); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_kic_aes_cmac_kdf) !=3D 48); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_kic_dkek_derive) !=3D 48); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_pke_rsa_enc) !=3D 48); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_pke_rsa_dec) !=3D 56); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_pke_rsa_crt_dec) !=3D 56); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_pke_rsa_keygen) !=3D 64); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_pke_ecdsa_sign) !=3D 40); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_pke_ecdh) !=3D 48); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_pke_ecdh_keygen) !=3D 24); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_pke_eddsa_sign) !=3D 40); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_pke_eddsa_verify) !=3D 40); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_pke_ec_keygen) !=3D 32); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_pke_ec_pubgen) !=3D 24); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_pke_eddsa_keygen_sca) !=3D 32); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_ml_kem_keygen) !=3D 64); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_ml_kem_enc) !=3D 64); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_ml_kem_dec) !=3D 56); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_ml_dsa_keygen) !=3D 56); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_ml_dsa_sign) !=3D 48); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_slhdsa_keygen) !=3D 56); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_slhdsa_sign) !=3D 56); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_slhdsa_sign_prehash) !=3D 64); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_sm2_ecdh_keygen) !=3D 24); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_sm2_ecdh) !=3D 56); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_sm2_dec_point) !=3D 32); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_sm2_enc_point) !=3D 40); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_sm2_id_digest) !=3D 32); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_sm2_ecdh_hash) !=3D 40); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_sm2_dec_hash) !=3D 32); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_sm2_enc_hash) !=3D 32); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_eac_read) !=3D 64); + BUILD_BUG_ON(sizeof(struct cmh_ioctl_drbg_config) !=3D 16); + + ret =3D misc_register(&cmh_mgmt_dev); + if (ret) { + dev_err(cmh_dev(), "mgmt: misc_register failed (rc=3D%d)\n", ret); + return ret; + } + + cmh_mgmt_registered =3D true; + return 0; +} + +/** + * cmh_mgmt_unregister() - Unregister the /dev/cmh_mgmt misc device + */ +void cmh_mgmt_unregister(void) +{ + if (!cmh_mgmt_registered) + return; + + misc_deregister(&cmh_mgmt_dev); + cmh_mgmt_registered =3D false; +} diff --git a/drivers/crypto/cmh/cmh_mgmt_pke.c b/drivers/crypto/cmh/cmh_mgm= t_pke.c new file mode 100644 index 000000000000..189af114fa43 --- /dev/null +++ b/drivers/crypto/cmh/cmh_mgmt_pke.c @@ -0,0 +1,1143 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH -- PKE ioctl handlers for /dev/cmh_mgmt + * + * RSA encrypt/decrypt/CRT/keygen, ECDSA sign, ECDH/keygen, + * EdDSA sign/verify, EC keygen/pubgen. + * + * Split from cmh_mgmt.c for maintainability. + */ + +#include +#include +#include +#include + +#include "cmh_mgmt.h" +#include "cmh_sys.h" +#include "cmh_txn.h" +#include "cmh_key.h" +#include "cmh_dma.h" +#include "cmh_config.h" +#include "cmh_pke.h" +#include "cmh_pke_abi.h" +#include "cmh_sys_abi.h" +#include + +#include + +/* -- PKE ioctl helpers ------------------- */ + +/* + * Maximum PKE operand size: 512 bytes (RSA 4096-bit), + * or 2 * 68 =3D 136 bytes (P-521 coordinate pair). + */ +#define PKE_MAX_OPERAND 512 + +/* Validate curve ID and return coordinate length; 0 =3D invalid */ +static u32 cmh_pke_validate_curve(u32 curve) +{ + return pke_curve_clen(curve); +} + +/** + * cmh_mgmt_pke_rsa_enc() - Handle CMH_MGMT_IOC_PKE_RSA_ENC ioctl + * @argp: User-space ioctl argument pointer + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_pke_rsa_enc(void __user *argp) +{ + u32 pke_cid =3D cmh_core_default_id(CMH_CORE_PKE); + + struct cmh_ioctl_pke_rsa_enc req; + struct vcq_cmd vcq[PKE_VCQ_CMDS_MIN]; + u32 n_len, e_padded; + u8 *e_buf, *n_buf, *m_buf, *c_buf; + dma_addr_t e_dma, n_dma, m_dma, c_dma; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.__reserved) + return -EINVAL; + if (req.bits < PKE_RSA_MIN_BITS || req.bits > PKE_RSA_MAX_BITS || + req.bits % 8) + return -EINVAL; + if (!req.e_len || req.e_len > PKE_MAX_OPERAND) + return -EINVAL; + + n_len =3D req.bits / 8; + e_padded =3D ALIGN(req.e_len, 4); + + e_buf =3D kzalloc(e_padded, GFP_KERNEL); + n_buf =3D kmalloc(n_len, GFP_KERNEL); + m_buf =3D kmalloc(n_len, GFP_KERNEL); + c_buf =3D kzalloc(n_len, GFP_KERNEL); + if (!e_buf || !n_buf || !m_buf || !c_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + /* Right-align exponent in zero-padded buffer for DMA alignment */ + if (copy_from_user(e_buf + e_padded - req.e_len, + u64_to_user_ptr(req.e), req.e_len) || + copy_from_user(n_buf, u64_to_user_ptr(req.n), n_len) || + copy_from_user(m_buf, u64_to_user_ptr(req.input), n_len)) { + ret =3D -EFAULT; + goto out_free; + } + + e_dma =3D cmh_dma_map_single(e_buf, e_padded, DMA_TO_DEVICE); + n_dma =3D cmh_dma_map_single(n_buf, n_len, DMA_TO_DEVICE); + m_dma =3D cmh_dma_map_single(m_buf, n_len, DMA_TO_DEVICE); + c_dma =3D cmh_dma_map_single(c_buf, n_len, DMA_FROM_DEVICE); + if (cmh_dma_map_error(e_dma) || cmh_dma_map_error(n_dma) || + cmh_dma_map_error(m_dma) || cmh_dma_map_error(c_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + vcq_set_header(&vcq[0], PKE_VCQ_CMDS_MIN); + vcq_add_pke_rsa_enc(&vcq[1], pke_cid, req.bits, e_padded, + e_dma, n_dma, m_dma, c_dma, PKE_SWAP_FLAGS); + vcq_add_pke_flush(&vcq[2], pke_cid); + + ret =3D cmh_tm_submit_sync_mbx(vcq, PKE_VCQ_CMDS_MIN, 1, MGMT_MBX); + +out_unmap: + if (!cmh_dma_map_error(c_dma)) + cmh_dma_unmap_single(c_dma, n_len, DMA_FROM_DEVICE); + if (!cmh_dma_map_error(m_dma)) + cmh_dma_unmap_single(m_dma, n_len, DMA_TO_DEVICE); + if (!cmh_dma_map_error(n_dma)) + cmh_dma_unmap_single(n_dma, n_len, DMA_TO_DEVICE); + if (!cmh_dma_map_error(e_dma)) + cmh_dma_unmap_single(e_dma, e_padded, DMA_TO_DEVICE); + + if (!ret) { + if (copy_to_user(u64_to_user_ptr(req.output), c_buf, n_len)) + ret =3D -EFAULT; + } + +out_free: + kfree(c_buf); + kfree_sensitive(m_buf); + kfree(n_buf); + kfree(e_buf); + return ret; +} + +/** + * cmh_mgmt_pke_rsa_dec() - Handle CMH_MGMT_IOC_PKE_RSA_DEC ioctl + * @argp: User-space ioctl argument pointer + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_pke_rsa_dec(void __user *argp) +{ + u32 pke_cid =3D cmh_core_default_id(CMH_CORE_PKE); + + struct cmh_ioctl_pke_rsa_dec req; + struct vcq_cmd vcq[PKE_VCQ_CMDS_MIN]; + u32 n_len, e_padded; + u8 *e_buf, *n_buf, *c_buf, *m_buf; + dma_addr_t e_dma, n_dma, c_dma, m_dma; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.__reserved) + return -EINVAL; + if (req.bits < PKE_RSA_MIN_BITS || req.bits > PKE_RSA_MAX_BITS || + req.bits % 8) + return -EINVAL; + if (!req.e_len || req.e_len > PKE_MAX_OPERAND) + return -EINVAL; + + n_len =3D req.bits / 8; + e_padded =3D ALIGN(req.e_len, 4); + + e_buf =3D kzalloc(e_padded, GFP_KERNEL); + n_buf =3D kmalloc(n_len, GFP_KERNEL); + c_buf =3D kmalloc(n_len, GFP_KERNEL); + m_buf =3D kzalloc(n_len, GFP_KERNEL); + if (!e_buf || !n_buf || !c_buf || !m_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + /* Right-align exponent in zero-padded buffer for DMA alignment */ + if (copy_from_user(e_buf + e_padded - req.e_len, + u64_to_user_ptr(req.e), req.e_len) || + copy_from_user(n_buf, u64_to_user_ptr(req.n), n_len) || + copy_from_user(c_buf, u64_to_user_ptr(req.input), n_len)) { + ret =3D -EFAULT; + goto out_free; + } + + e_dma =3D cmh_dma_map_single(e_buf, e_padded, DMA_TO_DEVICE); + n_dma =3D cmh_dma_map_single(n_buf, n_len, DMA_TO_DEVICE); + c_dma =3D cmh_dma_map_single(c_buf, n_len, DMA_TO_DEVICE); + m_dma =3D cmh_dma_map_single(m_buf, n_len, DMA_FROM_DEVICE); + if (cmh_dma_map_error(e_dma) || cmh_dma_map_error(n_dma) || + cmh_dma_map_error(c_dma) || cmh_dma_map_error(m_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + vcq_set_header(&vcq[0], PKE_VCQ_CMDS_MIN); + vcq_add_pke_rsa_dec(&vcq[1], pke_cid, req.bits, e_padded, + e_dma, n_dma, c_dma, m_dma, req.key_ref, + PKE_SWAP_FLAGS); + vcq_add_pke_flush(&vcq[2], pke_cid); + + ret =3D cmh_tm_submit_sync_mbx(vcq, PKE_VCQ_CMDS_MIN, 1, MGMT_MBX); + +out_unmap: + if (!cmh_dma_map_error(m_dma)) + cmh_dma_unmap_single(m_dma, n_len, DMA_FROM_DEVICE); + if (!cmh_dma_map_error(c_dma)) + cmh_dma_unmap_single(c_dma, n_len, DMA_TO_DEVICE); + if (!cmh_dma_map_error(n_dma)) + cmh_dma_unmap_single(n_dma, n_len, DMA_TO_DEVICE); + if (!cmh_dma_map_error(e_dma)) + cmh_dma_unmap_single(e_dma, e_padded, DMA_TO_DEVICE); + + if (!ret) { + if (copy_to_user(u64_to_user_ptr(req.output), m_buf, n_len)) + ret =3D -EFAULT; + } + +out_free: + kfree_sensitive(m_buf); + kfree(c_buf); + kfree(n_buf); + kfree(e_buf); + return ret; +} + +/** + * cmh_mgmt_pke_rsa_crt_dec() - Handle CMH_MGMT_IOC_PKE_RSA_CRT_DEC ioctl + * @argp: User-space ioctl argument pointer + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_pke_rsa_crt_dec(void __user *argp) +{ + u32 pke_cid =3D cmh_core_default_id(CMH_CORE_PKE); + + struct cmh_ioctl_pke_rsa_crt_dec req; + struct vcq_cmd vcq[PKE_VCQ_CMDS_MIN]; + u32 n_len, e_padded; + u8 *e_buf, *n_buf, *c_buf, *m_buf; + dma_addr_t e_dma, n_dma, c_dma, m_dma; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.__reserved) + return -EINVAL; + if (req.bits < PKE_RSA_MIN_BITS || req.bits > PKE_RSA_MAX_BITS || + req.bits % 8) + return -EINVAL; + if (!req.e_len || req.e_len > PKE_MAX_OPERAND) + return -EINVAL; + + n_len =3D req.bits / 8; + e_padded =3D ALIGN(req.e_len, 4); + + e_buf =3D kzalloc(e_padded, GFP_KERNEL); + n_buf =3D kmalloc(n_len, GFP_KERNEL); + c_buf =3D kmalloc(n_len, GFP_KERNEL); + m_buf =3D kzalloc(n_len, GFP_KERNEL); + if (!e_buf || !n_buf || !c_buf || !m_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + /* Right-align exponent in zero-padded buffer for DMA alignment */ + if (copy_from_user(e_buf + e_padded - req.e_len, + u64_to_user_ptr(req.e), req.e_len) || + copy_from_user(n_buf, u64_to_user_ptr(req.n), n_len) || + copy_from_user(c_buf, u64_to_user_ptr(req.input), n_len)) { + ret =3D -EFAULT; + goto out_free; + } + + e_dma =3D cmh_dma_map_single(e_buf, e_padded, DMA_TO_DEVICE); + n_dma =3D cmh_dma_map_single(n_buf, n_len, DMA_TO_DEVICE); + c_dma =3D cmh_dma_map_single(c_buf, n_len, DMA_TO_DEVICE); + m_dma =3D cmh_dma_map_single(m_buf, n_len, DMA_FROM_DEVICE); + if (cmh_dma_map_error(e_dma) || cmh_dma_map_error(n_dma) || + cmh_dma_map_error(c_dma) || cmh_dma_map_error(m_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + vcq_set_header(&vcq[0], PKE_VCQ_CMDS_MIN); + vcq_add_pke_rsa_crt_dec(&vcq[1], pke_cid, req.bits, e_padded, + e_dma, n_dma, c_dma, m_dma, req.crt_ref, + PKE_SWAP_FLAGS); + vcq_add_pke_flush(&vcq[2], pke_cid); + + ret =3D cmh_tm_submit_sync_mbx(vcq, PKE_VCQ_CMDS_MIN, 1, MGMT_MBX); + +out_unmap: + if (!cmh_dma_map_error(m_dma)) + cmh_dma_unmap_single(m_dma, n_len, DMA_FROM_DEVICE); + if (!cmh_dma_map_error(c_dma)) + cmh_dma_unmap_single(c_dma, n_len, DMA_TO_DEVICE); + if (!cmh_dma_map_error(n_dma)) + cmh_dma_unmap_single(n_dma, n_len, DMA_TO_DEVICE); + if (!cmh_dma_map_error(e_dma)) + cmh_dma_unmap_single(e_dma, e_padded, DMA_TO_DEVICE); + + if (!ret) { + if (copy_to_user(u64_to_user_ptr(req.output), m_buf, n_len)) + ret =3D -EFAULT; + } + +out_free: + kfree_sensitive(m_buf); + kfree(c_buf); + kfree(n_buf); + kfree(e_buf); + return ret; +} + +/** + * cmh_mgmt_pke_rsa_keygen() - Handle CMH_MGMT_IOC_PKE_RSA_KEYGEN ioctl + * @argp: User-space ioctl argument pointer + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_pke_rsa_keygen(void __user *argp) +{ + u32 pke_cid =3D cmh_core_default_id(CMH_CORE_PKE); + + struct cmh_ioctl_pke_rsa_keygen req; + /* + * When has_crt, we use a two-VCQ approach (CRI pattern): + * VCQ #1: header + SYS_NEW(d) + SYS_NEW(crt) + SYS_FLUSH (4 slots) + * VCQ #2: header + RSA_KEYGEN + PKE_FLUSH + SYS_FLUSH (4 slots) + * Without CRT, single VCQ: + * header + SYS_NEW(d) + RSA_KEYGEN + PKE_FLUSH + SYS_FLUSH (5 slots) + */ + struct vcq_cmd vcq[5]; + u32 n_len, e_padded, key_flags, d_ds_len, crt_ds_len; + u8 *e_buf, *n_buf; + u64 *d_ref_buf, *crt_ref_buf; + dma_addr_t e_dma =3D DMA_MAPPING_ERROR, n_dma =3D DMA_MAPPING_ERROR; + dma_addr_t d_ref_dma =3D DMA_MAPPING_ERROR; + dma_addr_t crt_ref_dma =3D DMA_MAPPING_ERROR; + int idx, ret; + bool has_crt, is_sca; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.bits < PKE_RSA_MIN_BITS || req.bits > PKE_RSA_MAX_BITS || + req.bits % 8) + return -EINVAL; + if (!req.e_len || req.e_len > PKE_MAX_OPERAND) + return -EINVAL; + if (req.flags & ~CMH_FLAG_MASK) + return -EINVAL; + + n_len =3D req.bits / 8; + has_crt =3D (req.crt_cid !=3D 0); + e_padded =3D ALIGN(req.e_len, 4); + key_flags =3D req.flags & CMH_FLAG_MASK; + is_sca =3D !!(req.flags & CMH_FLAG_SCA); + + /* + * SCA keys are stored in 2 shares -- DS allocation must be enlarged. + * CRI reference formulas: cmh_pke_rsa_private_key_size(). + */ + if (is_sca) { + d_ds_len =3D n_len * 2; + crt_ds_len =3D (7 + n_len / 2) * 4; + } else { + d_ds_len =3D n_len; + crt_ds_len =3D 5 * (n_len / 2); + } + + e_buf =3D kzalloc(e_padded, GFP_KERNEL); + n_buf =3D kzalloc(n_len, GFP_KERNEL); + d_ref_buf =3D kzalloc_obj(u64, GFP_KERNEL); + crt_ref_buf =3D kzalloc_obj(u64, GFP_KERNEL); + if (!e_buf || !n_buf || !d_ref_buf || !crt_ref_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + if (copy_from_user(e_buf + e_padded - req.e_len, + u64_to_user_ptr(req.e), req.e_len)) { + ret =3D -EFAULT; + goto out_free; + } + + e_dma =3D cmh_dma_map_single(e_buf, e_padded, DMA_TO_DEVICE); + n_dma =3D cmh_dma_map_single(n_buf, n_len, DMA_FROM_DEVICE); + d_ref_dma =3D cmh_dma_map_single(d_ref_buf, sizeof(u64), DMA_FROM_DEVICE); + crt_ref_dma =3D cmh_dma_map_single(crt_ref_buf, sizeof(u64), + DMA_FROM_DEVICE); + if (cmh_dma_map_error(e_dma) || cmh_dma_map_error(n_dma) || + cmh_dma_map_error(d_ref_dma) || cmh_dma_map_error(crt_ref_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + if (has_crt) { + /* + * Two-VCQ approach (CRI pattern): SYS_REF_LAST can only + * refer to the most recently created DS object. When we + * need both d and crt refs, we must first allocate DS + * objects, read back the opaque refs, then pass them by + * value in the keygen VCQ. + * + * VCQ #1: allocate both DS objects. + */ + idx =3D 0; + vcq_set_header(&vcq[idx++], 4); + vcq_add_sys_new(&vcq[idx++], req.d_cid, d_ref_dma, d_ds_len); + vcq_add_sys_new(&vcq[idx++], req.crt_cid, crt_ref_dma, + crt_ds_len); + vcq_add_sys_flush(&vcq[idx++]); + + ret =3D cmh_tm_submit_sync_mbx(vcq, 4, 1, MGMT_MBX); + if (ret) + goto out_unmap; + + /* Sync DMA so we can read back the opaque refs */ + cmh_dma_unmap_single(d_ref_dma, sizeof(u64), DMA_FROM_DEVICE); + cmh_dma_unmap_single(crt_ref_dma, sizeof(u64), + DMA_FROM_DEVICE); + d_ref_dma =3D DMA_MAPPING_ERROR; + crt_ref_dma =3D DMA_MAPPING_ERROR; + + /* + * VCQ #2: keygen with resolved refs. + */ + idx =3D 0; + memset(vcq, 0, sizeof(vcq)); + vcq_set_header(&vcq[idx++], 4); + + vcq[idx].magic =3D VCQ_CMD_MAGIC; + vcq[idx].id =3D VCQ_CMD_ID(pke_cid, PKE_SWAP_FLAGS, 1, + PKE_CMD_RSA_KEYGEN); + vcq[idx].hwc.pke.cmd_rsa_keygen.bits =3D req.bits; + vcq[idx].hwc.pke.cmd_rsa_keygen.e =3D e_dma; + vcq[idx].hwc.pke.cmd_rsa_keygen.n =3D n_dma; + vcq[idx].hwc.pke.cmd_rsa_keygen.d =3D *d_ref_buf; + vcq[idx].hwc.pke.cmd_rsa_keygen.d_type =3D + SYS_TYPE_SET(key_flags, CORE_ID_PKE); + vcq[idx].hwc.pke.cmd_rsa_keygen.crt =3D *crt_ref_buf; + vcq[idx].hwc.pke.cmd_rsa_keygen.crt_type =3D + SYS_TYPE_SET(key_flags, CORE_ID_PKE); + idx++; + + vcq_add_pke_flush(&vcq[idx++], pke_cid); + vcq_add_sys_flush(&vcq[idx++]); + + ret =3D cmh_tm_submit_sync_tmo(vcq, 4, 1, MGMT_MBX, + cmh_tm_slow_op_timeout_jiffies()); + } else { + /* + * Single-VCQ: only d, so SYS_REF_LAST is unambiguous. + */ + idx =3D 0; + vcq_set_header(&vcq[idx++], 5); + vcq_add_sys_new(&vcq[idx++], req.d_cid, d_ref_dma, d_ds_len); + + vcq[idx].magic =3D VCQ_CMD_MAGIC; + vcq[idx].id =3D VCQ_CMD_ID(pke_cid, PKE_SWAP_FLAGS, 1, + PKE_CMD_RSA_KEYGEN); + vcq[idx].hwc.pke.cmd_rsa_keygen.bits =3D req.bits; + vcq[idx].hwc.pke.cmd_rsa_keygen.e =3D e_dma; + vcq[idx].hwc.pke.cmd_rsa_keygen.n =3D n_dma; + vcq[idx].hwc.pke.cmd_rsa_keygen.d =3D SYS_REF_LAST; + vcq[idx].hwc.pke.cmd_rsa_keygen.d_type =3D + SYS_TYPE_SET(key_flags, CORE_ID_PKE); + vcq[idx].hwc.pke.cmd_rsa_keygen.crt =3D SYS_REF_NONE; + vcq[idx].hwc.pke.cmd_rsa_keygen.crt_type =3D 0; + idx++; + + vcq_add_pke_flush(&vcq[idx++], pke_cid); + vcq_add_sys_flush(&vcq[idx++]); + + ret =3D cmh_tm_submit_sync_tmo(vcq, 5, 1, MGMT_MBX, + cmh_tm_slow_op_timeout_jiffies()); + } + +out_unmap: + if (!cmh_dma_map_error(crt_ref_dma)) + cmh_dma_unmap_single(crt_ref_dma, sizeof(u64), + DMA_FROM_DEVICE); + if (!cmh_dma_map_error(d_ref_dma)) + cmh_dma_unmap_single(d_ref_dma, sizeof(u64), DMA_FROM_DEVICE); + if (!cmh_dma_map_error(n_dma)) + cmh_dma_unmap_single(n_dma, n_len, DMA_FROM_DEVICE); + if (!cmh_dma_map_error(e_dma)) + cmh_dma_unmap_single(e_dma, e_padded, DMA_TO_DEVICE); + + /* + * Keygen failed after SYS_CMD_NEW created the d/crt objects (VCQ #2 in + * the CRT path, or the keygen command in the single-VCQ path); scrub + * the orphans. Refs are zero unless the eSW wrote them back. + */ + if (ret && (*d_ref_buf || *crt_ref_buf)) + cmh_mgmt_ds_scrub(*d_ref_buf, *crt_ref_buf); + + if (!ret) { + /* Copy generated modulus and refs back */ + if (copy_to_user(u64_to_user_ptr(req.n), n_buf, n_len)) { + ret =3D -EFAULT; + } else { + req.d_ref =3D *d_ref_buf; + req.crt_ref =3D has_crt ? *crt_ref_buf : 0; + if (copy_to_user(argp, &req, sizeof(req))) + ret =3D -EFAULT; + } + + if (ret =3D=3D -EFAULT) { + /* + * The d (and CRT) DS slots were created but their + * refs never reached userspace, so the caller cannot + * free them. Logically delete the orphaned slots. + */ + cmh_mgmt_ds_scrub(*d_ref_buf, + has_crt ? *crt_ref_buf : 0); + dev_warn(cmh_dev(), + "mgmt: RSA_KEYGEN copy_to_user failed, DS slots cleaned up\n"); + } + } + +out_free: + kfree(crt_ref_buf); + kfree(d_ref_buf); + kfree(n_buf); + kfree(e_buf); + return ret; +} + +/** + * cmh_mgmt_pke_ecdsa_sign() - Handle CMH_MGMT_IOC_PKE_ECDSA_SIGN ioctl + * @argp: User-space ioctl argument pointer + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_pke_ecdsa_sign(void __user *argp) +{ + u32 pke_cid =3D cmh_core_default_id(CMH_CORE_PKE); + + struct cmh_ioctl_pke_ecdsa_sign req; + struct vcq_cmd vcq[PKE_VCQ_CMDS_MIN]; + u32 clen, sig_len, dig_map_len; + u8 *dig_buf, *sig_buf; + dma_addr_t dig_dma, sig_dma; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.__reserved) + return -EINVAL; + clen =3D cmh_pke_validate_curve(req.curve); + if (!clen || !req.digest_len || + req.digest_len > CMH_MGMT_MAX_DATA_LEN) + return -EINVAL; + + sig_len =3D 2 * clen; + + /* + * eSW requires digest_len >=3D clen. A hash shorter than clen must + * be LEFT-padded with zeros (right-aligned) so it keeps its value + * as a big-endian integer -- trailing zeros would multiply it by + * 256^k. When digest_len >=3D clen the offset is 0 (unchanged) and + * the eSW applies the ECDSA bits2int leftmost-bits truncation. + */ + dig_map_len =3D max_t(u32, req.digest_len, clen); + + dig_buf =3D kzalloc(dig_map_len, GFP_KERNEL); + sig_buf =3D kzalloc(sig_len, GFP_KERNEL); + if (!dig_buf || !sig_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + if (copy_from_user(dig_buf + (dig_map_len - req.digest_len), + u64_to_user_ptr(req.digest), req.digest_len)) { + ret =3D -EFAULT; + goto out_free; + } + + dig_dma =3D cmh_dma_map_single(dig_buf, dig_map_len, DMA_TO_DEVICE); + sig_dma =3D cmh_dma_map_single(sig_buf, sig_len, DMA_FROM_DEVICE); + if (cmh_dma_map_error(dig_dma) || cmh_dma_map_error(sig_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + vcq_set_header(&vcq[0], PKE_VCQ_CMDS_MIN); + vcq_add_pke_ecdsa_sign(&vcq[1], pke_cid, req.curve, clen, + dig_dma, sig_dma, req.key_ref, + dig_map_len, pke_swap_flags(req.curve)); + vcq_add_pke_flush(&vcq[2], pke_cid); + + ret =3D cmh_tm_submit_sync_mbx(vcq, PKE_VCQ_CMDS_MIN, 1, MGMT_MBX); + +out_unmap: + if (!cmh_dma_map_error(sig_dma)) + cmh_dma_unmap_single(sig_dma, sig_len, DMA_FROM_DEVICE); + if (!cmh_dma_map_error(dig_dma)) + cmh_dma_unmap_single(dig_dma, dig_map_len, DMA_TO_DEVICE); + + if (!ret) { + if (copy_to_user(u64_to_user_ptr(req.signature), + sig_buf, sig_len)) + ret =3D -EFAULT; + } + +out_free: + kfree(sig_buf); + kfree(dig_buf); + return ret; +} + +/** + * cmh_mgmt_pke_ecdh() - Handle CMH_MGMT_IOC_PKE_ECDH ioctl + * @argp: User-space ioctl argument pointer + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_pke_ecdh(void __user *argp) +{ + u32 pke_cid =3D cmh_core_default_id(CMH_CORE_PKE); + + struct cmh_ioctl_pke_ecdh req; + /* hdr + pke_ecdh(->SYS_REF_TEMP) + pke_flush + sys_data + sys_flush */ + struct vcq_cmd vcq[5]; + u32 clen, swap, ss_type; + u8 *peer_buf, *ss_buf; + dma_addr_t peer_dma, ss_dma; + int ret, idx; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + /* flags and reserved fields must be zero (raw host output only). */ + if (req.flags || req.__reserved || req.__reserved2) + return -EINVAL; + clen =3D cmh_pke_validate_curve(req.curve); + if (!clen) + return -EINVAL; + + swap =3D PKE_SWAP_FLAGS; + ss_type =3D SYS_TYPE_SET(SYS_TYPE_FLAG_PT, CORE_ID_PKE); + + peer_buf =3D kmalloc(clen, GFP_KERNEL); + ss_buf =3D kzalloc(clen, GFP_KERNEL); + if (!peer_buf || !ss_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + if (copy_from_user(peer_buf, u64_to_user_ptr(req.peer_key_x), clen)) { + ret =3D -EFAULT; + goto out_free; + } + + peer_dma =3D cmh_dma_map_single(peer_buf, clen, DMA_TO_DEVICE); + ss_dma =3D cmh_dma_map_single(ss_buf, clen, DMA_FROM_DEVICE); + if (cmh_dma_map_error(peer_dma) || cmh_dma_map_error(ss_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + /* + * Compute the shared secret into SYS_REF_TEMP and read it back in a + * single VCQ. The key is a caller-owned persistent DS ref, so it + * never collides with the temp output. Doing the temp push (pke_ecdh) + * and pop (sys_data) in one submission bounds the temp to a single + * mailbox occupation -- no concurrent mgmt op can interleave a temp -- + * and the read reclaims the slot, so no datastore object leaks. + */ + idx =3D 0; + vcq_set_header(&vcq[idx++], 5); + vcq_add_pke_ecdh(&vcq[idx++], pke_cid, req.curve, clen, clen, + ss_type, peer_dma, req.key_ref, + SYS_REF_TEMP, swap); + vcq_add_pke_flush(&vcq[idx++], pke_cid); + vcq_add_sys_data(&vcq[idx], SYS_REF_TEMP, ss_dma, clen); + vcq[idx++].id |=3D pke_swap_flags(req.curve); + vcq_add_sys_flush(&vcq[idx++]); + + ret =3D cmh_tm_submit_sync_mbx(vcq, 5, 1, MGMT_MBX); + +out_unmap: + if (!cmh_dma_map_error(ss_dma)) + cmh_dma_unmap_single(ss_dma, clen, DMA_FROM_DEVICE); + if (!cmh_dma_map_error(peer_dma)) + cmh_dma_unmap_single(peer_dma, clen, DMA_TO_DEVICE); + + if (!ret) { + if (copy_to_user(u64_to_user_ptr(req.output), ss_buf, clen)) + ret =3D -EFAULT; + } + +out_free: + kfree_sensitive(ss_buf); + kfree(peer_buf); + return ret; +} + +/** + * cmh_mgmt_pke_ecdh_keygen() - Handle CMH_MGMT_IOC_PKE_ECDH_KEYGEN ioctl + * @argp: User-space ioctl argument pointer + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_pke_ecdh_keygen(void __user *argp) +{ + u32 pke_cid =3D cmh_core_default_id(CMH_CORE_PKE); + + struct cmh_ioctl_pke_ecdh_keygen req; + struct vcq_cmd vcq[PKE_VCQ_CMDS_MIN]; + u32 clen, out_len; + u8 *pkx_buf; + dma_addr_t pkx_dma; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + clen =3D cmh_pke_validate_curve(req.curve); + if (!clen) + return -EINVAL; + + /* + * ECDH_KEYGEN always outputs both X and Y coordinates + * (2 * clen bytes total) even though only X is useful for + * the ECDH exchange. Allocate the full output size to avoid + * a DMA buffer overflow, but copy only X back to userspace. + */ + out_len =3D 2 * clen; + + pkx_buf =3D kzalloc(out_len, GFP_KERNEL); + if (!pkx_buf) + return -ENOMEM; + + pkx_dma =3D cmh_dma_map_single(pkx_buf, out_len, DMA_FROM_DEVICE); + if (cmh_dma_map_error(pkx_dma)) { + kfree(pkx_buf); + return -ENOMEM; + } + + vcq_set_header(&vcq[0], PKE_VCQ_CMDS_MIN); + vcq_add_pke_ecdh_keygen(&vcq[1], pke_cid, req.curve, clen, + pkx_dma, req.key_ref, + PKE_SWAP_FLAGS); + vcq_add_pke_flush(&vcq[2], pke_cid); + + ret =3D cmh_tm_submit_sync_mbx(vcq, PKE_VCQ_CMDS_MIN, 1, MGMT_MBX); + + cmh_dma_unmap_single(pkx_dma, out_len, DMA_FROM_DEVICE); + + if (!ret) { + if (copy_to_user(u64_to_user_ptr(req.public_key_x), + pkx_buf, clen)) + ret =3D -EFAULT; + } + + kfree(pkx_buf); + return ret; +} + +/** + * cmh_mgmt_pke_eddsa_sign() - Handle CMH_MGMT_IOC_PKE_EDDSA_SIGN ioctl + * @argp: User-space ioctl argument pointer + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_pke_eddsa_sign(void __user *argp) +{ + u32 pke_cid =3D cmh_core_default_id(CMH_CORE_PKE); + + struct cmh_ioctl_pke_eddsa_sign req; + struct vcq_cmd vcq[PKE_VCQ_CMDS_MIN]; + u32 klen, sig_len; + u8 *msg_buf, *sig_buf; + dma_addr_t msg_dma, sig_dma; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.__reserved) + return -EINVAL; + if (!cmh_pke_validate_curve(req.curve) || !req.digest_len || + req.digest_len > CMH_MGMT_MAX_DATA_LEN) + return -EINVAL; + if (!pke_curve_is_edwards(req.curve)) + return -EINVAL; + + klen =3D pke_eddsa_key_len(req.curve); + sig_len =3D 2 * klen; + + msg_buf =3D kmalloc(req.digest_len, GFP_KERNEL); + sig_buf =3D kzalloc(sig_len, GFP_KERNEL); + if (!msg_buf || !sig_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + if (copy_from_user(msg_buf, u64_to_user_ptr(req.digest), + req.digest_len)) { + ret =3D -EFAULT; + goto out_free; + } + + msg_dma =3D cmh_dma_map_single(msg_buf, req.digest_len, DMA_TO_DEVICE); + sig_dma =3D cmh_dma_map_single(sig_buf, sig_len, DMA_FROM_DEVICE); + if (cmh_dma_map_error(msg_dma) || cmh_dma_map_error(sig_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + vcq_set_header(&vcq[0], PKE_VCQ_CMDS_MIN); + vcq_add_pke_eddsa_sign(&vcq[1], pke_cid, req.curve, klen, + msg_dma, sig_dma, req.key_ref, + req.digest_len, pke_swap_flags(req.curve)); + vcq_add_pke_flush(&vcq[2], pke_cid); + + ret =3D cmh_tm_submit_sync_mbx(vcq, PKE_VCQ_CMDS_MIN, 1, MGMT_MBX); + +out_unmap: + if (!cmh_dma_map_error(sig_dma)) + cmh_dma_unmap_single(sig_dma, sig_len, DMA_FROM_DEVICE); + if (!cmh_dma_map_error(msg_dma)) + cmh_dma_unmap_single(msg_dma, req.digest_len, DMA_TO_DEVICE); + + if (!ret) { + if (copy_to_user(u64_to_user_ptr(req.signature), + sig_buf, sig_len)) + ret =3D -EFAULT; + } + +out_free: + kfree(sig_buf); + kfree(msg_buf); + return ret; +} + +/** + * cmh_mgmt_pke_eddsa_verify() - Handle CMH_MGMT_IOC_PKE_EDDSA_VERIFY ioctl + * @argp: User-space ioctl argument pointer + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_pke_eddsa_verify(void __user *argp) +{ + u32 pke_cid =3D cmh_core_default_id(CMH_CORE_PKE); + + struct cmh_ioctl_pke_eddsa_verify req; + struct vcq_cmd vcq[PKE_VCQ_CMDS_MIN]; + u32 clen, klen, sig_len; + u8 *msg_buf, *sig_buf, *pky_buf, *rp_buf; + dma_addr_t msg_dma, sig_dma, pky_dma, rp_dma; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.__reserved) + return -EINVAL; + clen =3D cmh_pke_validate_curve(req.curve); + if (!clen || !req.digest_len || + req.digest_len > CMH_MGMT_MAX_DATA_LEN) + return -EINVAL; + if (!pke_curve_is_edwards(req.curve)) + return -EINVAL; + + klen =3D pke_eddsa_key_len(req.curve); + sig_len =3D 2 * klen; + + msg_buf =3D kmalloc(req.digest_len, GFP_KERNEL); + sig_buf =3D kmalloc(sig_len, GFP_KERNEL); + pky_buf =3D kmalloc(klen, GFP_KERNEL); + rp_buf =3D kzalloc(clen, GFP_KERNEL); + if (!msg_buf || !sig_buf || !pky_buf || !rp_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + if (copy_from_user(msg_buf, u64_to_user_ptr(req.digest), + req.digest_len) || + copy_from_user(sig_buf, u64_to_user_ptr(req.signature), + sig_len) || + copy_from_user(pky_buf, u64_to_user_ptr(req.public_key_y), + klen)) { + ret =3D -EFAULT; + goto out_free; + } + + msg_dma =3D cmh_dma_map_single(msg_buf, req.digest_len, DMA_TO_DEVICE); + sig_dma =3D cmh_dma_map_single(sig_buf, sig_len, DMA_TO_DEVICE); + pky_dma =3D cmh_dma_map_single(pky_buf, klen, DMA_TO_DEVICE); + rp_dma =3D cmh_dma_map_single(rp_buf, clen, DMA_FROM_DEVICE); + if (cmh_dma_map_error(msg_dma) || cmh_dma_map_error(sig_dma) || + cmh_dma_map_error(pky_dma) || cmh_dma_map_error(rp_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + vcq_set_header(&vcq[0], PKE_VCQ_CMDS_MIN); + vcq_add_pke_eddsa_verify(&vcq[1], pke_cid, req.curve, req.digest_len, + pky_dma, msg_dma, sig_dma, rp_dma, + pke_swap_flags(req.curve)); + vcq_add_pke_flush(&vcq[2], pke_cid); + + ret =3D cmh_tm_submit_sync_mbx(vcq, PKE_VCQ_CMDS_MIN, 1, MGMT_MBX); + +out_unmap: + if (!cmh_dma_map_error(rp_dma)) + cmh_dma_unmap_single(rp_dma, clen, DMA_FROM_DEVICE); + if (!cmh_dma_map_error(pky_dma)) + cmh_dma_unmap_single(pky_dma, klen, DMA_TO_DEVICE); + if (!cmh_dma_map_error(sig_dma)) + cmh_dma_unmap_single(sig_dma, sig_len, DMA_TO_DEVICE); + if (!cmh_dma_map_error(msg_dma)) + cmh_dma_unmap_single(msg_dma, req.digest_len, DMA_TO_DEVICE); + +out_free: + kfree(rp_buf); + kfree(pky_buf); + kfree(sig_buf); + kfree(msg_buf); + return ret; +} + +/** + * cmh_mgmt_pke_ec_keygen() - Handle CMH_MGMT_IOC_PKE_EC_KEYGEN ioctl + * @argp: User-space ioctl argument pointer + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_pke_ec_keygen(void __user *argp) +{ + u32 pke_cid =3D cmh_core_default_id(CMH_CORE_PKE); + + struct cmh_ioctl_pke_ec_keygen req; + /* header + SYS_NEW + ECDSA_KEYGEN + flush_pke + flush_sys */ + struct vcq_cmd vcq[5]; + u32 clen, key_flags, ds_len; + u64 *ref_buf; + dma_addr_t ref_dma; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.__reserved) + return -EINVAL; + if (req.flags & ~CMH_FLAG_MASK) + return -EINVAL; + clen =3D cmh_pke_validate_curve(req.curve); + if (!clen) + return -EINVAL; + + key_flags =3D req.flags & CMH_FLAG_MASK; + /* SCA keys are stored in 2 shares -- allocate double the curve length */ + ds_len =3D (req.flags & CMH_FLAG_SCA) ? clen * 2 : clen; + + ref_buf =3D kzalloc_obj(u64, GFP_KERNEL); + if (!ref_buf) + return -ENOMEM; + + ref_dma =3D cmh_dma_map_single(ref_buf, sizeof(u64), DMA_FROM_DEVICE); + if (cmh_dma_map_error(ref_dma)) { + kfree(ref_buf); + return -ENOMEM; + } + + vcq_set_header(&vcq[0], 5); + vcq_add_sys_new(&vcq[1], req.cid, ref_dma, ds_len); + vcq_add_pke_ecdsa_keygen(&vcq[2], pke_cid, req.curve, clen, + SYS_REF_LAST, + SYS_TYPE_SET(key_flags, CORE_ID_PKE), + pke_swap_flags(req.curve)); + vcq_add_pke_flush(&vcq[3], pke_cid); + vcq_add_sys_flush(&vcq[4]); + + ret =3D cmh_tm_submit_sync_mbx(vcq, 5, 1, MGMT_MBX); + + cmh_dma_unmap_single(ref_dma, sizeof(u64), DMA_FROM_DEVICE); + + /* Keygen failed after SYS_CMD_NEW; scrub the orphaned slot. */ + if (ret && *ref_buf) + cmh_mgmt_ds_scrub(*ref_buf, 0); + + if (!ret) { + req.ref =3D *ref_buf; + if (copy_to_user(argp, &req, sizeof(req))) { + /* + * The DS slot was created but its ref never reached + * userspace, so the caller cannot free it. Logically + * delete the orphaned slot before returning. + */ + cmh_mgmt_ds_scrub(*ref_buf, 0); + dev_warn(cmh_dev(), + "mgmt: PKE_EC_KEYGEN copy_to_user failed, DS slot cleaned up\n"); + ret =3D -EFAULT; + } + } + + kfree(ref_buf); + return ret; +} + +/** + * cmh_mgmt_pke_ec_pubgen() - Handle CMH_MGMT_IOC_PKE_EC_PUBGEN ioctl + * @argp: User-space ioctl argument pointer + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_pke_ec_pubgen(void __user *argp) +{ + u32 pke_cid =3D cmh_core_default_id(CMH_CORE_PKE); + + struct cmh_ioctl_pke_ec_pubgen req; + struct vcq_cmd vcq[PKE_VCQ_CMDS_MIN]; + u32 clen, pk_len; + u8 *pk_buf; + dma_addr_t pk_dma; + bool is_ed; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + clen =3D cmh_pke_validate_curve(req.curve); + if (!clen) + return -EINVAL; + + is_ed =3D pke_curve_is_edwards(req.curve); + pk_len =3D is_ed ? pke_eddsa_key_len(req.curve) : 2 * clen; + + pk_buf =3D kzalloc(pk_len, GFP_KERNEL); + if (!pk_buf) + return -ENOMEM; + + pk_dma =3D cmh_dma_map_single(pk_buf, pk_len, DMA_FROM_DEVICE); + if (cmh_dma_map_error(pk_dma)) { + kfree(pk_buf); + return -ENOMEM; + } + + vcq_set_header(&vcq[0], PKE_VCQ_CMDS_MIN); + if (is_ed) + vcq_add_pke_eddsa_pubgen(&vcq[1], pke_cid, req.curve, + pke_eddsa_key_len(req.curve), + pk_dma, req.key_ref, + pke_swap_flags(req.curve)); + else + vcq_add_pke_ecdsa_pubgen(&vcq[1], pke_cid, req.curve, clen, + pk_dma, req.key_ref, + pke_swap_flags(req.curve)); + vcq_add_pke_flush(&vcq[2], pke_cid); + + ret =3D cmh_tm_submit_sync_mbx(vcq, PKE_VCQ_CMDS_MIN, 1, MGMT_MBX); + + cmh_dma_unmap_single(pk_dma, pk_len, DMA_FROM_DEVICE); + + if (!ret) { + if (copy_to_user(u64_to_user_ptr(req.public_key), + pk_buf, pk_len)) + ret =3D -EFAULT; + } + + kfree(pk_buf); + return ret; +} + +/** + * cmh_mgmt_pke_eddsa_keygen_sca() - Handle CMH_MGMT_IOC_PKE_EDDSA_KEYGEN_= SCA ioctl + * @argp: User-space ioctl argument pointer + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_pke_eddsa_keygen_sca(void __user *argp) +{ + u32 pke_cid =3D cmh_core_default_id(CMH_CORE_PKE); + + struct cmh_ioctl_pke_eddsa_keygen_sca req; + /* header + SYS_NEW + EDDSA_KEYGEN_SCA + flush_pke + flush_sys */ + struct vcq_cmd vcq[5]; + u64 *ref_buf; + dma_addr_t ref_dma; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + /* EdDSA SCA keygen is only supported for Ed448 */ + if (req.curve !=3D PKE_CURVE_448) + return -EINVAL; + + ref_buf =3D kzalloc_obj(u64, GFP_KERNEL); + if (!ref_buf) + return -ENOMEM; + + ref_dma =3D cmh_dma_map_single(ref_buf, sizeof(u64), DMA_FROM_DEVICE); + if (cmh_dma_map_error(ref_dma)) { + kfree(ref_buf); + return -ENOMEM; + } + + vcq_set_header(&vcq[0], 5); + vcq_add_sys_new(&vcq[1], req.cid, ref_dma, PKE_ED448_SK_SCA_LEN); + vcq_add_pke_eddsa_keygen_sca(&vcq[2], pke_cid, req.curve, req.key_ref, + SYS_REF_LAST); + vcq_add_pke_flush(&vcq[3], pke_cid); + vcq_add_sys_flush(&vcq[4]); + + ret =3D cmh_tm_submit_sync_mbx(vcq, 5, 1, MGMT_MBX); + + cmh_dma_unmap_single(ref_dma, sizeof(u64), DMA_FROM_DEVICE); + + /* Keygen failed after SYS_CMD_NEW; scrub the orphaned SCA key slot. */ + if (ret && *ref_buf) + cmh_mgmt_ds_scrub(*ref_buf, 0); + + if (!ret) { + req.sca_ref =3D *ref_buf; + if (copy_to_user(argp, &req, sizeof(req))) { + /* Ref never reached user space; scrub the orphan. */ + cmh_mgmt_ds_scrub(*ref_buf, 0); + dev_warn(cmh_dev(), + "mgmt: EDDSA_KEYGEN_SCA copy_to_user failed, DS slot cleaned up\n"); + ret =3D -EFAULT; + } + } + + kfree(ref_buf); + return ret; +} diff --git a/drivers/crypto/cmh/cmh_mgmt_pqc.c b/drivers/crypto/cmh/cmh_mgm= t_pqc.c new file mode 100644 index 000000000000..bc78e09078e2 --- /dev/null +++ b/drivers/crypto/cmh/cmh_mgmt_pqc.c @@ -0,0 +1,1365 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH -- PQC ioctl handlers for /dev/cmh_mgmt + * + * ML-KEM keygen/encapsulate/decapsulate, ML-DSA keygen/sign, + * SLH-DSA keygen/sign (pure + prehash). + * + * Split from cmh_mgmt.c for maintainability. + */ + +#include +#include +#include +#include + +#include "cmh_mgmt.h" +#include "cmh_sys.h" +#include "cmh_txn.h" +#include "cmh_key.h" +#include "cmh_dma.h" +#include "cmh_config.h" +#include "cmh_pqc.h" +#include "cmh_qse_abi.h" +#include "cmh_sys_abi.h" +#include + +#include + +/* -- PQC -- ML-KEM -- */ + +/** + * cmh_mgmt_ml_kem_keygen() - Handle CMH_MGMT_IOC_ML_KEM_KEYGEN ioctl + * @argp: User-space ioctl argument pointer + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_ml_kem_keygen(void __user *argp) +{ + u32 qse_cid =3D cmh_core_default_id(CMH_CORE_QSE); + + struct cmh_ioctl_ml_kem_keygen req; + struct vcq_cmd vcq[QSE_VCQ_CMDS_MAX]; + u32 ek_len, dk_len, seed_len, key_flags; + u32 qse_flags =3D 0; + bool masked, ds_ref, hw_rng; + u8 *seed_buf =3D NULL, *z_buf =3D NULL, *ek_buf, *dk_buf =3D NULL; + u64 *ref_buf =3D NULL; + dma_addr_t seed_dma =3D DMA_MAPPING_ERROR, z_dma =3D DMA_MAPPING_ERROR; + dma_addr_t ek_dma, dk_dma =3D DMA_MAPPING_ERROR, ref_dma =3D DMA_MAPPING_= ERROR; + int ret, idx; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.__reserved) + return -EINVAL; + if (ml_kem_k_idx(req.k) < 0) + return -EINVAL; + if (req.flags & ~(CMH_QSE_FLAG_MASK | CMH_FLAG_MASK)) + return -EINVAL; + + masked =3D !!(req.flags & CMH_QSE_FLAG_MASKED); + ds_ref =3D !!(req.flags & CMH_QSE_FLAG_DS_REF); + hw_rng =3D !!(req.flags & CMH_QSE_FLAG_HW_RNG); + + /* + * QSE keys only support PT storage -- the eSW dec/sign paths + * hardcode SYS_TYPE_FLAG_PT when reading the key back. + * QSE SCA protection uses masking (CMH_QSE_FLAG_MASKED), + * not the 2-share mechanism (CMH_FLAG_SCA). + */ + key_flags =3D req.flags & CMH_FLAG_MASK; + if (key_flags && key_flags !=3D CMH_FLAG_PT) + return -EINVAL; + key_flags =3D CMH_FLAG_PT; + + /* Masked keygen must store dk in DS -- polynomial unmasking not supporte= d */ + if (masked && !ds_ref) + return -EINVAL; + + ek_len =3D ML_KEM_EK_SIZE(req.k); + dk_len =3D masked ? ML_KEM_DK_SIZE_MASKED(req.k) + : ML_KEM_DK_SIZE(req.k); + seed_len =3D masked ? QSE_SEED_LEN_MASKED : QSE_SEED_LEN; + + if (hw_rng) + qse_flags |=3D QSE_FLAG_USE_RNG; + if (ds_ref) + qse_flags |=3D QSE_FLAG_USE_REF; + + /* + * Without HW RNG the caller must supply both seed and z; a 0 pointer + * with non-zero length is a wild eSW DMA read (F3), not "absent". + */ + if (!hw_rng && (!req.seed || !req.z)) + return -EINVAL; + + ek_buf =3D kzalloc(ek_len, GFP_KERNEL); + if (!ek_buf) + return -ENOMEM; + + if (!hw_rng && req.seed && req.z) { + seed_buf =3D kmalloc(seed_len, GFP_KERNEL); + z_buf =3D kmalloc(seed_len, GFP_KERNEL); + if (!seed_buf || !z_buf) { + ret =3D -ENOMEM; + goto out_free; + } + if (copy_from_user(seed_buf, u64_to_user_ptr(req.seed), + seed_len) || + copy_from_user(z_buf, u64_to_user_ptr(req.z), seed_len)) { + ret =3D -EFAULT; + goto out_free; + } + } + + if (ds_ref) { + ref_buf =3D kzalloc_obj(u64, GFP_KERNEL); + if (!ref_buf) { + ret =3D -ENOMEM; + goto out_free; + } + } else { + dk_buf =3D kzalloc(dk_len, GFP_KERNEL); + if (!dk_buf) { + ret =3D -ENOMEM; + goto out_free; + } + } + + /* DMA map */ + ek_dma =3D cmh_dma_map_single(ek_buf, ek_len, DMA_FROM_DEVICE); + if (cmh_dma_map_error(ek_dma)) { + ret =3D -ENOMEM; + goto out_free; + } + + if (seed_buf) { + seed_dma =3D cmh_dma_map_single(seed_buf, seed_len, + DMA_TO_DEVICE); + z_dma =3D cmh_dma_map_single(z_buf, seed_len, DMA_TO_DEVICE); + if (cmh_dma_map_error(seed_dma) || cmh_dma_map_error(z_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + } + + if (ds_ref) { + ref_dma =3D cmh_dma_map_single(ref_buf, sizeof(u64), + DMA_FROM_DEVICE); + if (cmh_dma_map_error(ref_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + } else { + dk_dma =3D cmh_dma_map_single(dk_buf, dk_len, DMA_FROM_DEVICE); + if (cmh_dma_map_error(dk_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + } + + idx =3D 0; + if (ds_ref) { + vcq_set_header(&vcq[0], QSE_VCQ_CMDS_MAX); + idx++; + vcq_add_sys_new(&vcq[idx++], req.dk_cid, ref_dma, dk_len); + vcq_add_qse_ml_kem_keygen(&vcq[idx++], qse_cid, req.k, qse_flags, + seed_dma, z_dma, + ek_dma, SYS_REF_LAST, + SYS_TYPE_SET(key_flags, + CORE_ID_QSE), + masked); + vcq_add_qse_flush(&vcq[idx++], qse_cid); + ret =3D cmh_tm_submit_sync_mbx(vcq, QSE_VCQ_CMDS_MAX, + 1, MGMT_MBX); + } else { + vcq_set_header(&vcq[0], QSE_VCQ_CMDS_MIN); + idx++; + vcq_add_qse_ml_kem_keygen(&vcq[idx++], qse_cid, req.k, qse_flags, + seed_dma, z_dma, + ek_dma, dk_dma, 0, masked); + vcq_add_qse_flush(&vcq[idx++], qse_cid); + ret =3D cmh_tm_submit_sync_mbx(vcq, QSE_VCQ_CMDS_MIN, + 1, MGMT_MBX); + } + +out_unmap: + if (ds_ref && !cmh_dma_map_error(ref_dma)) + cmh_dma_unmap_single(ref_dma, sizeof(u64), DMA_FROM_DEVICE); + if (!ds_ref && dk_buf && !cmh_dma_map_error(dk_dma)) + cmh_dma_unmap_single(dk_dma, dk_len, DMA_FROM_DEVICE); + if (z_buf && !cmh_dma_map_error(z_dma)) + cmh_dma_unmap_single(z_dma, seed_len, DMA_TO_DEVICE); + if (seed_buf && !cmh_dma_map_error(seed_dma)) + cmh_dma_unmap_single(seed_dma, seed_len, DMA_TO_DEVICE); + if (!cmh_dma_map_error(ek_dma)) + cmh_dma_unmap_single(ek_dma, ek_len, DMA_FROM_DEVICE); + + /* Keygen failed after SYS_CMD_NEW (ds_ref mode); scrub the orphan. */ + if (ret && ds_ref && *ref_buf) + cmh_mgmt_ds_scrub(*ref_buf, 0); + + if (!ret) { + if (copy_to_user(u64_to_user_ptr(req.ek), ek_buf, ek_len)) { + ret =3D -EFAULT; + } else if (ds_ref) { + req.dk_ref =3D *ref_buf; + } else if (copy_to_user(u64_to_user_ptr(req.dk), + dk_buf, dk_len)) { + ret =3D -EFAULT; + } + if (!ret && copy_to_user(argp, &req, sizeof(req))) + ret =3D -EFAULT; + /* + * A copy_to_user faulted after the DS slot was created, so + * its ref never reached userspace and the caller cannot free + * it. Logically delete the orphaned slot: grant-with-no-access + * wipes the key material and clears the CID (it does not + * reclaim the stack space -- only a datastore reset does). + */ + if (ret && ds_ref) { + cmh_mgmt_ds_scrub(*ref_buf, 0); + dev_warn(cmh_dev(), + "mgmt: ML_KEM_KEYGEN copy_to_user failed, DS slot cleaned up\n"); + } + } + +out_free: + kfree_sensitive(dk_buf); + kfree(ref_buf); + kfree_sensitive(z_buf); + kfree_sensitive(seed_buf); + kfree(ek_buf); + return ret; +} + +/** + * cmh_mgmt_ml_kem_enc() - Handle CMH_MGMT_IOC_ML_KEM_ENC ioctl + * @argp: User-space ioctl argument pointer + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_ml_kem_enc(void __user *argp) +{ + u32 qse_cid =3D cmh_core_default_id(CMH_CORE_QSE); + + struct cmh_ioctl_ml_kem_enc req; + struct vcq_cmd vcq[QSE_VCQ_CMDS_MIN]; + u32 ek_len, ct_len, ss_out_len; + u32 qse_flags =3D 0; + bool masked, hw_rng; + u8 *ek_buf, *coin_buf =3D NULL, *ct_buf, *ss_buf; + dma_addr_t ek_dma, coin_dma =3D DMA_MAPPING_ERROR, ct_dma, ss_dma; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.__reserved || req.__reserved2[0] || req.__reserved2[1]) + return -EINVAL; + if (ml_kem_k_idx(req.k) < 0) + return -EINVAL; + + masked =3D !!(req.flags & CMH_QSE_FLAG_MASKED); + hw_rng =3D !!(req.flags & CMH_QSE_FLAG_HW_RNG); + + ek_len =3D ML_KEM_EK_SIZE(req.k); + ct_len =3D ML_KEM_CT_SIZE(req.k); + ss_out_len =3D masked ? ML_KEM_SS_LEN_MASKED : ML_KEM_SS_LEN; + + if (hw_rng) + qse_flags |=3D QSE_FLAG_USE_RNG; + + /* + * Without HW RNG the caller must supply the coin; a 0 pointer with + * non-zero length is a wild eSW DMA read (F3), not "absent". + */ + if (!hw_rng && !req.coin) + return -EINVAL; + + ek_buf =3D kmalloc(ek_len, GFP_KERNEL); + ct_buf =3D kzalloc(ct_len, GFP_KERNEL); + ss_buf =3D kzalloc(ss_out_len, GFP_KERNEL); + if (!ek_buf || !ct_buf || !ss_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + if (copy_from_user(ek_buf, u64_to_user_ptr(req.ek), ek_len)) { + ret =3D -EFAULT; + goto out_free; + } + + if (!hw_rng && req.coin) { + u32 coin_len =3D masked ? QSE_SEED_LEN_MASKED : QSE_SEED_LEN; + + coin_buf =3D kmalloc(coin_len, GFP_KERNEL); + if (!coin_buf) { + ret =3D -ENOMEM; + goto out_free; + } + if (copy_from_user(coin_buf, u64_to_user_ptr(req.coin), + coin_len)) { + ret =3D -EFAULT; + goto out_free; + } + coin_dma =3D cmh_dma_map_single(coin_buf, coin_len, + DMA_TO_DEVICE); + if (cmh_dma_map_error(coin_dma)) { + ret =3D -ENOMEM; + goto out_free; + } + } + + ek_dma =3D cmh_dma_map_single(ek_buf, ek_len, DMA_TO_DEVICE); + ct_dma =3D cmh_dma_map_single(ct_buf, ct_len, DMA_FROM_DEVICE); + ss_dma =3D cmh_dma_map_single(ss_buf, ss_out_len, DMA_FROM_DEVICE); + if (cmh_dma_map_error(ek_dma) || cmh_dma_map_error(ct_dma) || + cmh_dma_map_error(ss_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + vcq_set_header(&vcq[0], QSE_VCQ_CMDS_MIN); + vcq_add_qse_ml_kem_enc(&vcq[1], qse_cid, req.k, qse_flags, + coin_dma, ek_dma, ct_dma, ss_dma, 0, masked); + vcq_add_qse_flush(&vcq[2], qse_cid); + + ret =3D cmh_tm_submit_sync_mbx(vcq, QSE_VCQ_CMDS_MIN, 1, MGMT_MBX); + +out_unmap: + if (!cmh_dma_map_error(ss_dma)) + cmh_dma_unmap_single(ss_dma, ss_out_len, DMA_FROM_DEVICE); + if (!cmh_dma_map_error(ct_dma)) + cmh_dma_unmap_single(ct_dma, ct_len, DMA_FROM_DEVICE); + if (!cmh_dma_map_error(ek_dma)) + cmh_dma_unmap_single(ek_dma, ek_len, DMA_TO_DEVICE); + if (coin_buf && !cmh_dma_map_error(coin_dma)) + cmh_dma_unmap_single(coin_dma, + masked ? QSE_SEED_LEN_MASKED + : QSE_SEED_LEN, + DMA_TO_DEVICE); + + if (!ret) { + if (copy_to_user(u64_to_user_ptr(req.ct), ct_buf, ct_len)) { + ret =3D -EFAULT; + goto out_free; + } + /* Unmask ss if masked: ss =3D share0 ^ share1 */ + if (masked) { + crypto_xor(ss_buf, ss_buf + ML_KEM_SS_LEN, + ML_KEM_SS_LEN); + } + if (copy_to_user(u64_to_user_ptr(req.ss), ss_buf, + ML_KEM_SS_LEN)) { + ret =3D -EFAULT; + goto out_free; + } + if (copy_to_user(argp, &req, sizeof(req))) + ret =3D -EFAULT; + } + +out_free: + kfree_sensitive(ss_buf); + kfree(ct_buf); + kfree(coin_buf); + kfree(ek_buf); + return ret; +} + +/** + * cmh_mgmt_ml_kem_dec() - Handle CMH_MGMT_IOC_ML_KEM_DEC ioctl + * @argp: User-space ioctl argument pointer + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_ml_kem_dec(void __user *argp) +{ + u32 qse_cid =3D cmh_core_default_id(CMH_CORE_QSE); + + struct cmh_ioctl_ml_kem_dec req; + /* DS_REF path: hdr + ml_kem_dec + qse_flush + sys_data + sys_flush */ + struct vcq_cmd vcq[5]; + u32 ct_len, dk_len, ss_out_len; + u32 qse_flags =3D 0; + bool masked, ds_ref; + u8 *ct_buf, *dk_buf =3D NULL, *ss_buf; + dma_addr_t ct_dma, dk_dma =3D DMA_MAPPING_ERROR, ss_dma; + u64 dk_ref; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.__reserved || req.__reserved2[0] || req.__reserved2[1]) + return -EINVAL; + if (ml_kem_k_idx(req.k) < 0) + return -EINVAL; + + masked =3D !!(req.flags & CMH_QSE_FLAG_MASKED); + ds_ref =3D !!(req.flags & CMH_QSE_FLAG_DS_REF); + + ct_len =3D ML_KEM_CT_SIZE(req.k); + dk_len =3D masked ? ML_KEM_DK_SIZE_MASKED(req.k) + : ML_KEM_DK_SIZE(req.k); + ss_out_len =3D masked ? ML_KEM_SS_LEN_MASKED : ML_KEM_SS_LEN; + + ct_buf =3D kmalloc(ct_len, GFP_KERNEL); + ss_buf =3D kzalloc(ss_out_len, GFP_KERNEL); + if (!ct_buf || !ss_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + if (copy_from_user(ct_buf, u64_to_user_ptr(req.ct), ct_len)) { + ret =3D -EFAULT; + goto out_free; + } + + ct_dma =3D cmh_dma_map_single(ct_buf, ct_len, DMA_TO_DEVICE); + ss_dma =3D cmh_dma_map_single(ss_buf, ss_out_len, DMA_FROM_DEVICE); + if (cmh_dma_map_error(ct_dma) || cmh_dma_map_error(ss_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + /* + * dk: if DS_REF flag is set, req.dk is a DS reference. + * Otherwise, copy raw dk from user-space and use extmem DMA. + * Masked decaps requires DS ref (polynomial unmasking not supported). + */ + if (ds_ref) { + dk_ref =3D req.dk; + qse_flags |=3D QSE_FLAG_USE_REF; + } else { + if (masked) { + ret =3D -EINVAL; + goto out_unmap; + } + dk_buf =3D kmalloc(dk_len, GFP_KERNEL); + if (!dk_buf) { + ret =3D -ENOMEM; + goto out_unmap; + } + if (copy_from_user(dk_buf, u64_to_user_ptr(req.dk), dk_len)) { + ret =3D -EFAULT; + goto out_unmap; + } + dk_dma =3D cmh_dma_map_single(dk_buf, dk_len, DMA_TO_DEVICE); + if (cmh_dma_map_error(dk_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + dk_ref =3D dk_dma; + } + + if (ds_ref) { + /* + * DS_REF decaps: dk and ss both live in the datastore. + * Compute ss into SYS_REF_TEMP and read it back within a + * single VCQ. dk is a caller-owned persistent ref, so it + * never collides with the temp output. Keeping the temp push + * (ml_kem_dec) and pop (sys_data) in one submission bounds the + * temp to a single mailbox occupation -- no concurrent mgmt op + * can interleave a temp -- and the read reclaims the slot, so + * nothing leaks. qse_flush drains the QSE sidecar (committing + * the shared-secret DMA) without resetting the temp stack. + */ + vcq_set_header(&vcq[0], 5); + vcq_add_qse_ml_kem_dec(&vcq[1], qse_cid, req.k, qse_flags, + ct_dma, dk_ref, SYS_REF_TEMP, + SYS_TYPE_SET(SYS_TYPE_FLAG_PT, + CORE_ID_QSE), + masked); + vcq_add_qse_flush(&vcq[2], qse_cid); + vcq_add_sys_data(&vcq[3], SYS_REF_TEMP, ss_dma, ss_out_len); + vcq_add_sys_flush(&vcq[4]); + + ret =3D cmh_tm_submit_sync_mbx(vcq, 5, 1, MGMT_MBX); + } else { + vcq_set_header(&vcq[0], QSE_VCQ_CMDS_MIN); + vcq_add_qse_ml_kem_dec(&vcq[1], qse_cid, req.k, qse_flags, + ct_dma, dk_ref, ss_dma, 0, masked); + vcq_add_qse_flush(&vcq[2], qse_cid); + + ret =3D cmh_tm_submit_sync_mbx(vcq, QSE_VCQ_CMDS_MIN, + 1, MGMT_MBX); + } + +out_unmap: + if (dk_buf && !cmh_dma_map_error(dk_dma)) + cmh_dma_unmap_single(dk_dma, dk_len, DMA_TO_DEVICE); + if (!cmh_dma_map_error(ss_dma)) + cmh_dma_unmap_single(ss_dma, ss_out_len, DMA_FROM_DEVICE); + if (!cmh_dma_map_error(ct_dma)) + cmh_dma_unmap_single(ct_dma, ct_len, DMA_TO_DEVICE); + + if (!ret) { + if (masked) { + crypto_xor(ss_buf, ss_buf + ML_KEM_SS_LEN, + ML_KEM_SS_LEN); + } + if (copy_to_user(u64_to_user_ptr(req.ss), ss_buf, + ML_KEM_SS_LEN)) { + ret =3D -EFAULT; + goto out_free; + } + if (copy_to_user(argp, &req, sizeof(req))) + ret =3D -EFAULT; + } + +out_free: + kfree_sensitive(dk_buf); + kfree_sensitive(ss_buf); + kfree(ct_buf); + return ret; +} + +/* -- PQC -- ML-DSA -- */ + +/** + * cmh_mgmt_ml_dsa_keygen() - Handle CMH_MGMT_IOC_ML_DSA_KEYGEN ioctl + * @argp: User-space ioctl argument pointer + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_ml_dsa_keygen(void __user *argp) +{ + u32 qse_cid =3D cmh_core_default_id(CMH_CORE_QSE); + + struct cmh_ioctl_ml_dsa_keygen req; + struct vcq_cmd vcq[QSE_VCQ_CMDS_MAX]; + u32 pk_size, sk_size, seed_len, key_flags; + u32 qse_flags =3D 0; + bool masked, ds_ref, hw_rng; + u8 *seed_buf =3D NULL, *pk_buf, *sk_buf =3D NULL; + u64 *ref_buf =3D NULL; + dma_addr_t seed_dma =3D DMA_MAPPING_ERROR, pk_dma; + dma_addr_t sk_dma =3D DMA_MAPPING_ERROR, ref_dma =3D DMA_MAPPING_ERROR; + int ret, idx, mi; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.__reserved) + return -EINVAL; + mi =3D ml_dsa_mode_idx(req.mode); + if (mi < 0) + return -EINVAL; + if (req.flags & ~(CMH_QSE_FLAG_MASK | CMH_FLAG_MASK)) + return -EINVAL; + + masked =3D !!(req.flags & CMH_QSE_FLAG_MASKED); + ds_ref =3D !!(req.flags & CMH_QSE_FLAG_DS_REF); + hw_rng =3D !!(req.flags & CMH_QSE_FLAG_HW_RNG); + + /* + * QSE keys only support PT storage -- the eSW sign path + * hardcodes SYS_TYPE_FLAG_PT when reading the key back. + * QSE SCA protection uses masking (CMH_QSE_FLAG_MASKED), + * not the 2-share mechanism (CMH_FLAG_SCA). + */ + key_flags =3D req.flags & CMH_FLAG_MASK; + if (key_flags && key_flags !=3D CMH_FLAG_PT) + return -EINVAL; + key_flags =3D CMH_FLAG_PT; + + if (masked && !ds_ref) + return -EINVAL; + + pk_size =3D ml_dsa_pk_size[mi]; + sk_size =3D masked ? ml_dsa_sk_size_masked[mi] : ml_dsa_sk_size[mi]; + seed_len =3D masked ? QSE_SEED_LEN_MASKED : QSE_SEED_LEN; + + if (hw_rng) + qse_flags |=3D QSE_FLAG_USE_RNG; + if (ds_ref) + qse_flags |=3D QSE_FLAG_USE_REF; + + /* + * Without HW RNG the caller must supply the seed; a 0 pointer with + * non-zero length is a wild eSW DMA read (F3), not "absent". + */ + if (!hw_rng && !req.seed) + return -EINVAL; + + pk_buf =3D kzalloc(pk_size, GFP_KERNEL); + if (!pk_buf) + return -ENOMEM; + + if (!hw_rng && req.seed) { + seed_buf =3D kmalloc(seed_len, GFP_KERNEL); + if (!seed_buf) { + ret =3D -ENOMEM; + goto out_free; + } + if (copy_from_user(seed_buf, u64_to_user_ptr(req.seed), + seed_len)) { + ret =3D -EFAULT; + goto out_free; + } + } + + if (ds_ref) { + ref_buf =3D kzalloc_obj(u64, GFP_KERNEL); + if (!ref_buf) { + ret =3D -ENOMEM; + goto out_free; + } + } else { + sk_buf =3D kzalloc(sk_size, GFP_KERNEL); + if (!sk_buf) { + ret =3D -ENOMEM; + goto out_free; + } + } + + pk_dma =3D cmh_dma_map_single(pk_buf, pk_size, DMA_FROM_DEVICE); + if (cmh_dma_map_error(pk_dma)) { + ret =3D -ENOMEM; + goto out_free; + } + + if (seed_buf) { + seed_dma =3D cmh_dma_map_single(seed_buf, seed_len, + DMA_TO_DEVICE); + if (cmh_dma_map_error(seed_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + } + + if (ds_ref) { + ref_dma =3D cmh_dma_map_single(ref_buf, sizeof(u64), + DMA_FROM_DEVICE); + if (cmh_dma_map_error(ref_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + } else { + sk_dma =3D cmh_dma_map_single(sk_buf, sk_size, DMA_FROM_DEVICE); + if (cmh_dma_map_error(sk_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + } + + idx =3D 0; + if (ds_ref) { + vcq_set_header(&vcq[0], QSE_VCQ_CMDS_MAX); + idx++; + vcq_add_sys_new(&vcq[idx++], req.sk_cid, ref_dma, sk_size); + vcq_add_qse_ml_dsa_keygen(&vcq[idx++], qse_cid, req.mode, qse_flags, + seed_dma, pk_dma, + SYS_REF_LAST, + SYS_TYPE_SET(key_flags, + CORE_ID_QSE), + masked); + vcq_add_qse_flush(&vcq[idx++], qse_cid); + ret =3D cmh_tm_submit_sync_mbx(vcq, QSE_VCQ_CMDS_MAX, + 1, MGMT_MBX); + } else { + vcq_set_header(&vcq[0], QSE_VCQ_CMDS_MIN); + idx++; + vcq_add_qse_ml_dsa_keygen(&vcq[idx++], qse_cid, req.mode, qse_flags, + seed_dma, pk_dma, + sk_dma, 0, masked); + vcq_add_qse_flush(&vcq[idx++], qse_cid); + ret =3D cmh_tm_submit_sync_mbx(vcq, QSE_VCQ_CMDS_MIN, + 1, MGMT_MBX); + } + +out_unmap: + if (ds_ref && !cmh_dma_map_error(ref_dma)) + cmh_dma_unmap_single(ref_dma, sizeof(u64), DMA_FROM_DEVICE); + if (!ds_ref && sk_buf && !cmh_dma_map_error(sk_dma)) + cmh_dma_unmap_single(sk_dma, sk_size, DMA_FROM_DEVICE); + if (seed_buf && !cmh_dma_map_error(seed_dma)) + cmh_dma_unmap_single(seed_dma, seed_len, DMA_TO_DEVICE); + if (!cmh_dma_map_error(pk_dma)) + cmh_dma_unmap_single(pk_dma, pk_size, DMA_FROM_DEVICE); + + /* Keygen failed after SYS_CMD_NEW (ds_ref mode); scrub the orphan. */ + if (ret && ds_ref && *ref_buf) + cmh_mgmt_ds_scrub(*ref_buf, 0); + + if (!ret) { + if (copy_to_user(u64_to_user_ptr(req.pk), pk_buf, pk_size)) { + ret =3D -EFAULT; + } else if (ds_ref) { + req.sk_ref =3D *ref_buf; + } else if (copy_to_user(u64_to_user_ptr(req.sk), + sk_buf, sk_size)) { + ret =3D -EFAULT; + } + if (!ret && copy_to_user(argp, &req, sizeof(req))) + ret =3D -EFAULT; + /* + * A copy_to_user faulted after the DS slot was created, so + * its ref never reached userspace and the caller cannot free + * it. Logically delete the orphaned slot: grant-with-no-access + * wipes the key material and clears the CID (it does not + * reclaim the stack space -- only a datastore reset does). + */ + if (ret && ds_ref) { + cmh_mgmt_ds_scrub(*ref_buf, 0); + dev_warn(cmh_dev(), + "mgmt: ML_DSA_KEYGEN copy_to_user failed, DS slot cleaned up\n"); + } + } + +out_free: + kfree_sensitive(sk_buf); + kfree(ref_buf); + kfree_sensitive(seed_buf); + kfree(pk_buf); + return ret; +} + +/** + * cmh_mgmt_ml_dsa_sign() - Handle CMH_MGMT_IOC_ML_DSA_SIGN ioctl + * @argp: User-space ioctl argument pointer + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_ml_dsa_sign(void __user *argp) +{ + u32 qse_cid =3D cmh_core_default_id(CMH_CORE_QSE); + + struct cmh_ioctl_ml_dsa_sign req; + struct vcq_cmd vcq[QSE_VCQ_CMDS_MIN]; + u32 sig_size, copy_len, rnd_len; + u32 qse_flags =3D 0; + bool masked; + u8 *m_buf, *sig_buf, *sk_buf =3D NULL, *rnd_buf =3D NULL; + dma_addr_t m_dma =3D DMA_MAPPING_ERROR, sig_dma =3D DMA_MAPPING_ERROR; + dma_addr_t sk_dma =3D DMA_MAPPING_ERROR; + /* + * 0 =3D "no added randomness": the eSW treats a zero rnd pointer as + * absent. A garbage DMA_MAPPING_ERROR would be taken as a real + * address and DMA'd from. + */ + dma_addr_t rnd_dma =3D 0; + u64 sk_ref; + int mi, ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + mi =3D ml_dsa_mode_idx(req.mode); + if (mi < 0) + return -EINVAL; + if (req.mlen > ML_DSA_MAX_MLEN && req.mlen !=3D ML_DSA_MLEN_EXTERNAL_MU) + return -EINVAL; + + masked =3D !!(req.flags & CMH_QSE_FLAG_MASKED); + rnd_len =3D masked ? QSE_SEED_LEN_MASKED : QSE_SEED_LEN; + sig_size =3D ml_dsa_sig_size[mi]; + copy_len =3D (req.mlen =3D=3D ML_DSA_MLEN_EXTERNAL_MU) + ? ML_DSA_EXTMU_LEN : req.mlen; + + /* + * sk: if DS_REF, req.sk is a DS reference (masked sk lives in DS). + * Otherwise, copy raw sk from user-space. + * Masked sign requires DS ref (polynomial unmasking not supported). + */ + if (req.flags & CMH_QSE_FLAG_DS_REF) { + sk_ref =3D req.sk; + qse_flags |=3D QSE_FLAG_USE_REF; + } else { + u32 sk_size; + + if (masked) + return -EINVAL; + sk_size =3D ml_dsa_sk_size[mi]; + sk_buf =3D kmalloc(sk_size, GFP_KERNEL); + if (!sk_buf) + return -ENOMEM; + if (copy_from_user(sk_buf, u64_to_user_ptr(req.sk), sk_size)) { + kfree_sensitive(sk_buf); + return -EFAULT; + } + } + + m_buf =3D kmalloc(max_t(u32, copy_len, 1), GFP_KERNEL); + sig_buf =3D kzalloc(sig_size, GFP_KERNEL); + if (!m_buf || !sig_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + if (copy_len > 0 && + copy_from_user(m_buf, u64_to_user_ptr(req.m), copy_len)) { + ret =3D -EFAULT; + goto out_free; + } + + if (req.rnd) { + rnd_buf =3D kmalloc(rnd_len, GFP_KERNEL); + if (!rnd_buf) { + ret =3D -ENOMEM; + goto out_free; + } + if (copy_from_user(rnd_buf, u64_to_user_ptr(req.rnd), + rnd_len)) { + ret =3D -EFAULT; + goto out_free; + } + } + + if (copy_len > 0) { + m_dma =3D cmh_dma_map_single(m_buf, copy_len, DMA_TO_DEVICE); + if (cmh_dma_map_error(m_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + } + sig_dma =3D cmh_dma_map_single(sig_buf, sig_size, DMA_FROM_DEVICE); + if (cmh_dma_map_error(sig_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + if (sk_buf) { + sk_dma =3D cmh_dma_map_single(sk_buf, ml_dsa_sk_size[mi], + DMA_TO_DEVICE); + if (cmh_dma_map_error(sk_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + sk_ref =3D sk_dma; + } + + if (rnd_buf) { + rnd_dma =3D cmh_dma_map_single(rnd_buf, rnd_len, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rnd_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + } + + vcq_set_header(&vcq[0], QSE_VCQ_CMDS_MIN); + vcq_add_qse_ml_dsa_sign(&vcq[1], qse_cid, req.mode, qse_flags, + rnd_dma, m_dma, sk_ref, sig_dma, + req.mlen, masked); + vcq_add_qse_flush(&vcq[2], qse_cid); + + ret =3D cmh_tm_submit_sync_mbx(vcq, QSE_VCQ_CMDS_MIN, 1, MGMT_MBX); + +out_unmap: + if (rnd_buf && !cmh_dma_map_error(rnd_dma)) + cmh_dma_unmap_single(rnd_dma, rnd_len, DMA_TO_DEVICE); + if (sk_buf && !cmh_dma_map_error(sk_dma)) + cmh_dma_unmap_single(sk_dma, ml_dsa_sk_size[mi], + DMA_TO_DEVICE); + if (!cmh_dma_map_error(sig_dma)) + cmh_dma_unmap_single(sig_dma, sig_size, DMA_FROM_DEVICE); + if (copy_len > 0 && !cmh_dma_map_error(m_dma)) + cmh_dma_unmap_single(m_dma, copy_len, DMA_TO_DEVICE); + + if (!ret) { + if (copy_to_user(u64_to_user_ptr(req.sig), sig_buf, sig_size)) + ret =3D -EFAULT; + } + +out_free: + kfree_sensitive(rnd_buf); + kfree(sig_buf); + kfree(m_buf); + kfree_sensitive(sk_buf); + return ret; +} + +/* -- PQC -- SLH-DSA -- */ + +/** + * cmh_mgmt_slhdsa_keygen() - Handle CMH_MGMT_IOC_SLHDSA_KEYGEN ioctl + * @argp: User-space ioctl argument pointer + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_slhdsa_keygen(void __user *argp) +{ + u32 hcq_cid =3D cmh_core_default_id(CMH_CORE_HCQ); + + struct cmh_ioctl_slhdsa_keygen req; + struct vcq_cmd vcq[HCQ_VCQ_CMDS_MAX]; + u32 pk_sz, sk_sz, seed_sz, sk_alloc, vcq_cnt, key_flags; + bool ds_ref; + u8 *seed_buf, *pk_buf, *sk_buf =3D NULL; + u64 *ref_buf =3D NULL; + dma_addr_t seed_dma, pk_dma, sk_dma =3D DMA_MAPPING_ERROR, ref_dma =3D DM= A_MAPPING_ERROR; + int ret, idx; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.__reserved) + return -EINVAL; + if (req.parameter_set < 1 || req.parameter_set > HCQ_SLHDSA_PARAM_MAX) + return -EINVAL; + if (req.flags & ~(CMH_QSE_FLAG_DS_REF | CMH_FLAG_MASK)) + return -EINVAL; + + ds_ref =3D !!(req.flags & CMH_QSE_FLAG_DS_REF); + + /* + * QSE keys only support PT storage -- the eSW sign path + * hardcodes SYS_TYPE_FLAG_PT when reading the key back. + * HCQ core sets key type internally during keygen. + */ + key_flags =3D req.flags & CMH_FLAG_MASK; + if (key_flags && key_flags !=3D CMH_FLAG_PT) + return -EINVAL; + (void)key_flags; + + pk_sz =3D slhdsa_pk_size(req.parameter_set); + sk_sz =3D slhdsa_sk_size(req.parameter_set); + seed_sz =3D slhdsa_seed_size(req.parameter_set); + + seed_buf =3D kmalloc(seed_sz, GFP_KERNEL); + pk_buf =3D kzalloc(pk_sz, GFP_KERNEL); + if (!seed_buf || !pk_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + if (copy_from_user(seed_buf, u64_to_user_ptr(req.seed), seed_sz)) { + ret =3D -EFAULT; + goto out_free; + } + + /* + * Both paths need ref_buf for sys_new output. Non-ds_ref also + * needs sk_buf (+16 for SYS header) to read back via sys_read. + */ + ref_buf =3D kzalloc(sizeof(u64), GFP_KERNEL); + if (!ref_buf) { + ret =3D -ENOMEM; + goto out_free; + } + if (!ds_ref) { + sk_alloc =3D sk_sz + SYS_WRAP_HDR_SIZE; + sk_buf =3D kzalloc(sk_alloc, GFP_KERNEL); + if (!sk_buf) { + ret =3D -ENOMEM; + goto out_free; + } + } + + seed_dma =3D cmh_dma_map_single(seed_buf, seed_sz, DMA_TO_DEVICE); + pk_dma =3D cmh_dma_map_single(pk_buf, pk_sz, DMA_FROM_DEVICE); + ref_dma =3D cmh_dma_map_single(ref_buf, sizeof(u64), DMA_FROM_DEVICE); + if (cmh_dma_map_error(seed_dma) || cmh_dma_map_error(pk_dma) || + cmh_dma_map_error(ref_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + if (!ds_ref) { + sk_dma =3D cmh_dma_map_single(sk_buf, sk_alloc, + DMA_FROM_DEVICE); + if (cmh_dma_map_error(sk_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + } + + /* + * SLH-DSA keygen requires seed and sk as DS references. + * VCQ: hdr + sys_new(sk) + sys_write(seed->TEMP) + keygen + [sys_read] += flush + */ + idx =3D 0; + if (ds_ref) { + vcq_cnt =3D HCQ_VCQ_CMDS_MAX - 1; /* hdr+new+write+keygen+flush */ + vcq_set_header(&vcq[idx++], vcq_cnt); + vcq_add_sys_new(&vcq[idx++], req.sk_cid, ref_dma, + sk_sz); + } else { + vcq_cnt =3D HCQ_VCQ_CMDS_MAX; /* hdr+new+write+keygen+read+flush */ + vcq_set_header(&vcq[idx++], vcq_cnt); + vcq_add_sys_new(&vcq[idx++], SYS_CID_NONE, ref_dma, + sk_sz); + } + vcq_add_sys_write(&vcq[idx++], SYS_REF_TEMP, seed_dma, 0, + seed_sz, + SYS_TYPE_SET(SYS_TYPE_FLAG_PT, CORE_ID_HCQ)); + vcq_add_hcq_slhdsa_keygen(&vcq[idx++], hcq_cid, req.parameter_set, + seed_sz, pk_sz, sk_sz, + SYS_REF_TEMP, pk_dma, SYS_REF_LAST); + if (!ds_ref) + vcq_add_sys_read(&vcq[idx++], SYS_REF_LAST, sk_dma, + 0, sk_sz + SYS_WRAP_HDR_SIZE); + vcq_add_hcq_flush(&vcq[idx++], hcq_cid); + + ret =3D cmh_tm_submit_sync_tmo(vcq, vcq_cnt, 1, MGMT_MBX, + cmh_tm_slow_op_timeout_jiffies()); + +out_unmap: + if (!ds_ref && sk_buf && !cmh_dma_map_error(sk_dma)) + cmh_dma_unmap_single(sk_dma, sk_alloc, DMA_FROM_DEVICE); + if (!cmh_dma_map_error(ref_dma)) + cmh_dma_unmap_single(ref_dma, sizeof(u64), DMA_FROM_DEVICE); + if (!cmh_dma_map_error(pk_dma)) + cmh_dma_unmap_single(pk_dma, pk_sz, DMA_FROM_DEVICE); + if (!cmh_dma_map_error(seed_dma)) + cmh_dma_unmap_single(seed_dma, seed_sz, DMA_TO_DEVICE); + + /* + * Raw-output mode exports the SK to host and keeps no datastore handle, + * so scrub the object once the readback VCQ has run: grant-with-no- + * access wipes the exported private key from datastore RAM and revokes + * its CID. The slot offset is not reclaimed (the datastore is a bump + * allocator; only a full reset frees space) -- acceptable on this + * CAP_SYS_ADMIN path, as for every other DS-allocating mgmt ioctl. + */ + if (!ds_ref && *ref_buf) + cmh_mgmt_ds_scrub(*ref_buf, 0); + + /* Keygen failed after SYS_CMD_NEW (ds_ref mode); scrub the orphan. */ + if (ret && ds_ref && *ref_buf) + cmh_mgmt_ds_scrub(*ref_buf, 0); + + if (!ret) { + if (copy_to_user(u64_to_user_ptr(req.pk), pk_buf, pk_sz)) { + ret =3D -EFAULT; + } else if (ds_ref) { + req.sk_ref =3D *ref_buf; + } else if (copy_to_user(u64_to_user_ptr(req.sk), + sk_buf + SYS_WRAP_HDR_SIZE, + sk_sz)) { + ret =3D -EFAULT; + } + if (!ret && copy_to_user(argp, &req, sizeof(req))) + ret =3D -EFAULT; + /* + * A copy_to_user faulted after the DS slot was created, so + * its ref never reached userspace and the caller cannot free + * it. Logically delete the orphaned slot: grant-with-no-access + * wipes the key material and clears the CID (it does not + * reclaim the stack space -- only a datastore reset does). + */ + if (ret && ds_ref) { + cmh_mgmt_ds_scrub(*ref_buf, 0); + dev_warn(cmh_dev(), + "mgmt: SLHDSA_KEYGEN copy_to_user failed, DS slot cleaned up\n"); + } + } + +out_free: + kfree_sensitive(sk_buf); + kfree(ref_buf); + kfree(pk_buf); + kfree_sensitive(seed_buf); + return ret; +} + +/** + * cmh_mgmt_slhdsa_sign() - Handle CMH_MGMT_IOC_SLHDSA_SIGN ioctl + * @argp: User-space ioctl argument pointer + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_slhdsa_sign(void __user *argp) +{ + u32 hcq_cid =3D cmh_core_default_id(CMH_CORE_HCQ); + + struct cmh_ioctl_slhdsa_sign req; + struct vcq_cmd vcq[HCQ_VCQ_CMDS_MIN]; + u32 sig_sz, n_val; + u8 *msg_buf, *ctx_buf =3D NULL, *sig_buf, *rnd_buf =3D NULL; + dma_addr_t msg_dma =3D DMA_MAPPING_ERROR, ctx_dma =3D DMA_MAPPING_ERROR; + /* + * 0 =3D "no added randomness": the eSW treats a zero rnd pointer as + * absent. A garbage DMA_MAPPING_ERROR would be taken as a real + * address and DMA'd from. + */ + dma_addr_t sig_dma =3D DMA_MAPPING_ERROR, rnd_dma =3D 0; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.parameter_set < 1 || req.parameter_set > HCQ_SLHDSA_PARAM_MAX) + return -EINVAL; + if (req.msg_len > SLHDSA_MAX_MSG_LEN) + return -EINVAL; + if (req.ctx_len > SLHDSA_MAX_CTX_LEN) + return -EINVAL; + + /* ctx_len > 0 with a 0 ctx pointer is a wild eSW DMA read (F3). */ + if (req.ctx_len > 0 && !req.ctx) + return -EINVAL; + + sig_sz =3D slhdsa_get_sig_size(req.parameter_set); + n_val =3D slhdsa_n[req.parameter_set - 1]; + + msg_buf =3D kmalloc(max_t(u32, req.msg_len, 1), GFP_KERNEL); + sig_buf =3D kzalloc(sig_sz, GFP_KERNEL); + if (!msg_buf || !sig_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + if (req.msg_len > 0 && + copy_from_user(msg_buf, u64_to_user_ptr(req.msg), req.msg_len)) { + ret =3D -EFAULT; + goto out_free; + } + + if (req.ctx_len > 0 && req.ctx) { + ctx_buf =3D kmalloc(req.ctx_len, GFP_KERNEL); + if (!ctx_buf) { + ret =3D -ENOMEM; + goto out_free; + } + if (copy_from_user(ctx_buf, u64_to_user_ptr(req.ctx), + req.ctx_len)) { + ret =3D -EFAULT; + goto out_free; + } + } + + if (req.add_random) { + rnd_buf =3D kmalloc(n_val, GFP_KERNEL); + if (!rnd_buf) { + ret =3D -ENOMEM; + goto out_free; + } + if (copy_from_user(rnd_buf, u64_to_user_ptr(req.add_random), + n_val)) { + ret =3D -EFAULT; + goto out_free; + } + } + + sig_dma =3D cmh_dma_map_single(sig_buf, sig_sz, DMA_FROM_DEVICE); + if (cmh_dma_map_error(sig_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + if (req.msg_len > 0) { + msg_dma =3D cmh_dma_map_single(msg_buf, req.msg_len, + DMA_TO_DEVICE); + if (cmh_dma_map_error(msg_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + } + + if (ctx_buf) { + ctx_dma =3D cmh_dma_map_single(ctx_buf, req.ctx_len, + DMA_TO_DEVICE); + if (cmh_dma_map_error(ctx_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + } + + if (rnd_buf) { + rnd_dma =3D cmh_dma_map_single(rnd_buf, n_val, DMA_TO_DEVICE); + if (cmh_dma_map_error(rnd_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + } + + vcq_set_header(&vcq[0], HCQ_VCQ_CMDS_MIN); + vcq_add_hcq_slhdsa_sign(&vcq[1], hcq_cid, req.parameter_set, + req.msg_len, req.ctx_len, + rnd_dma, msg_dma, ctx_dma, + req.sk, sig_dma); + vcq_add_hcq_flush(&vcq[2], hcq_cid); + + ret =3D cmh_tm_submit_sync_tmo(vcq, HCQ_VCQ_CMDS_MIN, 1, MGMT_MBX, + cmh_tm_slow_op_timeout_jiffies()); + +out_unmap: + if (rnd_buf && !cmh_dma_map_error(rnd_dma)) + cmh_dma_unmap_single(rnd_dma, n_val, DMA_TO_DEVICE); + if (ctx_buf && !cmh_dma_map_error(ctx_dma)) + cmh_dma_unmap_single(ctx_dma, req.ctx_len, DMA_TO_DEVICE); + if (!cmh_dma_map_error(sig_dma)) + cmh_dma_unmap_single(sig_dma, sig_sz, DMA_FROM_DEVICE); + if (req.msg_len > 0 && !cmh_dma_map_error(msg_dma)) + cmh_dma_unmap_single(msg_dma, req.msg_len, DMA_TO_DEVICE); + + if (!ret) { + if (copy_to_user(u64_to_user_ptr(req.sig), sig_buf, sig_sz)) + ret =3D -EFAULT; + } + +out_free: + kfree_sensitive(rnd_buf); + kfree(ctx_buf); + kfree(sig_buf); + kfree(msg_buf); + return ret; +} + +/* -- PQC -- SLH-DSA prehash -- */ + +/** + * cmh_mgmt_slhdsa_sign_prehash() - Handle CMH_MGMT_IOC_SLHDSA_SIGN_PREHAS= H ioctl + * @argp: User-space ioctl argument pointer + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_slhdsa_sign_prehash(void __user *argp) +{ + u32 hcq_cid =3D cmh_core_default_id(CMH_CORE_HCQ); + + struct cmh_ioctl_slhdsa_sign_prehash req; + struct vcq_cmd vcq[HCQ_VCQ_CMDS_MIN]; + u32 sig_sz, n_val, hcq_cmd; + u8 *msg_buf, *ctx_buf =3D NULL, *sig_buf, *rnd_buf =3D NULL; + dma_addr_t msg_dma =3D DMA_MAPPING_ERROR, ctx_dma =3D DMA_MAPPING_ERROR; + /* + * 0 =3D "no added randomness": the eSW treats a zero rnd pointer as + * absent. A garbage DMA_MAPPING_ERROR would be taken as a real + * address and DMA'd from. + */ + dma_addr_t sig_dma =3D DMA_MAPPING_ERROR, rnd_dma =3D 0; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.parameter_set < 1 || req.parameter_set > HCQ_SLHDSA_PARAM_MAX) + return -EINVAL; + if (req.prehash_algo < 1 || req.prehash_algo > HCQ_SLHDSA_PREHASH_SHAKE25= 6) + return -EINVAL; + if (req.msg_len > SLHDSA_MAX_MSG_LEN) + return -EINVAL; + if (req.ctx_len > SLHDSA_MAX_CTX_LEN) + return -EINVAL; + + /* ctx_len > 0 with a 0 ctx pointer is a wild eSW DMA read (F3). */ + if (req.ctx_len > 0 && !req.ctx) + return -EINVAL; + + hcq_cmd =3D req.digest ? HCQ_CMD_SLHDSA_SIGN_PREHASH_DIGEST + : HCQ_CMD_SLHDSA_SIGN_PREHASH; + + sig_sz =3D slhdsa_get_sig_size(req.parameter_set); + n_val =3D slhdsa_n[req.parameter_set - 1]; + + msg_buf =3D kmalloc(max_t(u32, req.msg_len, 1), GFP_KERNEL); + sig_buf =3D kzalloc(sig_sz, GFP_KERNEL); + if (!msg_buf || !sig_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + if (req.msg_len > 0 && + copy_from_user(msg_buf, u64_to_user_ptr(req.msg), req.msg_len)) { + ret =3D -EFAULT; + goto out_free; + } + + if (req.ctx_len > 0 && req.ctx) { + ctx_buf =3D kmalloc(req.ctx_len, GFP_KERNEL); + if (!ctx_buf) { + ret =3D -ENOMEM; + goto out_free; + } + if (copy_from_user(ctx_buf, u64_to_user_ptr(req.ctx), + req.ctx_len)) { + ret =3D -EFAULT; + goto out_free; + } + } + + if (req.add_random) { + rnd_buf =3D kmalloc(n_val, GFP_KERNEL); + if (!rnd_buf) { + ret =3D -ENOMEM; + goto out_free; + } + if (copy_from_user(rnd_buf, u64_to_user_ptr(req.add_random), + n_val)) { + ret =3D -EFAULT; + goto out_free; + } + } + + sig_dma =3D cmh_dma_map_single(sig_buf, sig_sz, DMA_FROM_DEVICE); + if (cmh_dma_map_error(sig_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + if (req.msg_len > 0) { + msg_dma =3D cmh_dma_map_single(msg_buf, req.msg_len, + DMA_TO_DEVICE); + if (cmh_dma_map_error(msg_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + } + + if (ctx_buf) { + ctx_dma =3D cmh_dma_map_single(ctx_buf, req.ctx_len, + DMA_TO_DEVICE); + if (cmh_dma_map_error(ctx_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + } + + if (rnd_buf) { + rnd_dma =3D cmh_dma_map_single(rnd_buf, n_val, DMA_TO_DEVICE); + if (cmh_dma_map_error(rnd_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + } + + vcq_set_header(&vcq[0], HCQ_VCQ_CMDS_MIN); + vcq_add_hcq_slhdsa_sign_prehash(&vcq[1], hcq_cid, + hcq_cmd, req.parameter_set, + req.prehash_algo, + req.msg_len, req.ctx_len, + rnd_dma, msg_dma, ctx_dma, + req.sk, sig_dma); + vcq_add_hcq_flush(&vcq[2], hcq_cid); + + ret =3D cmh_tm_submit_sync_tmo(vcq, HCQ_VCQ_CMDS_MIN, 1, MGMT_MBX, + cmh_tm_slow_op_timeout_jiffies()); + +out_unmap: + if (rnd_buf && !cmh_dma_map_error(rnd_dma)) + cmh_dma_unmap_single(rnd_dma, n_val, DMA_TO_DEVICE); + if (ctx_buf && !cmh_dma_map_error(ctx_dma)) + cmh_dma_unmap_single(ctx_dma, req.ctx_len, DMA_TO_DEVICE); + if (!cmh_dma_map_error(sig_dma)) + cmh_dma_unmap_single(sig_dma, sig_sz, DMA_FROM_DEVICE); + if (req.msg_len > 0 && !cmh_dma_map_error(msg_dma)) + cmh_dma_unmap_single(msg_dma, req.msg_len, DMA_TO_DEVICE); + + if (!ret) { + if (copy_to_user(u64_to_user_ptr(req.sig), sig_buf, sig_sz)) + ret =3D -EFAULT; + } + +out_free: + kfree_sensitive(rnd_buf); + kfree(ctx_buf); + kfree(sig_buf); + kfree(msg_buf); + return ret; +} + +/* -- EAC (Error and Alarm Controller) ---- */ + diff --git a/drivers/crypto/cmh/cmh_pke_sm2.c b/drivers/crypto/cmh/cmh_pke_= sm2.c new file mode 100644 index 000000000000..66d818636406 --- /dev/null +++ b/drivers/crypto/cmh/cmh_pke_sm2.c @@ -0,0 +1,862 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- SM2 PKE Ioctl Handlers + * + * SM2 (GM/T 0003) is the Chinese national public-key standard over the + * sm2p256v1 curve (256-bit). It defines three protocols: + * + * - Signature: reuses ECDSA sign/verify with SM2_CURVE (0x18), handled + * by the existing cmh_mgmt_pke_ecdsa_{sign,verify}() paths. + * - Encryption: two-step (ENC_POINT + ENC_HASH / DEC_POINT + DEC_HASH). + * - Key Exchange: four-step (ECDH_KEYGEN + ID_DIGEST + ECDH + ECDH_HASH= ). + * + * This file implements the 8 SM2-specific ioctl handlers (0x16--0x1D). + * Sign/verify/keygen/pubgen use the existing ECDSA/EC paths unchanged. + * + * VCQ flag convention (from eSW API): + * - Most SM2 commands use flags=3D0 (no swap). + * - SM2_DEC_POINT and SM2_ECDH_HASH use PKE_SWAP_FLAGS on the + * PKE command itself. + * - SM2_ECDH and SM2_ECDH_HASH also apply PKE_SWAP_FLAGS on + * their sys_new/sys_data VCQ phases: the swap byte-reverses the + * DS object *payload* into Weierstrass format, not the ref-id + * writeback, so the datastore ref is read back intact (covered by + * the SM2 KAT). + */ + +#include +#include + +#include "cmh_pke.h" +#include "cmh_pke_sm2.h" +#include "cmh_sys.h" +#include "cmh_dma.h" +#include "cmh_txn.h" +#include "cmh_mgmt.h" +#include "cmh_sys_abi.h" +#include + +/* SM2 fixed sizes (sm2p256v1: 256-bit curve) */ +#define SM2_CLEN 32U /* coordinate length */ +#define SM2_POINT_LEN 64U /* uncompressed EC point (x||y) */ +#define SM2_SHARED_KEY_LEN 16U /* ECDH shared key output */ +#define SM2_DIGEST_LEN 32U /* SM3 ZA digest */ +#define SM2_NONCE_LEN 32U /* nonce (when caller-provided) */ +/* + * SM2 enc_hash/dec_hash payload limit. + * + * The eSW PKE driver expands the GM/T 0003.4 KDF by issuing a single SM3 + * invocation per command (one 32-byte block of key stream). Messages + * longer than 32 bytes would require ceil(msg_len / 32) SM3 invocations + * with an incremented counter, which the eSW does not perform; longer + * inputs would silently produce incorrect ciphertext / plaintext. + * + * The eSW PKE SRAM can physically hold up to 4000 bytes of payload, but + * that capacity is unusable until a future eSW change implements the full + * KDF expansion. Until then we cap the LKM at the 32-byte limit + * documented in Documentation/ABI/testing/cmh-mgmt. + */ +#define SM2_MAX_MSG_LEN 32U /* max plaintext for encrypt/decrypt */ +#define SM2_MAX_ID_LEN 32U /* max identity string */ +#define SM2_CT_OVERHEAD 96U /* C1(64) + C3(32) */ +#define SM2_MAX_CT_LEN (SM2_CT_OVERHEAD + SM2_MAX_MSG_LEN) /* 128 */ + +/* -- SM2_ECDH_KEYGEN ------------------- */ + +/** + * cmh_mgmt_sm2_ecdh_keygen() - Handle CMH_MGMT_IOC_SM2_ECDH_KEYGEN ioctl + * @argp: User-space ioctl argument pointer + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_sm2_ecdh_keygen(void __user *argp) +{ + struct cmh_ioctl_sm2_ecdh_keygen req; + struct vcq_cmd vcq[PKE_VCQ_CMDS_MIN]; + u32 core_id =3D cmh_core_default_id(CMH_CORE_PKE); + u8 *nonce_buf, *sk_buf; + dma_addr_t nonce_dma, sk_dma; + int nonce_dir; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.nonce_len !=3D 0 && req.nonce_len !=3D SM2_NONCE_LEN) + return -EINVAL; + + sk_buf =3D kzalloc(SM2_POINT_LEN, GFP_KERNEL); + nonce_buf =3D kzalloc(SM2_NONCE_LEN, GFP_KERNEL); + if (!sk_buf || !nonce_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + /* + * nonce_len=3D32: caller provides ephemeral scalar r (DMA_TO_DEVICE). + * nonce_len=3D0: HW generates r and writes it back (DMA_FROM_DEVICE). + * The caller MUST supply a valid nonce pointer in both cases. + */ + if (req.nonce_len) { + if (copy_from_user(nonce_buf, u64_to_user_ptr(req.nonce), + SM2_NONCE_LEN)) { + ret =3D -EFAULT; + goto out_free; + } + nonce_dir =3D DMA_TO_DEVICE; + } else { + nonce_dir =3D DMA_FROM_DEVICE; + } + + sk_dma =3D cmh_dma_map_single(sk_buf, SM2_POINT_LEN, DMA_FROM_DEVICE); + nonce_dma =3D cmh_dma_map_single(nonce_buf, SM2_NONCE_LEN, nonce_dir); + if (cmh_dma_map_error(sk_dma) || cmh_dma_map_error(nonce_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + vcq_set_header(&vcq[0], PKE_VCQ_CMDS_MIN); + vcq_add_pke_sm2_ecdh_keygen(&vcq[1], core_id, nonce_dma, sk_dma, + req.nonce_len, 0); + vcq_add_pke_flush(&vcq[2], core_id); + + ret =3D cmh_tm_submit_sync_mbx(vcq, PKE_VCQ_CMDS_MIN, 1, MGMT_MBX); + +out_unmap: + if (!cmh_dma_map_error(nonce_dma)) + cmh_dma_unmap_single(nonce_dma, SM2_NONCE_LEN, nonce_dir); + if (!cmh_dma_map_error(sk_dma)) + cmh_dma_unmap_single(sk_dma, SM2_POINT_LEN, DMA_FROM_DEVICE); + + if (!ret) { + if (copy_to_user(u64_to_user_ptr(req.session_key), + sk_buf, SM2_POINT_LEN)) + ret =3D -EFAULT; + /* Write back HW-generated nonce when nonce_len=3D0 */ + if (!ret && !req.nonce_len) { + if (copy_to_user(u64_to_user_ptr(req.nonce), + nonce_buf, SM2_NONCE_LEN)) + ret =3D -EFAULT; + } + } + +out_free: + kfree_sensitive(nonce_buf); + kfree_sensitive(sk_buf); + return ret; +} + +/* -- SM2_ECDH -------------------------- */ + +/** + * cmh_mgmt_sm2_ecdh() - Handle CMH_MGMT_IOC_SM2_ECDH ioctl + * @argp: User-space ioctl argument pointer + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_sm2_ecdh(void __user *argp) +{ + struct cmh_ioctl_sm2_ecdh req; + /* keep_ds: hdr+sys_new+sm2_ecdh+pke_flush; readback adds sys_data+flush = */ + struct vcq_cmd vcq[5]; + u32 sp_type, core_id; + u8 *nonce_buf, *peer_pk_buf, *peer_sk_buf, *sp_buf; + u64 *ref_buf; + dma_addr_t nonce_dma, peer_pk_dma, peer_sk_dma, sp_dma, ref_dma; + int nonce_dir, ret, idx; + bool keep_ds; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.nonce_len !=3D 0 && req.nonce_len !=3D SM2_NONCE_LEN) + return -EINVAL; + + keep_ds =3D (req.shared_point_ref !=3D 0); + sp_type =3D SYS_TYPE_SET(SYS_TYPE_FLAG_PT, CORE_ID_PKE); + core_id =3D cmh_core_default_id(CMH_CORE_PKE); + + peer_pk_buf =3D kmalloc(SM2_POINT_LEN, GFP_KERNEL); + peer_sk_buf =3D kmalloc(SM2_POINT_LEN, GFP_KERNEL); + sp_buf =3D kzalloc(SM2_POINT_LEN, GFP_KERNEL); + ref_buf =3D kzalloc_obj(u64, GFP_KERNEL); + nonce_buf =3D kzalloc(SM2_NONCE_LEN, GFP_KERNEL); + if (!peer_pk_buf || !peer_sk_buf || !sp_buf || !ref_buf || + !nonce_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + if (copy_from_user(peer_pk_buf, u64_to_user_ptr(req.peer_public_key), + SM2_POINT_LEN) || + copy_from_user(peer_sk_buf, u64_to_user_ptr(req.peer_session_key), + SM2_POINT_LEN)) { + ret =3D -EFAULT; + goto out_free; + } + + if (req.nonce_len) { + if (copy_from_user(nonce_buf, u64_to_user_ptr(req.nonce), + SM2_NONCE_LEN)) { + ret =3D -EFAULT; + goto out_free; + } + nonce_dir =3D DMA_TO_DEVICE; + } else { + nonce_dir =3D DMA_FROM_DEVICE; + } + + peer_pk_dma =3D cmh_dma_map_single(peer_pk_buf, SM2_POINT_LEN, + DMA_TO_DEVICE); + peer_sk_dma =3D cmh_dma_map_single(peer_sk_buf, SM2_POINT_LEN, + DMA_TO_DEVICE); + sp_dma =3D cmh_dma_map_single(sp_buf, SM2_POINT_LEN, DMA_FROM_DEVICE); + ref_dma =3D cmh_dma_map_single(ref_buf, sizeof(u64), DMA_FROM_DEVICE); + nonce_dma =3D cmh_dma_map_single(nonce_buf, SM2_NONCE_LEN, nonce_dir); + + if (cmh_dma_map_error(peer_pk_dma) || cmh_dma_map_error(peer_sk_dma) || + cmh_dma_map_error(sp_dma) || cmh_dma_map_error(ref_dma) || + cmh_dma_map_error(nonce_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + if (keep_ds) { + /* + * keep_ds: retain the shared point as a persistent DS object + * (sys_new) whose ref is handed back for SM2_ECDH_HASH to + * consume. The caller owns and later releases it. + */ + idx =3D 0; + vcq_set_header(&vcq[idx++], 4); + vcq_add_sys_new(&vcq[idx], 0, ref_dma, SM2_POINT_LEN); + vcq[idx++].id |=3D PKE_SWAP_FLAGS; + vcq_add_pke_sm2_ecdh(&vcq[idx++], core_id, req.nonce_len, + SM2_CLEN, nonce_dma, peer_pk_dma, + peer_sk_dma, req.key_ref, SYS_REF_LAST, + sp_type, 0); + vcq_add_pke_flush(&vcq[idx++], core_id); + + ret =3D cmh_tm_submit_sync_mbx(vcq, 4, 1, MGMT_MBX); + if (ret) + goto out_unmap; + + /* Sync bounce buffer so CPU sees the DMA-written ref */ + cmh_dma_sync_for_cpu(ref_dma, sizeof(u64), DMA_FROM_DEVICE); + } else { + /* + * Read-back mode: compute into SYS_REF_TEMP and pop it in the + * same VCQ. Bounding the temp push (sm2_ecdh) and pop + * (sys_data) to one submission keeps them within a single + * mailbox occupation -- no concurrent mgmt op can interleave a + * temp -- and the read reclaims the slot, so nothing leaks. + */ + idx =3D 0; + vcq_set_header(&vcq[idx++], 5); + vcq_add_pke_sm2_ecdh(&vcq[idx++], core_id, req.nonce_len, + SM2_CLEN, nonce_dma, peer_pk_dma, + peer_sk_dma, req.key_ref, SYS_REF_TEMP, + sp_type, 0); + vcq_add_pke_flush(&vcq[idx++], core_id); + vcq_add_sys_data(&vcq[idx], SYS_REF_TEMP, sp_dma, + SM2_POINT_LEN); + vcq[idx++].id |=3D PKE_SWAP_FLAGS; + vcq_add_sys_flush(&vcq[idx++]); + + ret =3D cmh_tm_submit_sync_mbx(vcq, 5, 1, MGMT_MBX); + } + +out_unmap: + if (!cmh_dma_map_error(nonce_dma)) + cmh_dma_unmap_single(nonce_dma, SM2_NONCE_LEN, nonce_dir); + if (!cmh_dma_map_error(ref_dma)) + cmh_dma_unmap_single(ref_dma, sizeof(u64), DMA_FROM_DEVICE); + if (!cmh_dma_map_error(sp_dma)) + cmh_dma_unmap_single(sp_dma, SM2_POINT_LEN, DMA_FROM_DEVICE); + if (!cmh_dma_map_error(peer_sk_dma)) + cmh_dma_unmap_single(peer_sk_dma, SM2_POINT_LEN, + DMA_TO_DEVICE); + if (!cmh_dma_map_error(peer_pk_dma)) + cmh_dma_unmap_single(peer_pk_dma, SM2_POINT_LEN, + DMA_TO_DEVICE); + + /* + * keep_ds sys_new may have created the DS slot before the VCQ + * failed; scrub it so a partial failure does not leak a slot. + */ + if (ret && keep_ds && *ref_buf) + cmh_mgmt_ds_scrub(*ref_buf, 0); + + if (!ret) { + if (!keep_ds) { + if (copy_to_user(u64_to_user_ptr(req.shared_point), + sp_buf, SM2_POINT_LEN)) + ret =3D -EFAULT; + } else { + /* Return DS ref for ECDH_HASH to consume */ + u64 __user *sp_refp =3D (__u64 __user *) + u64_to_user_ptr(req.shared_point_ref); + + if (put_user(*ref_buf, sp_refp)) { + /* + * Failed to deliver the DS ref to + * userspace. Logically delete the + * orphaned slot so it does not leak. + */ + vcq_set_header(&vcq[0], 3); + vcq_add_sys_grant(&vcq[1], *ref_buf, + 0, 0, 0); + vcq_add_sys_flush(&vcq[2]); + cmh_tm_submit_sync_mbx(vcq, 3, 1, + MGMT_MBX); + dev_warn(cmh_dev(), "SM2 ECDH put_user failed, DS slot cleaned up\n"); + ret =3D -EFAULT; + } + } + /* Write back HW-generated nonce when nonce_len=3D0 */ + if (!ret && !req.nonce_len) { + if (copy_to_user(u64_to_user_ptr(req.nonce), + nonce_buf, SM2_NONCE_LEN)) + ret =3D -EFAULT; + } + } + +out_free: + kfree_sensitive(nonce_buf); + kfree(ref_buf); + kfree_sensitive(sp_buf); + kfree(peer_sk_buf); + kfree(peer_pk_buf); + return ret; +} + +/* -- SM2_DEC_POINT --------------------- */ + +/** + * cmh_mgmt_sm2_dec_point() - Handle CMH_MGMT_IOC_SM2_DEC_POINT ioctl + * @argp: User-space ioctl argument pointer + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_sm2_dec_point(void __user *argp) +{ + struct cmh_ioctl_sm2_dec_point req; + struct vcq_cmd vcq[PKE_VCQ_CMDS_MIN]; + u32 core_id =3D cmh_core_default_id(CMH_CORE_PKE); + u8 *ct_buf, *dp_buf; + dma_addr_t ct_dma, dp_dma; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.ciphertext_len <=3D SM2_CT_OVERHEAD || + req.ciphertext_len > SM2_MAX_CT_LEN) + return -EINVAL; + + /* Only need C1 (first 64 bytes) for the sidecar */ + ct_buf =3D kmalloc(SM2_POINT_LEN, GFP_KERNEL); + dp_buf =3D kzalloc(SM2_POINT_LEN, GFP_KERNEL); + if (!ct_buf || !dp_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + if (copy_from_user(ct_buf, u64_to_user_ptr(req.ciphertext), + SM2_POINT_LEN)) { + ret =3D -EFAULT; + goto out_free; + } + + ct_dma =3D cmh_dma_map_single(ct_buf, SM2_POINT_LEN, DMA_TO_DEVICE); + dp_dma =3D cmh_dma_map_single(dp_buf, SM2_POINT_LEN, DMA_FROM_DEVICE); + if (cmh_dma_map_error(ct_dma) || cmh_dma_map_error(dp_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + vcq_set_header(&vcq[0], PKE_VCQ_CMDS_MIN); + vcq_add_pke_sm2_dec_point(&vcq[1], core_id, req.ciphertext_len, SM2_CLEN, + ct_dma, dp_dma, req.key_ref, + PKE_SWAP_FLAGS); + vcq_add_pke_flush(&vcq[2], core_id); + + ret =3D cmh_tm_submit_sync_mbx(vcq, PKE_VCQ_CMDS_MIN, 1, MGMT_MBX); + +out_unmap: + if (!cmh_dma_map_error(dp_dma)) + cmh_dma_unmap_single(dp_dma, SM2_POINT_LEN, DMA_FROM_DEVICE); + if (!cmh_dma_map_error(ct_dma)) + cmh_dma_unmap_single(ct_dma, SM2_POINT_LEN, DMA_TO_DEVICE); + + if (!ret) { + if (copy_to_user(u64_to_user_ptr(req.dec_point), + dp_buf, SM2_POINT_LEN)) + ret =3D -EFAULT; + } + +out_free: + kfree_sensitive(dp_buf); + kfree(ct_buf); + return ret; +} + +/* -- SM2_ENC_POINT --------------------- */ + +/** + * cmh_mgmt_sm2_enc_point() - Handle CMH_MGMT_IOC_SM2_ENC_POINT ioctl + * @argp: User-space ioctl argument pointer + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_sm2_enc_point(void __user *argp) +{ + struct cmh_ioctl_sm2_enc_point req; + struct vcq_cmd vcq[PKE_VCQ_CMDS_MIN]; + u32 core_id =3D cmh_core_default_id(CMH_CORE_PKE); + u8 *nonce_buf =3D NULL, *pk_buf, *ct_buf, *ep_buf; + dma_addr_t nonce_dma =3D DMA_MAPPING_ERROR, pk_dma, ct_dma, ep_dma; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.nonce_len !=3D 0 && req.nonce_len !=3D SM2_NONCE_LEN) + return -EINVAL; + + pk_buf =3D kmalloc(SM2_POINT_LEN, GFP_KERNEL); + ct_buf =3D kzalloc(SM2_POINT_LEN, GFP_KERNEL); + ep_buf =3D kzalloc(SM2_POINT_LEN, GFP_KERNEL); + if (!pk_buf || !ct_buf || !ep_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + if (copy_from_user(pk_buf, u64_to_user_ptr(req.public_key), + SM2_POINT_LEN)) { + ret =3D -EFAULT; + goto out_free; + } + + if (req.nonce_len) { + nonce_buf =3D kmalloc(SM2_NONCE_LEN, GFP_KERNEL); + if (!nonce_buf) { + ret =3D -ENOMEM; + goto out_free; + } + if (copy_from_user(nonce_buf, u64_to_user_ptr(req.nonce), + SM2_NONCE_LEN)) { + ret =3D -EFAULT; + goto out_free; + } + } + + pk_dma =3D cmh_dma_map_single(pk_buf, SM2_POINT_LEN, DMA_TO_DEVICE); + ct_dma =3D cmh_dma_map_single(ct_buf, SM2_POINT_LEN, DMA_FROM_DEVICE); + ep_dma =3D cmh_dma_map_single(ep_buf, SM2_POINT_LEN, DMA_FROM_DEVICE); + if (nonce_buf) + nonce_dma =3D cmh_dma_map_single(nonce_buf, SM2_NONCE_LEN, + DMA_TO_DEVICE); + if (cmh_dma_map_error(pk_dma) || cmh_dma_map_error(ct_dma) || + cmh_dma_map_error(ep_dma) || + (nonce_buf && cmh_dma_map_error(nonce_dma))) { + ret =3D -ENOMEM; + goto out_unmap; + } + + vcq_set_header(&vcq[0], PKE_VCQ_CMDS_MIN); + vcq_add_pke_sm2_enc_point(&vcq[1], core_id, nonce_dma, pk_dma, ct_dma, + ep_dma, req.nonce_len, 0); + vcq_add_pke_flush(&vcq[2], core_id); + + ret =3D cmh_tm_submit_sync_mbx(vcq, PKE_VCQ_CMDS_MIN, 1, MGMT_MBX); + +out_unmap: + if (nonce_buf && !cmh_dma_map_error(nonce_dma)) + cmh_dma_unmap_single(nonce_dma, SM2_NONCE_LEN, DMA_TO_DEVICE); + if (!cmh_dma_map_error(ep_dma)) + cmh_dma_unmap_single(ep_dma, SM2_POINT_LEN, DMA_FROM_DEVICE); + if (!cmh_dma_map_error(ct_dma)) + cmh_dma_unmap_single(ct_dma, SM2_POINT_LEN, DMA_FROM_DEVICE); + if (!cmh_dma_map_error(pk_dma)) + cmh_dma_unmap_single(pk_dma, SM2_POINT_LEN, DMA_TO_DEVICE); + + if (!ret) { + if (copy_to_user(u64_to_user_ptr(req.ciphertext), + ct_buf, SM2_POINT_LEN) || + copy_to_user(u64_to_user_ptr(req.enc_point), + ep_buf, SM2_POINT_LEN)) + ret =3D -EFAULT; + } + +out_free: + kfree_sensitive(nonce_buf); + kfree(ep_buf); + kfree(ct_buf); + kfree(pk_buf); + return ret; +} + +/* -- SM2_ID_DIGEST --------------------- */ + +/** + * cmh_mgmt_sm2_id_digest() - Handle CMH_MGMT_IOC_SM2_ID_DIGEST ioctl + * @argp: User-space ioctl argument pointer + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_sm2_id_digest(void __user *argp) +{ + struct cmh_ioctl_sm2_id_digest req; + struct vcq_cmd vcq[PKE_VCQ_CMDS_MIN]; + u32 core_id =3D cmh_core_default_id(CMH_CORE_PKE); + u8 *id_buf, *pk_buf, *dig_buf; + dma_addr_t id_dma, pk_dma, dig_dma; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (!req.id_len || req.id_len > SM2_MAX_ID_LEN) + return -EINVAL; + + id_buf =3D kmalloc(req.id_len, GFP_KERNEL); + pk_buf =3D kmalloc(SM2_POINT_LEN, GFP_KERNEL); + dig_buf =3D kzalloc(SM2_DIGEST_LEN, GFP_KERNEL); + if (!id_buf || !pk_buf || !dig_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + if (copy_from_user(id_buf, u64_to_user_ptr(req.id), req.id_len) || + copy_from_user(pk_buf, u64_to_user_ptr(req.public_key), + SM2_POINT_LEN)) { + ret =3D -EFAULT; + goto out_free; + } + + id_dma =3D cmh_dma_map_single(id_buf, req.id_len, DMA_TO_DEVICE); + pk_dma =3D cmh_dma_map_single(pk_buf, SM2_POINT_LEN, DMA_TO_DEVICE); + dig_dma =3D cmh_dma_map_single(dig_buf, SM2_DIGEST_LEN, + DMA_FROM_DEVICE); + if (cmh_dma_map_error(id_dma) || cmh_dma_map_error(pk_dma) || + cmh_dma_map_error(dig_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + vcq_set_header(&vcq[0], PKE_VCQ_CMDS_MIN); + vcq_add_pke_sm2_id_digest(&vcq[1], core_id, id_dma, pk_dma, dig_dma, + req.id_len, 0); + vcq_add_pke_flush(&vcq[2], core_id); + + ret =3D cmh_tm_submit_sync_mbx(vcq, PKE_VCQ_CMDS_MIN, 1, MGMT_MBX); + +out_unmap: + if (!cmh_dma_map_error(dig_dma)) + cmh_dma_unmap_single(dig_dma, SM2_DIGEST_LEN, + DMA_FROM_DEVICE); + if (!cmh_dma_map_error(pk_dma)) + cmh_dma_unmap_single(pk_dma, SM2_POINT_LEN, DMA_TO_DEVICE); + if (!cmh_dma_map_error(id_dma)) + cmh_dma_unmap_single(id_dma, req.id_len, DMA_TO_DEVICE); + + if (!ret) { + if (copy_to_user(u64_to_user_ptr(req.digest), + dig_buf, SM2_DIGEST_LEN)) + ret =3D -EFAULT; + } + +out_free: + kfree(dig_buf); + kfree(pk_buf); + kfree(id_buf); + return ret; +} + +/* -- SM2_ECDH_HASH --------------------- */ + +/** + * cmh_mgmt_sm2_ecdh_hash() - Handle CMH_MGMT_IOC_SM2_ECDH_HASH ioctl + * @argp: User-space ioctl argument pointer + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_sm2_ecdh_hash(void __user *argp) +{ + struct cmh_ioctl_sm2_ecdh_hash req; + /* Phase 1: hdr + sys_new + sm2_ecdh_hash + pke_flush; reused for Phase 2= */ + struct vcq_cmd vcq[4]; + u32 sk_type, core_id; + u8 *peer_dig_buf, *dig_buf, *sk_buf; + u64 *ref_buf; + dma_addr_t peer_dig_dma, dig_dma, sk_dma, ref_dma; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.__reserved) + return -EINVAL; + + sk_type =3D SYS_TYPE_SET(SYS_TYPE_FLAG_PT, CORE_ID_PKE); + core_id =3D cmh_core_default_id(CMH_CORE_PKE); + + peer_dig_buf =3D kmalloc(SM2_DIGEST_LEN, GFP_KERNEL); + dig_buf =3D kmalloc(SM2_DIGEST_LEN, GFP_KERNEL); + sk_buf =3D kzalloc(SM2_SHARED_KEY_LEN, GFP_KERNEL); + ref_buf =3D kzalloc_obj(u64, GFP_KERNEL); + if (!peer_dig_buf || !dig_buf || !sk_buf || !ref_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + if (copy_from_user(peer_dig_buf, u64_to_user_ptr(req.peer_id_digest), + SM2_DIGEST_LEN) || + copy_from_user(dig_buf, u64_to_user_ptr(req.id_digest), + SM2_DIGEST_LEN)) { + ret =3D -EFAULT; + goto out_free; + } + + peer_dig_dma =3D cmh_dma_map_single(peer_dig_buf, SM2_DIGEST_LEN, + DMA_TO_DEVICE); + dig_dma =3D cmh_dma_map_single(dig_buf, SM2_DIGEST_LEN, DMA_TO_DEVICE); + sk_dma =3D cmh_dma_map_single(sk_buf, SM2_SHARED_KEY_LEN, + DMA_FROM_DEVICE); + ref_dma =3D cmh_dma_map_single(ref_buf, sizeof(u64), DMA_FROM_DEVICE); + if (cmh_dma_map_error(peer_dig_dma) || cmh_dma_map_error(dig_dma) || + cmh_dma_map_error(sk_dma) || cmh_dma_map_error(ref_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + /* + * Phase 1: sys_new(shared_key_ref) + SM2_ECDH_HASH + * The shared_point_ref from the ECDH step is passed directly + * as a DS reference -- the eSW hub reads it from DS. + */ + vcq_set_header(&vcq[0], 4); + vcq_add_sys_new(&vcq[1], 0, ref_dma, SM2_SHARED_KEY_LEN); + vcq[1].id |=3D PKE_SWAP_FLAGS; + vcq_add_pke_sm2_ecdh_hash(&vcq[2], core_id, peer_dig_dma, dig_dma, + req.shared_point_ref, SYS_REF_LAST, + sk_type, PKE_SWAP_FLAGS); + vcq_add_pke_flush(&vcq[3], core_id); + + ret =3D cmh_tm_submit_sync_mbx(vcq, 4, 1, MGMT_MBX); + if (ret) + goto out_unmap; + + /* Sync bounce buffer so CPU sees the DMA-written ref */ + cmh_dma_sync_for_cpu(ref_dma, sizeof(u64), DMA_FROM_DEVICE); + + /* Phase 2: read shared key from DS -> DMA */ + vcq_set_header(&vcq[0], 3); + vcq_add_sys_data(&vcq[1], *ref_buf, sk_dma, SM2_SHARED_KEY_LEN); + vcq_add_sys_flush(&vcq[2]); + + ret =3D cmh_tm_submit_sync_mbx(vcq, 3, 1, MGMT_MBX); + + /* + * Scrub the transient shared-key DS object created in phase 1 -- it + * is internal to this ioctl and never handed to the caller, so + * leaving it would leak a finite datastore slot on every call. + */ + cmh_mgmt_ds_scrub(*ref_buf, 0); + +out_unmap: + if (!cmh_dma_map_error(ref_dma)) + cmh_dma_unmap_single(ref_dma, sizeof(u64), DMA_FROM_DEVICE); + if (!cmh_dma_map_error(sk_dma)) + cmh_dma_unmap_single(sk_dma, SM2_SHARED_KEY_LEN, + DMA_FROM_DEVICE); + if (!cmh_dma_map_error(dig_dma)) + cmh_dma_unmap_single(dig_dma, SM2_DIGEST_LEN, DMA_TO_DEVICE); + if (!cmh_dma_map_error(peer_dig_dma)) + cmh_dma_unmap_single(peer_dig_dma, SM2_DIGEST_LEN, + DMA_TO_DEVICE); + + if (!ret) { + if (copy_to_user(u64_to_user_ptr(req.shared_key), + sk_buf, SM2_SHARED_KEY_LEN)) + ret =3D -EFAULT; + } + +out_free: + kfree(ref_buf); + kfree_sensitive(sk_buf); + kfree(dig_buf); + kfree(peer_dig_buf); + return ret; +} + +/* -- SM2_DEC_HASH ---------------------- */ + +/** + * cmh_mgmt_sm2_dec_hash() - Handle CMH_MGMT_IOC_SM2_DEC_HASH ioctl + * @argp: User-space ioctl argument pointer + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_sm2_dec_hash(void __user *argp) +{ + struct cmh_ioctl_sm2_dec_hash req; + struct vcq_cmd vcq[PKE_VCQ_CMDS_MIN]; + u32 msg_len, core_id; + u8 *ct_buf, *dp_buf, *pt_buf; + dma_addr_t ct_dma, dp_dma, pt_dma; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (req.ciphertext_len <=3D SM2_CT_OVERHEAD || + req.ciphertext_len > SM2_MAX_CT_LEN) + return -EINVAL; + + msg_len =3D req.ciphertext_len - SM2_CT_OVERHEAD; + core_id =3D cmh_core_default_id(CMH_CORE_PKE); + + ct_buf =3D kmalloc(req.ciphertext_len, GFP_KERNEL); + dp_buf =3D kmalloc(SM2_POINT_LEN, GFP_KERNEL); + pt_buf =3D kzalloc(msg_len, GFP_KERNEL); + if (!ct_buf || !dp_buf || !pt_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + if (copy_from_user(ct_buf, u64_to_user_ptr(req.ciphertext), + req.ciphertext_len) || + copy_from_user(dp_buf, u64_to_user_ptr(req.dec_point), + SM2_POINT_LEN)) { + ret =3D -EFAULT; + goto out_free; + } + + ct_dma =3D cmh_dma_map_single(ct_buf, req.ciphertext_len, + DMA_TO_DEVICE); + dp_dma =3D cmh_dma_map_single(dp_buf, SM2_POINT_LEN, DMA_TO_DEVICE); + pt_dma =3D cmh_dma_map_single(pt_buf, msg_len, DMA_FROM_DEVICE); + if (cmh_dma_map_error(ct_dma) || cmh_dma_map_error(dp_dma) || + cmh_dma_map_error(pt_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + vcq_set_header(&vcq[0], PKE_VCQ_CMDS_MIN); + vcq_add_pke_sm2_dec_hash(&vcq[1], core_id, ct_dma, dp_dma, pt_dma, + req.ciphertext_len, 0); + vcq_add_pke_flush(&vcq[2], core_id); + + ret =3D cmh_tm_submit_sync_mbx(vcq, PKE_VCQ_CMDS_MIN, 1, MGMT_MBX); + +out_unmap: + if (!cmh_dma_map_error(pt_dma)) + cmh_dma_unmap_single(pt_dma, msg_len, DMA_FROM_DEVICE); + if (!cmh_dma_map_error(dp_dma)) + cmh_dma_unmap_single(dp_dma, SM2_POINT_LEN, DMA_TO_DEVICE); + if (!cmh_dma_map_error(ct_dma)) + cmh_dma_unmap_single(ct_dma, req.ciphertext_len, + DMA_TO_DEVICE); + + if (!ret) { + if (copy_to_user(u64_to_user_ptr(req.plaintext), + pt_buf, msg_len)) + ret =3D -EFAULT; + } + +out_free: + kfree_sensitive(pt_buf); + kfree_sensitive(dp_buf); + kfree(ct_buf); + return ret; +} + +/* -- SM2_ENC_HASH ---------------------- */ + +/** + * cmh_mgmt_sm2_enc_hash() - Handle CMH_MGMT_IOC_SM2_ENC_HASH ioctl + * @argp: User-space ioctl argument pointer + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_mgmt_sm2_enc_hash(void __user *argp) +{ + struct cmh_ioctl_sm2_enc_hash req; + struct vcq_cmd vcq[PKE_VCQ_CMDS_MIN]; + u32 ct_len, core_id; + u8 *msg_buf, *ep_buf, *ct_buf; + dma_addr_t msg_dma, ep_dma, ct_dma; + int ret; + + if (copy_from_user(&req, argp, sizeof(req))) + return -EFAULT; + if (req.version !=3D CMH_MGMT_V1) + return -EINVAL; + if (!req.message_len || req.message_len > SM2_MAX_MSG_LEN) + return -EINVAL; + + ct_len =3D SM2_CT_OVERHEAD + req.message_len; + core_id =3D cmh_core_default_id(CMH_CORE_PKE); + + msg_buf =3D kmalloc(req.message_len, GFP_KERNEL); + ep_buf =3D kmalloc(SM2_POINT_LEN, GFP_KERNEL); + ct_buf =3D kzalloc(ct_len, GFP_KERNEL); + if (!msg_buf || !ep_buf || !ct_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + if (copy_from_user(msg_buf, u64_to_user_ptr(req.message), + req.message_len) || + copy_from_user(ep_buf, u64_to_user_ptr(req.enc_point), + SM2_POINT_LEN)) { + ret =3D -EFAULT; + goto out_free; + } + + msg_dma =3D cmh_dma_map_single(msg_buf, req.message_len, DMA_TO_DEVICE); + ep_dma =3D cmh_dma_map_single(ep_buf, SM2_POINT_LEN, DMA_TO_DEVICE); + ct_dma =3D cmh_dma_map_single(ct_buf, ct_len, DMA_FROM_DEVICE); + if (cmh_dma_map_error(msg_dma) || cmh_dma_map_error(ep_dma) || + cmh_dma_map_error(ct_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + vcq_set_header(&vcq[0], PKE_VCQ_CMDS_MIN); + vcq_add_pke_sm2_enc_hash(&vcq[1], core_id, msg_dma, ep_dma, ct_dma, + req.message_len, 0); + vcq_add_pke_flush(&vcq[2], core_id); + + ret =3D cmh_tm_submit_sync_mbx(vcq, PKE_VCQ_CMDS_MIN, 1, MGMT_MBX); + +out_unmap: + if (!cmh_dma_map_error(ct_dma)) + cmh_dma_unmap_single(ct_dma, ct_len, DMA_FROM_DEVICE); + if (!cmh_dma_map_error(ep_dma)) + cmh_dma_unmap_single(ep_dma, SM2_POINT_LEN, DMA_TO_DEVICE); + if (!cmh_dma_map_error(msg_dma)) + cmh_dma_unmap_single(msg_dma, req.message_len, DMA_TO_DEVICE); + + if (!ret) { + if (copy_to_user(u64_to_user_ptr(req.ciphertext), + ct_buf, ct_len)) + ret =3D -EFAULT; + } + +out_free: + kfree(ct_buf); + kfree(ep_buf); + kfree_sensitive(msg_buf); + return ret; +} diff --git a/drivers/crypto/cmh/cmh_sys.c b/drivers/crypto/cmh/cmh_sys.c new file mode 100644 index 000000000000..b01d058e6d89 --- /dev/null +++ b/drivers/crypto/cmh/cmh_sys.c @@ -0,0 +1,376 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- SYS Core VCQ Builders + * + * VCQ builder functions for SYS core datastore commands. Each function + * populates a single vcq_cmd slot. Callers (cmh_mgmt.c, cmh_key.c) + * assemble complete VCQs by combining header + command(s) + flush, + * then submit via cmh_tm_submit_sync(). + * + * Hardware-required datastore semantics + * -------------------------------------- + * The commands below (NEW, WRITE, DATA, FIND, DELETE, FLUSH) are + * direct mappings of the eSW firmware SYS core command set. The + * eSW maintains per-mailbox datastore namespaces with two object + * classes: + * + * SYS_REF_TEMP -- Temporary objects. Lifetime is scoped to the + * current mailbox slot; reclaimed automatically + * when the slot is reused or on explicit FLUSH. + * Used for raw-key provisioning on every VCQ. + * + * SYS_REF_PERSIST -- Persistent objects. Survive across slots; + * require explicit DELETE to reclaim. Identified + * by a 64-bit Content ID (CID) and resolved to + * a per-MBX ref via SYS_CMD_FIND. + * + * These semantics are hardware requirements, not driver policy. + * The per-MBX temp-stack and per-MBX ref namespace are eSW firmware + * design constraints that cannot be changed by the kernel driver. + */ + +#include + +#include "cmh_sys.h" + +/** + * vcq_add_sys_flush() - Build a SYS_FLUSH VCQ command + * @slot: VCQ command slot to populate + */ +void vcq_add_sys_flush(struct vcq_cmd *slot) +{ + vcq_add_flush(slot, CORE_ID_SYS); +} + +/** + * vcq_add_sys_new() - Build a SYS_NEW VCQ command + * @slot: VCQ command slot to populate + * @cid: Content identifier for the new datastore object + * @ref_dma: DMA address of the object reference buffer + * @len: Length of the object data in bytes + */ +void vcq_add_sys_new(struct vcq_cmd *slot, u64 cid, u64 ref_dma, u32 len) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(CORE_ID_SYS, 0, 1, SYS_CMD_NEW); + slot->hwc.sys.cmd_new.cid =3D cid; + slot->hwc.sys.cmd_new.ref =3D ref_dma; + slot->hwc.sys.cmd_new.len =3D len; +} + +/** + * vcq_add_sys_write() - Build a SYS_WRITE VCQ command + * @slot: VCQ command slot to populate + * @ref: Datastore object reference handle + * @src_dma: DMA address of source data buffer + * @wrap_key: Wrapping key reference (0 if none) + * @len: Length of data to write in bytes + * @sys_type: Datastore object type identifier + */ +void vcq_add_sys_write(struct vcq_cmd *slot, u64 ref, u64 src_dma, + u64 wrap_key, u32 len, u32 sys_type) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(CORE_ID_SYS, 0, 1, SYS_CMD_WRITE); + slot->hwc.sys.cmd_write.ref =3D ref; + slot->hwc.sys.cmd_write.src =3D src_dma; + slot->hwc.sys.cmd_write.key =3D wrap_key; + slot->hwc.sys.cmd_write.len =3D len; + slot->hwc.sys.cmd_write.type =3D sys_type; +} + +/** + * vcq_add_sys_read() - Build a SYS_READ VCQ command + * @slot: VCQ command slot to populate + * @ref: Datastore object reference handle + * @dst_dma: DMA address of destination buffer + * @wrap_key: Wrapping key reference (0 if none) + * @len: Length of data to read in bytes + */ +void vcq_add_sys_read(struct vcq_cmd *slot, u64 ref, u64 dst_dma, + u64 wrap_key, u32 len) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(CORE_ID_SYS, 0, 1, SYS_CMD_READ); + slot->hwc.sys.cmd_read.ref =3D ref; + slot->hwc.sys.cmd_read.dst =3D dst_dma; + slot->hwc.sys.cmd_read.key =3D wrap_key; + slot->hwc.sys.cmd_read.len =3D len; +} + +/** + * vcq_add_sys_data() - Build a SYS_DATA VCQ command + * @slot: VCQ command slot to populate + * @ref: Datastore object reference handle + * @dst_dma: DMA address of destination buffer + * @len: Length of data section to read in bytes + */ +void vcq_add_sys_data(struct vcq_cmd *slot, u64 ref, u64 dst_dma, u32 len) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(CORE_ID_SYS, 0, 1, SYS_CMD_DATA); + slot->hwc.sys.cmd_data.ref =3D ref; + slot->hwc.sys.cmd_data.dst =3D dst_dma; + slot->hwc.sys.cmd_data.len =3D len; +} + +/** + * vcq_add_sys_find() - Build a SYS_FIND VCQ command + * @slot: VCQ command slot to populate + * @cid: Content identifier to search for + * @dst_dma: DMA address of destination buffer for result + * @len: Length of destination buffer in bytes + */ +void vcq_add_sys_find(struct vcq_cmd *slot, u64 cid, u64 dst_dma, u32 len) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(CORE_ID_SYS, 0, 1, SYS_CMD_FIND); + slot->hwc.sys.cmd_find.cid =3D cid; + slot->hwc.sys.cmd_find.dst =3D dst_dma; + slot->hwc.sys.cmd_find.len =3D len; +} + +/** + * vcq_add_sys_list() - Build a SYS_LIST VCQ command + * @slot: VCQ command slot to populate + * @ref: Datastore object reference for enumeration start + * @dst_dma: DMA address of destination buffer for list + * @len: Length of destination buffer in bytes + */ +void vcq_add_sys_list(struct vcq_cmd *slot, u64 ref, u64 dst_dma, u32 len) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(CORE_ID_SYS, 0, 1, SYS_CMD_LIST); + slot->hwc.sys.cmd_list.ref =3D ref; + slot->hwc.sys.cmd_list.dst =3D dst_dma; + slot->hwc.sys.cmd_list.len =3D len; +} + +/** + * vcq_add_sys_grant() - Build a SYS_GRANT VCQ command + * @slot: VCQ command slot to populate + * @ref: Datastore object reference handle + * @read: Read permission bitmask + * @write: Write permission bitmask + * @execute: Execute permission bitmask + */ +void vcq_add_sys_grant(struct vcq_cmd *slot, u64 ref, u64 read, + u64 write, u64 execute) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(CORE_ID_SYS, 0, 1, SYS_CMD_GRANT); + slot->hwc.sys.cmd_grant.ref =3D ref; + slot->hwc.sys.cmd_grant.read =3D read; + slot->hwc.sys.cmd_grant.write =3D write; + slot->hwc.sys.cmd_grant.execute =3D execute; +} + +/** + * vcq_add_sys_export() - Build a SYS_EXPORT VCQ command + * @slot: VCQ command slot to populate + * @cid: Content identifier of object to export + * @dst_dma: DMA address of destination buffer for wrapped blob + * @wrap_key: Wrapping key reference for export + * @len: Length of destination buffer in bytes + */ +void vcq_add_sys_export(struct vcq_cmd *slot, u64 cid, u64 dst_dma, + u64 wrap_key, u32 len) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(CORE_ID_SYS, 0, 1, SYS_CMD_EXPORT); + slot->hwc.sys.cmd_export.cid =3D cid; + slot->hwc.sys.cmd_export.dst =3D dst_dma; + slot->hwc.sys.cmd_export.key =3D wrap_key; + slot->hwc.sys.cmd_export.len =3D len; +} + +/** + * vcq_add_sys_import() - Build a SYS_IMPORT VCQ command + * @slot: VCQ command slot to populate + * @src_dma: DMA address of wrapped datastore blob to import + * @wrap_key: Wrapping key reference for unwrapping + * @len: Length of wrapped blob in bytes + */ +void vcq_add_sys_import(struct vcq_cmd *slot, u64 src_dma, + u64 wrap_key, u32 len) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(CORE_ID_SYS, 0, 1, SYS_CMD_IMPORT); + slot->hwc.sys.cmd_import.src =3D src_dma; + slot->hwc.sys.cmd_import.key =3D wrap_key; + slot->hwc.sys.cmd_import.len =3D len; +} + +/* -- KIC Core VCQ Builders --------------------- */ + +/** + * vcq_add_kic_hkdf1() - Build a KIC HKDF-Expand VCQ command + * @slot: VCQ command slot to populate + * @dst: Datastore reference for derived key output + * @base: Datastore reference for base key input + * @label_dma: DMA address of HKDF label/info buffer + * @key_len: Derived key length in bytes + * @label_len: Length of label buffer in bytes + * @type: Derived key datastore type + */ +void vcq_add_kic_hkdf1(struct vcq_cmd *slot, u64 dst, u64 base, + u64 label_dma, u32 key_len, u32 label_len, u32 type) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(CORE_ID_KIC, 0, 1, KIC_CMD_HKDF1); + slot->hwc.kic.cmd_hkdf1.dst =3D dst; + slot->hwc.kic.cmd_hkdf1.base =3D base; + slot->hwc.kic.cmd_hkdf1.label =3D label_dma; + slot->hwc.kic.cmd_hkdf1.llen =3D label_len; + slot->hwc.kic.cmd_hkdf1.len =3D key_len; + slot->hwc.kic.cmd_hkdf1.type =3D type; +} + +/** + * vcq_add_kic_hkdf2() - Build a KIC HKDF-with-salt VCQ command + * @slot: VCQ command slot to populate + * @dst: Datastore reference for derived key output + * @base: Datastore reference for base key input + * @salt: Datastore reference for HKDF salt key + * @label_dma: DMA address of HKDF label/info buffer + * @key_len: Derived key length in bytes + * @label_len: Length of label buffer in bytes + * @type: Derived key datastore type + */ +void vcq_add_kic_hkdf2(struct vcq_cmd *slot, u64 dst, u64 base, u64 salt, + u64 label_dma, u32 key_len, u32 label_len, u32 type) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(CORE_ID_KIC, 0, 1, KIC_CMD_HKDF2); + slot->hwc.kic.cmd_hkdf2.dst =3D dst; + slot->hwc.kic.cmd_hkdf2.base =3D base; + slot->hwc.kic.cmd_hkdf2.salt =3D salt; + slot->hwc.kic.cmd_hkdf2.label =3D label_dma; + slot->hwc.kic.cmd_hkdf2.llen =3D label_len; + slot->hwc.kic.cmd_hkdf2.len =3D key_len; + slot->hwc.kic.cmd_hkdf2.type =3D type; +} + +/** + * vcq_add_kic_aes_cmac_kdf() - Build a KIC AES-CMAC KDF VCQ command + * @slot: VCQ command slot to populate + * @out_key: Datastore reference for derived key output + * @base_key: Datastore reference for base key input + * @label_dma: DMA address of KDF label buffer + * @key_len: Derived key length in bytes + * @label_len: Length of label buffer in bytes + * @type: Derived key datastore type + */ +void vcq_add_kic_aes_cmac_kdf(struct vcq_cmd *slot, u64 out_key, u64 base_= key, + u64 label_dma, u32 key_len, u32 label_len, + u32 type) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(CORE_ID_KIC, 0, 1, KIC_CMD_AES_CMAC_KDF); + slot->hwc.kic.cmd_aes_cmac_kdf.base_key =3D base_key; + slot->hwc.kic.cmd_aes_cmac_kdf.out_key =3D out_key; + slot->hwc.kic.cmd_aes_cmac_kdf.label =3D label_dma; + slot->hwc.kic.cmd_aes_cmac_kdf.key_len =3D key_len; + slot->hwc.kic.cmd_aes_cmac_kdf.label_len =3D label_len; + slot->hwc.kic.cmd_aes_cmac_kdf.type =3D type; +} + +/** + * vcq_add_kic_dkek_derive() - Build a KIC DKEK derivation VCQ command + * @slot: VCQ command slot to populate + * @out_key: Datastore reference for derived DKEK output + * @base_key: Datastore reference for base key input + * @host_id: Host identifier for key binding + * @metadata_dma: DMA address of derivation metadata buffer + * @metadata_len: Length of metadata buffer in bytes + */ +void vcq_add_kic_dkek_derive(struct vcq_cmd *slot, u64 out_key, u64 base_k= ey, + u32 host_id, u64 metadata_dma, u32 metadata_len) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(CORE_ID_KIC, 0, 1, KIC_CMD_DKEK_DERIVE); + slot->hwc.kic.cmd_dkek_derive.base_key =3D base_key; + slot->hwc.kic.cmd_dkek_derive.out_key =3D out_key; + slot->hwc.kic.cmd_dkek_derive.host_id =3D host_id; + slot->hwc.kic.cmd_dkek_derive.metadata =3D metadata_dma; + slot->hwc.kic.cmd_dkek_derive.metadata_len =3D metadata_len; +} + +/* -- DRBG Core VCQ Builders -------------------- */ + +/** + * vcq_add_drbg_reset() - Build a DRBG reset VCQ command + * @slot: VCQ command slot to populate + * + * Issues DRBG_CMD_RESET which clears the instantiated state, allowing + * a subsequent CONFIG to proceed without a double-instantiate error. + */ +void vcq_add_drbg_reset(struct vcq_cmd *slot) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(CORE_ID_DRBG, 0, 1, DRBG_CMD_RESET); +} + +/** + * vcq_add_drbg_config() - Build a DRBG configuration VCQ command + * @slot: VCQ command slot to populate + * @ratio: Entropy-to-output ratio + * @strength: Security strength in bits + */ +void vcq_add_drbg_config(struct vcq_cmd *slot, u32 ratio, u32 strength) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(CORE_ID_DRBG, 0, 1, DRBG_CMD_CONFIG); + slot->hwc.drbg.cmd_config.entropy_ratio =3D ratio; + slot->hwc.drbg.cmd_config.security_strength =3D strength; +} + +/** + * vcq_add_drbg_datastore() - Build a DRBG datastore setup VCQ command + * @slot: VCQ command slot to populate + * @ref: Datastore object reference handle + * @len: Length of datastore allocation in bytes + * @type: Datastore object type + */ +void vcq_add_drbg_datastore(struct vcq_cmd *slot, u64 ref, u32 len, u32 ty= pe) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(CORE_ID_DRBG, 0, 1, DRBG_CMD_DATASTORE); + slot->hwc.drbg.cmd_datastore.ref =3D ref; + slot->hwc.drbg.cmd_datastore.len =3D len; + slot->hwc.drbg.cmd_datastore.type =3D type; +} + +/* -- EAC Core VCQ Builder ---------------------- */ + +/** + * vcq_add_eac_read() - Build an EAC read VCQ command + * @slot: VCQ command slot to populate + * @dst_dma: DMA address of destination buffer + * @len: Length of data to read in bytes + */ +void vcq_add_eac_read(struct vcq_cmd *slot, u64 dst_dma, u32 len) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(CORE_ID_EAC, 0, 1, EAC_CMD_READ); + slot->hwc.eac.cmd_read.dst =3D dst_dma; + slot->hwc.eac.cmd_read.len =3D len; +} diff --git a/drivers/crypto/cmh/include/cmh_key.h b/drivers/crypto/cmh/incl= ude/cmh_key.h new file mode 100644 index 000000000000..bad69c92b892 --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_key.h @@ -0,0 +1,82 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Per-transform key context + * + * Per-transform key context used by all keyed crypto algorithms (AES, + * SM4, CCP, HMAC, KMAC). Stores raw key bytes supplied via the crypto + * API .setkey() callback: the key is DMA-mapped once at setkey time and + * written to SYS_REF_TEMP in every VCQ. + * + * Each keyed algorithm driver embeds a struct cmh_key_ctx in its + * per-transform context and calls cmh_key_setkey_raw() from its + * .setkey() callback. + * + * Raw-key atomicity (SYS_REF_TEMP) + * --------------------------------- + * SYS_CMD_WRITE to SYS_REF_TEMP is packed into the same VCQ as the + * algorithm commands (AES_CMD_INIT, HC_CMD_HMAC, etc.). SYS_REF_TEMP + * is per-MBX -- the CMH eSW allocates it in the tail of each mailbox's + * own VCQ buffer (mbx_alloc_temp), so concurrent raw-key requests on + * different MBXes do not interfere. + */ + +#ifndef CMH_KEY_H +#define CMH_KEY_H + +#include +#include "cmh_config.h" +#include "cmh_vcq.h" + +/* Key context mode */ +enum cmh_key_mode { + CMH_KEY_NONE =3D 0, /* no key set yet */ + CMH_KEY_RAW, /* raw key bytes in memory */ +}; + +/* Per-transform key context */ +struct cmh_key_ctx { + enum cmh_key_mode mode; + struct { + u8 *data; /* kmemdup'd raw key bytes */ + u32 len; /* key length in bytes */ + u32 sys_type; /* SYS_TYPE_SET(flags, core_id) */ + dma_addr_t dma; /* pre-mapped DMA addr (DMA_TO_DEVICE) */ + } raw; +}; + +/** + * cmh_key_setkey_raw() - Store raw key bytes in the transform context + * @ctx: Per-transform key context + * @key: Raw key bytes + * @keylen: Key length in bytes + * @core_id: Target algorithm core (e.g. CORE_ID_AES) + * + * SYS_TYPE_FLAG_PT is set so the written temp key + * can be read back as plaintext if needed. The actual SYS_CMD_WRITE + * to SYS_REF_TEMP is deferred to each encrypt/decrypt VCQ, where it + * is packed inline for atomicity. + * + * Return: 0 on success, -ENOMEM on allocation failure. + */ +int cmh_key_setkey_raw(struct cmh_key_ctx *ctx, const u8 *key, + u32 keylen, u32 core_id); + +/** + * cmh_key_destroy() - Free key resources + * @ctx: Per-transform key context + * + * Zeroises and frees the raw key buffer. + */ +void cmh_key_destroy(struct cmh_key_ctx *ctx); + +/** + * cmh_ds_type_to_core_id() - Map datastore key type to core ID + * @ds_type: CMH_DS_* key type constant + * + * Return: Corresponding CORE_ID_*, or CORE_ID_NUM (0x1F) on + * unrecognised type (caller should return -EINVAL). + */ +u32 cmh_ds_type_to_core_id(u32 ds_type); + +#endif /* CMH_KEY_H */ diff --git a/drivers/crypto/cmh/include/cmh_mgmt.h b/drivers/crypto/cmh/inc= lude/cmh_mgmt.h new file mode 100644 index 000000000000..a0da71462bf8 --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_mgmt.h @@ -0,0 +1,69 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH -- Key Management misc_device (/dev/cmh_mgmt) + * + * ioctl interface for key CRUD + datastore export/import, + * PKE operations (RSA, ECDSA, ECDH, EdDSA), + * and PQC operations (ML-KEM, ML-DSA, SLH-DSA). + * + * Registered alongside crypto algorithms in module_init, + * unregistered before them in module_exit. + */ + +#ifndef CMH_MGMT_H +#define CMH_MGMT_H + +#ifdef CONFIG_CRYPTO_DEV_CMH_MGMT + +/* + * Pin all mgmt ioctls to MBX 0 for DS ownership and SYS_REF_TEMP scope. + * Shared by cmh_mgmt.c, cmh_mgmt_pke.c, cmh_mgmt_pqc.c, cmh_pke_sm2.c. + */ +#define MGMT_MBX 0 + +/* Maximum DMA buffer size for key data / datastore blobs */ +#define CMH_MGMT_MAX_DATA_LEN (256 * 1024) /* 256 KB */ + +/* + * Scrub up to two orphaned datastore objects on an ioctl error path. + * Grant-with-no-access wipes the key material; the datastore stack space + * itself is only reclaimed by a full reset. Pass 0 to skip. + */ +void cmh_mgmt_ds_scrub(u64 ref0, u64 ref1); + +int cmh_mgmt_register(void); +void cmh_mgmt_unregister(void); + +/* -- PKE ioctl handlers (cmh_mgmt_pke.c) -- */ +int cmh_mgmt_pke_rsa_enc(void __user *argp); +int cmh_mgmt_pke_rsa_dec(void __user *argp); +int cmh_mgmt_pke_rsa_crt_dec(void __user *argp); +int cmh_mgmt_pke_rsa_keygen(void __user *argp); +int cmh_mgmt_pke_ecdsa_sign(void __user *argp); +int cmh_mgmt_pke_ecdh(void __user *argp); +int cmh_mgmt_pke_ecdh_keygen(void __user *argp); +int cmh_mgmt_pke_eddsa_sign(void __user *argp); +int cmh_mgmt_pke_eddsa_verify(void __user *argp); +int cmh_mgmt_pke_ec_keygen(void __user *argp); +int cmh_mgmt_pke_ec_pubgen(void __user *argp); +int cmh_mgmt_pke_eddsa_keygen_sca(void __user *argp); + +/* -- PQC ioctl handlers (cmh_mgmt_pqc.c) -- */ +int cmh_mgmt_ml_kem_keygen(void __user *argp); +int cmh_mgmt_ml_kem_enc(void __user *argp); +int cmh_mgmt_ml_kem_dec(void __user *argp); +int cmh_mgmt_ml_dsa_keygen(void __user *argp); +int cmh_mgmt_ml_dsa_sign(void __user *argp); +int cmh_mgmt_slhdsa_keygen(void __user *argp); +int cmh_mgmt_slhdsa_sign(void __user *argp); +int cmh_mgmt_slhdsa_sign_prehash(void __user *argp); + +#else /* !CONFIG_CRYPTO_DEV_CMH_MGMT */ + +static inline int cmh_mgmt_register(void) { return 0; } +static inline void cmh_mgmt_unregister(void) { } + +#endif /* CONFIG_CRYPTO_DEV_CMH_MGMT */ + +#endif /* CMH_MGMT_H */ diff --git a/drivers/crypto/cmh/include/cmh_pke.h b/drivers/crypto/cmh/incl= ude/cmh_pke.h new file mode 100644 index 000000000000..dcfdb3fc3cd6 --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_pke.h @@ -0,0 +1,245 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- PKE Common Types and Helpers + * + * Shared definitions for RSA, ECDSA, ECDH, EdDSA, and SM2 drivers. + * Curve -> coordinate-length mapping, VCQ byte-swap flags, and + * common VCQ builder prototypes. + */ + +#ifndef CMH_PKE_H +#define CMH_PKE_H + +#include +#include "cmh_vcq.h" +#include "cmh_pke_abi.h" + +/* VCQ byte-swap flags for DMA transfers (per CMH VCQ ABI) */ +#define VCQ_FLAG_SWAP_BYTES 0x400000U +#define VCQ_FLAG_SWAP_WORDS 0x200000U + +/* VCQ byte-swap flags for PKE -- big-endian data on LE bus */ +#define PKE_SWAP_FLAGS (VCQ_FLAG_SWAP_BYTES | VCQ_FLAG_SWAP_WORDS) + +/* VCQ layout: header + [SYS_WRITE] + PKE_CMD + flush */ +#define PKE_VCQ_CMDS_MIN 3 /* header + cmd + flush */ +#define PKE_VCQ_CMDS_MAX 4 /* header + SYS_WRITE + cmd + flush */ + +/* Max RSA key size in bytes (4096 bits) */ +#define PKE_RSA_MAX_BYTES 512 +#define PKE_RSA_MIN_BITS 1024 +#define PKE_RSA_MAX_BITS 4096 + +/* EdDSA SCA: Ed448 blinded private key length (bytes) */ +#define PKE_ED448_SK_SCA_LEN 226 + +/** + * pke_curve_clen() - Get EC curve coordinate length in bytes + * @curve: PKE curve identifier (PKE_CURVE_*) + * + * Return: Coordinate length in bytes, or 0 for unknown curves. + */ +static inline u32 pke_curve_clen(u32 curve) +{ + switch (curve) { + case PKE_CURVE_P192: + case PKE_CURVE_BP192R1: + return 24; + case PKE_CURVE_P224: + case PKE_CURVE_BP224R1: + return 28; + case PKE_CURVE_P256: + case PKE_CURVE_SECP256K1: + case PKE_CURVE_BP256R1: + case PKE_CURVE_ANSSI_FRP256V1: + case PKE_CURVE_SM2: + case PKE_CURVE_25519: + return 32; + case PKE_CURVE_BP320R1: + return 40; + case PKE_CURVE_P384: + case PKE_CURVE_BP384R1: + return 48; + case PKE_CURVE_BP512R1: + return 64; + case PKE_CURVE_P521: + return 68; /* ceil(521/8) =3D 66, ABI uses ALIGN(66, 4) =3D 68 */ + case PKE_CURVE_448: + return 56; + default: + return 0; + } +} + +/** + * pke_curve_bits() - Get EC curve size in bits + * @curve: PKE curve identifier (PKE_CURVE_*) + * + * Return: Curve size in bits, or 0 for unknown curves. + */ +static inline u32 pke_curve_bits(u32 curve) +{ + switch (curve) { + case PKE_CURVE_P192: + case PKE_CURVE_BP192R1: + return 192; + case PKE_CURVE_P224: + case PKE_CURVE_BP224R1: + return 224; + case PKE_CURVE_P256: + case PKE_CURVE_SECP256K1: + case PKE_CURVE_BP256R1: + case PKE_CURVE_ANSSI_FRP256V1: + case PKE_CURVE_SM2: + case PKE_CURVE_25519: + return 256; + case PKE_CURVE_BP320R1: + return 320; + case PKE_CURVE_P384: + case PKE_CURVE_BP384R1: + return 384; + case PKE_CURVE_BP512R1: + return 512; + case PKE_CURVE_P521: + return 521; + case PKE_CURVE_448: + return 448; + default: + return 0; + } +} + +/** + * pke_eddsa_key_len() - Get EdDSA key/pubkey length + * @curve: PKE curve identifier (PKE_CURVE_25519 or PKE_CURVE_448) + * + * Ed25519 uses 32 bytes (=3D=3D clen), Ed448 uses 57 bytes (clen + 1 + * flag byte per RFC 8032). Signature length is 2 * pke_eddsa_key_len(). + * + * Return: Key length in bytes. + */ +static inline u32 pke_eddsa_key_len(u32 curve) +{ + u32 clen =3D pke_curve_clen(curve); + + return (curve =3D=3D PKE_CURVE_448) ? clen + 1 : clen; +} + +/** + * pke_curve_is_edwards() - Check if curve uses Edwards form + * @curve: PKE curve identifier (PKE_CURVE_*) + * + * Return: true for Curve25519 and Curve448, false otherwise. + */ +static inline bool pke_curve_is_edwards(u32 curve) +{ + return curve =3D=3D PKE_CURVE_25519 || curve =3D=3D PKE_CURVE_448; +} + +/** + * pke_swap_flags() - Get VCQ byte-swap flags for a given curve + * @curve: PKE curve identifier (PKE_CURVE_*) + * + * Weierstrass curves need byte+word swap; Edwards curves do not. + * + * Return: VCQ swap flags to OR into the command ID. + */ +static inline u32 pke_swap_flags(u32 curve) +{ + return pke_curve_is_edwards(curve) ? 0 : PKE_SWAP_FLAGS; +} + +/* Common VCQ builder prototypes */ + +void vcq_add_pke_flush(struct vcq_cmd *slot, u32 core_id); + +void vcq_add_pke_rsa_enc(struct vcq_cmd *slot, u32 core_id, u32 bits, u32 = e_len, + u64 e_dma, u64 n_dma, u64 m_dma, u64 c_dma, + u32 flags); + +void vcq_add_pke_rsa_dec(struct vcq_cmd *slot, u32 core_id, u32 bits, u32 = e_len, + u64 e_dma, u64 n_dma, u64 c_dma, u64 m_dma, + u64 d_ref, u32 flags); + +void vcq_add_pke_rsa_crt_dec(struct vcq_cmd *slot, u32 core_id, u32 bits, = u32 e_len, + u64 e_dma, u64 n_dma, u64 c_dma, u64 m_dma, + u64 crt_ref, u32 flags); + +void vcq_add_pke_ecdsa_verify(struct vcq_cmd *slot, u32 core_id, u32 curve= , u32 dlen, + u64 pk_dma, u64 dig_dma, u64 sig_dma, + u64 rp_dma, u32 flags); + +void vcq_add_pke_ecdsa_sign(struct vcq_cmd *slot, u32 core_id, u32 curve, = u32 sklen, + u64 dig_dma, u64 sig_dma, u64 sk_ref, + u32 dlen, u32 flags); + +void vcq_add_pke_ecdsa_pubgen(struct vcq_cmd *slot, u32 core_id, u32 curve= , u32 sklen, + u64 pk_dma, u64 sk_ref, u32 flags); + +void vcq_add_pke_ecdsa_keygen(struct vcq_cmd *slot, u32 core_id, u32 curve= , u32 sklen, + u64 sk_ref, u32 sk_type, u32 flags); + +void vcq_add_pke_ecdh_keygen(struct vcq_cmd *slot, u32 core_id, u32 curve,= u32 sklen, + u64 pkx_dma, u64 sk_ref, u32 flags); + +void vcq_add_pke_ecdh(struct vcq_cmd *slot, u32 core_id, u32 curve, u32 sk= len, + u32 sslen, u32 ss_type, u64 peer_dma, u64 sk_ref, + u64 ss_ref, u32 flags); + +void vcq_add_pke_eddsa_verify(struct vcq_cmd *slot, u32 core_id, u32 curve= , u32 dlen, + u64 pky_dma, u64 dig_dma, u64 sig_dma, + u64 rp_dma, u32 flags); + +void vcq_add_pke_eddsa_sign(struct vcq_cmd *slot, u32 core_id, u32 curve, = u32 sklen, + u64 dig_dma, u64 sig_dma, u64 sk_ref, + u32 dlen, u32 flags); + +void vcq_add_pke_eddsa_pubgen(struct vcq_cmd *slot, u32 core_id, u32 curve= , u32 sklen, + u64 pky_dma, u64 sk_ref, u32 flags); + +void vcq_add_pke_eddsa_keygen_sca(struct vcq_cmd *slot, u32 core_id, u32 c= urve, + u64 sk_ref, u64 sca_sk_ref); + +/* SM2 VCQ builders */ + +void vcq_add_pke_sm2_ecdh_keygen(struct vcq_cmd *slot, u32 core_id, u64 no= nce_dma, + u64 session_key_dma, u32 nonce_len, u32 flags); + +void vcq_add_pke_sm2_ecdh(struct vcq_cmd *slot, u32 core_id, u32 nonce_len, + u32 private_key_len, u64 nonce_dma, + u64 peer_pk_dma, u64 peer_sk_dma, + u64 priv_ref, u64 sp_ref, u32 sp_type, u32 flags); + +void vcq_add_pke_sm2_dec_point(struct vcq_cmd *slot, u32 core_id, u32 ct_l= en, + u32 pk_len, u64 ct_dma, u64 dp_dma, + u64 priv_ref, u32 flags); + +void vcq_add_pke_sm2_enc_point(struct vcq_cmd *slot, u32 core_id, u64 nonc= e_dma, + u64 pk_dma, u64 ct_dma, u64 ep_dma, + u32 nonce_len, u32 flags); + +void vcq_add_pke_sm2_id_digest(struct vcq_cmd *slot, u32 core_id, u64 id_d= ma, + u64 pk_dma, u64 dig_dma, u32 id_len, + u32 flags); + +void vcq_add_pke_sm2_ecdh_hash(struct vcq_cmd *slot, u32 core_id, u64 peer= _dig_dma, + u64 dig_dma, u64 sp_ref, u64 sk_ref, + u32 sk_type, u32 flags); + +void vcq_add_pke_sm2_dec_hash(struct vcq_cmd *slot, u32 core_id, u64 ct_dm= a, + u64 dp_dma, u64 pt_dma, u32 ct_len, u32 flags); + +void vcq_add_pke_sm2_enc_hash(struct vcq_cmd *slot, u32 core_id, u64 msg_d= ma, + u64 ep_dma, u64 ct_dma, u32 msg_len, u32 flags); + +/* Registration */ + +int cmh_pke_rsa_register(void); +void cmh_pke_rsa_unregister(void); +int cmh_pke_ecdsa_register(void); +void cmh_pke_ecdsa_unregister(void); +int cmh_pke_ecdh_register(void); +void cmh_pke_ecdh_unregister(void); + +#endif /* CMH_PKE_H */ diff --git a/drivers/crypto/cmh/include/cmh_pke_sm2.h b/drivers/crypto/cmh/= include/cmh_pke_sm2.h new file mode 100644 index 000000000000..a2c7164b8d49 --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_pke_sm2.h @@ -0,0 +1,30 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- SM2 PKE Ioctl Handler Declarations + * + * SM2 signature (GM/T 0003.2) requires the caller to compute + * ZA =3D SM3(ENTLA || IDA || a || b || xG || yG || xA || yA) + * and pass SM3(ZA || M) as the digest to the sign/verify path. + * The CMH eSW does NOT compute ZA internally; the full + * identity pre-hash is the caller's responsibility. + * + * For the in-kernel akcipher "sm2" algorithm this means the + * caller (e.g. asymmetric_key subsystem) must pre-hash with ZA + * before invoking verify. The SM2_ID_DIGEST ioctl below can + * compute ZA for userspace callers of the misc-device path. + */ + +#ifndef CMH_PKE_SM2_H +#define CMH_PKE_SM2_H + +int cmh_mgmt_sm2_ecdh_keygen(void __user *argp); +int cmh_mgmt_sm2_ecdh(void __user *argp); +int cmh_mgmt_sm2_dec_point(void __user *argp); +int cmh_mgmt_sm2_enc_point(void __user *argp); +int cmh_mgmt_sm2_id_digest(void __user *argp); +int cmh_mgmt_sm2_ecdh_hash(void __user *argp); +int cmh_mgmt_sm2_dec_hash(void __user *argp); +int cmh_mgmt_sm2_enc_hash(void __user *argp); + +#endif /* CMH_PKE_SM2_H */ diff --git a/drivers/crypto/cmh/include/cmh_pqc.h b/drivers/crypto/cmh/incl= ude/cmh_pqc.h new file mode 100644 index 000000000000..cd4761a0ce5c --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_pqc.h @@ -0,0 +1,25 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- PQC Algorithm Registration + * + * Registration/unregistration functions for PQC akcipher algorithms: + * ML-DSA, SLH-DSA, LMS, XMSS. + */ + +#ifndef CMH_PQC_H +#define CMH_PQC_H + +int cmh_pqc_mldsa_register(void); +void cmh_pqc_mldsa_unregister(void); + +int cmh_pqc_slhdsa_register(void); +void cmh_pqc_slhdsa_unregister(void); + +int cmh_pqc_lms_register(void); +void cmh_pqc_lms_unregister(void); + +int cmh_pqc_xmss_register(void); +void cmh_pqc_xmss_unregister(void); + +#endif /* CMH_PQC_H */ diff --git a/drivers/crypto/cmh/include/cmh_sys.h b/drivers/crypto/cmh/incl= ude/cmh_sys.h new file mode 100644 index 000000000000..dd336b67bd65 --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_sys.h @@ -0,0 +1,111 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- SYS Core VCQ Builders + * + * VCQ builder functions for SYS core commands (NEW, WRITE, READ, + * FIND, GRANT, DATA, EXPORT, IMPORT). Each builder populates one + * vcq_cmd slot with the appropriate magic, command ID, and payload. + * + * Callers combine these with vcq_set_header() + vcq_add_flush() + * and submit via cmh_tm_submit_sync(). + */ + +#ifndef CMH_SYS_H +#define CMH_SYS_H + +#include "cmh_vcq.h" + +void vcq_add_sys_new(struct vcq_cmd *slot, u64 cid, u64 ref_dma, u32 len); +void vcq_add_sys_write(struct vcq_cmd *slot, u64 ref, u64 src_dma, + u64 wrap_key, u32 len, u32 sys_type); +void vcq_add_sys_read(struct vcq_cmd *slot, u64 ref, u64 dst_dma, + u64 wrap_key, u32 len); +void vcq_add_sys_data(struct vcq_cmd *slot, u64 ref, u64 dst_dma, u32 len); +void vcq_add_sys_find(struct vcq_cmd *slot, u64 cid, u64 dst_dma, u32 len); +void vcq_add_sys_list(struct vcq_cmd *slot, u64 ref, u64 dst_dma, u32 len); +void vcq_add_sys_grant(struct vcq_cmd *slot, u64 ref, u64 read, + u64 write, u64 execute); +void vcq_add_sys_export(struct vcq_cmd *slot, u64 cid, u64 dst_dma, + u64 wrap_key, u32 len); +void vcq_add_sys_import(struct vcq_cmd *slot, u64 src_dma, + u64 wrap_key, u32 len); + +/* KIC core VCQ builders */ +void vcq_add_kic_hkdf1(struct vcq_cmd *slot, u64 dst, u64 base, + u64 label_dma, u32 key_len, u32 label_len, u32 type); +void vcq_add_kic_hkdf2(struct vcq_cmd *slot, u64 dst, u64 base, u64 salt, + u64 label_dma, u32 key_len, u32 label_len, u32 type); +void vcq_add_kic_aes_cmac_kdf(struct vcq_cmd *slot, u64 out_key, u64 base_= key, + u64 label_dma, u32 key_len, u32 label_len, + u32 type); +void vcq_add_kic_dkek_derive(struct vcq_cmd *slot, u64 out_key, u64 base_k= ey, + u32 host_id, u64 metadata_dma, u32 metadata_len); + +/* DRBG core VCQ builders */ +void vcq_add_drbg_reset(struct vcq_cmd *slot); +void vcq_add_drbg_config(struct vcq_cmd *slot, u32 ratio, u32 strength); +void vcq_add_drbg_datastore(struct vcq_cmd *slot, u64 ref, u32 len, u32 ty= pe); + +/* QSE core VCQ builders */ +void vcq_add_qse_flush(struct vcq_cmd *slot, u32 core_id); +void vcq_add_qse_ml_kem_keygen(struct vcq_cmd *slot, u32 core_id, u32 k, u= 32 flags, + u64 seed, u64 z, u64 ek, u64 dk, u32 dk_type, + bool masked); +void vcq_add_qse_ml_kem_enc(struct vcq_cmd *slot, u32 core_id, u32 k, u32 = flags, + u64 coin, u64 ek, u64 ct, u64 ss, u32 ss_type, + bool masked); +void vcq_add_qse_ml_kem_dec(struct vcq_cmd *slot, u32 core_id, u32 k, u32 = flags, + u64 ct, u64 dk, u64 ss, u32 ss_type, + bool masked); +void vcq_add_qse_ml_dsa_keygen(struct vcq_cmd *slot, u32 core_id, u32 mode= , u32 flags, + u64 seed, u64 pk, u64 sk, u32 sk_type, + bool masked); +void vcq_add_qse_ml_dsa_sign(struct vcq_cmd *slot, u32 core_id, u32 mode, = u32 flags, + u64 rnd, u64 m, u64 sk, u64 sig, u32 mlen, + bool masked); +void vcq_add_qse_ml_dsa_verify(struct vcq_cmd *slot, u32 core_id, u32 mode= , u32 flags, + u64 m, u64 pk, u64 sig, u32 mlen); + +/* HCQ core VCQ builders */ +void vcq_add_hcq_flush(struct vcq_cmd *slot, u32 core_id); +void vcq_add_hcq_slhdsa_keygen(struct vcq_cmd *slot, u32 core_id, u32 para= m_set, + u32 seed_len, u32 pk_len, u32 sk_len, + u64 seed, u64 pk, u64 sk); +void vcq_add_hcq_slhdsa_sign(struct vcq_cmd *slot, u32 core_id, u32 param_= set, + u32 msg_len, u32 ctx_len, + u64 add_random, u64 msg, u64 ctx, + u64 sk, u64 sig); +void vcq_add_hcq_slhdsa_sign_internal(struct vcq_cmd *slot, u32 core_id, u= 32 param_set, + u32 msg_len, u64 add_random, + u64 msg, u64 sk, u64 sig); +void vcq_add_hcq_slhdsa_verify(struct vcq_cmd *slot, u32 core_id, u32 para= m_set, + u32 msg_len, u32 ctx_len, + u64 msg, u64 ctx, u64 pk, u64 sig); +void vcq_add_hcq_slhdsa_sign_prehash(struct vcq_cmd *slot, u32 core_id, + u32 cmd, u32 param_set, u32 prehash_algo, + u32 msg_len, u32 ctx_len, + u64 add_random, u64 msg, u64 ctx, + u64 sk, u64 sig); +void vcq_add_hcq_slhdsa_verify_prehash(struct vcq_cmd *slot, u32 core_id, + u32 cmd, u32 param_set, u32 prehash_algo, + u32 msg_len, u32 ctx_len, + u64 msg, u64 ctx, u64 pk, u64 sig); +void vcq_add_hcq_slhdsa_verify_internal(struct vcq_cmd *slot, u32 core_id,= u32 param_set, + u32 msg_len, u64 msg, u64 pk, u64 sig); +void vcq_add_hcq_slhdsa_pubgen(struct vcq_cmd *slot, u32 core_id, u32 para= m_set, + u32 sk_len, u64 sk, u64 pk); +void vcq_add_hcq_lms_verify(struct vcq_cmd *slot, u32 core_id, u32 lms_hss, + u32 pk_len, u32 sig_len, u32 dig_len, + u64 pk, u64 sig, u64 dig); +void vcq_add_hcq_xmss_verify(struct vcq_cmd *slot, u32 core_id, u32 xmss_m= t, + u32 pk_len, u32 sig_len, u32 dig_len, + u64 pk, u64 sig, u64 dig); + +/* SYS core flush */ +void vcq_add_sys_flush(struct vcq_cmd *slot); + +/* EAC core VCQ builder */ +void vcq_add_eac_read(struct vcq_cmd *slot, u64 dst_dma, u32 len); + +#endif /* CMH_SYS_H */ diff --git a/include/uapi/linux/cmh_mgmt_ioctl.h b/include/uapi/linux/cmh_m= gmt_ioctl.h new file mode 100644 index 000000000000..19748260404a --- /dev/null +++ b/include/uapi/linux/cmh_mgmt_ioctl.h @@ -0,0 +1,902 @@ +/* SPDX-License-Identifier: GPL-2.0 WITH Linux-syscall-note */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Key Management ioctl Interface (User-Space API) + * + * ioctl commands for /dev/cmh_mgmt -- key CRUD, datastore + * export/import, KIC key derivation, PKE, SM2, and PQC operations. + * + * Relationship to the in-kernel crypto API + * ----------------------------------------- + * Most commands here have no crypto API representation (no transform + * type or verb exists): keystore CRUD, key generation, KIC key + * derivation, ML-KEM encapsulate/decapsulate, SM2 multi-step + * encrypt/decrypt/key-exchange, EdDSA, EAC, and DRBG configuration. + * For these the character device is the only available UAPI. + * + * A bounded subset names primitives the driver ALSO registers with + * the crypto API, and the overlap is intentional: + * - Hardware-held-key operations (RSA decrypt, ECDSA/ML-DSA/SLH-DSA + * sign, ECDH) reference a private key by datastore handle. The + * crypto API set_priv_key()/set_secret() take only raw key bytes + * and cannot name a key that never leaves the hardware; these + * ioctls keep the key hardware-resident. The registered + * transforms serve raw-key in-kernel users -- the paths are + * complementary. + * + * Multi-step protocol flows are documented above the PKE and SM2 + * struct sections. Single-command ioctls are self-documenting. + * + * Versioned structs: user space sets .version =3D CMH_MGMT_V1 so the + * driver can extend structs in the future without breaking ABI. + */ + +#ifndef _UAPI_CMH_MGMT_IOCTL_H +#define _UAPI_CMH_MGMT_IOCTL_H + +#include +#include +#include + +#define CMH_MGMT_V1 1 + +/* Special reference values */ +#define CMH_REF_NONE 0x0000000000000000ULL /* no key (plaintext) */ + +/* Flags for cmh_ioctl_key_new.flags / cmh_ioctl_key_write.flags */ +#define CMH_FLAG_PT _BITUL(16) /* key can be read as plaintext */ +#define CMH_FLAG_XC _BITUL(17) /* key can be exported over XC bus */ +#define CMH_FLAG_SCA _BITUL(18) /* SCA key stored in 2 shares */ +#define CMH_FLAG_MASK (CMH_FLAG_PT | CMH_FLAG_XC | CMH_FLAG_SCA) + +/* + * Datastore key types -- the LKM maps these to core IDs internally. + * User space passes these in cmh_ioctl_key_new.ds_type. + */ +#define CMH_DS_RAW_VALUE 1 +#define CMH_DS_AES_KEY 2 +#define CMH_DS_AES_XTS_KEY 3 +#define CMH_DS_HMAC_KEY 4 +#define CMH_DS_KMAC_KEY 5 +#define CMH_DS_SM4_KEY 6 +#define CMH_DS_CHACHA20_KEY 7 + +/* PKE key types -- all map to CORE_ID_PKE (0x0A) */ +#define CMH_DS_RSA_PRIV_KEY 10 +#define CMH_DS_RSA_PUB_KEY 11 +#define CMH_DS_RSA_CRT_KEY 12 +#define CMH_DS_ECDSA_PRIV_KEY 13 +#define CMH_DS_ECDSA_PUB_KEY 14 +#define CMH_DS_ECDH_PRIV_KEY 15 +#define CMH_DS_EDDSA_PRIV_KEY 16 +#define CMH_DS_SHARED_SECRET 17 +#define CMH_DS_SM2_PRIV_KEY 18 + +/* QSE key types -- map to CORE_ID_QSE (0x09) */ +#define CMH_DS_ML_KEM_DK 20 +#define CMH_DS_ML_DSA_SK 21 + +/* HCQ key types -- map to CORE_ID_HCQ (0x08) */ +#define CMH_DS_SLHDSA_SK 25 + +/* ioctl argument structures */ + +struct cmh_ioctl_key_new { + __u32 version; /* must be CMH_MGMT_V1 */ + __u32 ds_type; /* CMH_DS_* key type */ + __u32 len; /* key length in bytes */ + __u32 flags; /* CMH_FLAG_* (e.g. CMH_FLAG_PT) */ + __u64 cid; /* caller ID (name) for the key */ + __u64 ref; /* [out] CMH eSW returns key_ref here */ +}; + +struct cmh_ioctl_key_write { + __u32 version; + __u32 len; /* key data length */ + __u32 ds_type; /* CMH_DS_* key type */ + __u32 flags; /* CMH_FLAG_* (e.g. CMH_FLAG_PT) */ + __u64 ref; /* key reference from KEY_NEW */ + __u64 wrap_key; /* wrapping key ref (CMH_REF_NONE =3D plaintext) */ + __u64 data; /* user-space pointer to key material */ +}; + +/* + * KEY_READ returns the datastore-internal representation of the object. + * For Weierstrass PKE and RSA key types, KEY_WRITE byte-swaps the input + * into the PKE sidecar's big-endian form before storing, so a plaintext + * KEY_READ of such a key returns those big-endian bytes -- it is NOT a + * byte-for-byte echo of the host-order buffer passed to KEY_WRITE. + * Symmetric keys are stored verbatim (no swap) and wrapped reads return + * an opaque blob. This asymmetry is intentional: KEY_READ exposes the + * stored form the hardware uses. + */ +struct cmh_ioctl_key_read { + __u32 version; + __u32 len; /* buffer length */ + __u64 ref; /* key reference */ + __u64 wrap_key; /* wrapping key ref (CMH_REF_NONE =3D plaintext) */ + __u64 data; /* user-space pointer to output buffer */ + __u32 out_len; /* [out] actual bytes written */ + __u32 __reserved; +}; + +struct cmh_ioctl_key_find { + __u32 version; + __u32 __reserved; + __u64 cid; /* caller ID to search for */ + __u64 ref; /* [out] resolved key reference */ + __u32 len; /* [out] key length */ + __u32 type; /* [out] key type */ +}; + +/* + * KEY_LIST -- iterate datastore objects. + * + * Pass start_ref=3D0 to begin from the first accessible object. + * On return, ref/cid/len/type describe that object. Pass the + * returned ref as start_ref in the next call to advance. Iteration + * ends when ref =3D=3D 0 (no more objects). + */ +struct cmh_ioctl_key_list { + __u32 version; + __u32 __reserved; + __u64 start_ref; /* starting DS reference (0 =3D first) */ + __u64 ref; /* [out] object reference */ + __u64 cid; /* [out] caller ID */ + __u32 len; /* [out] object length */ + __u32 type; /* [out] object type */ +}; + +struct cmh_ioctl_key_grant { + __u32 version; + __u32 __reserved; + __u64 ref; /* key reference */ + __u64 read; /* per-MBX read permission bitfield */ + __u64 write; /* per-MBX write permission bitfield */ + __u64 execute; /* per-MBX execute permission bitfield */ +}; + +/* Export blob overhead beyond the raw object data (bytes) */ +#define CMH_DS_EXPORT_OVERHEAD_WRAPPED 48 /* 16B hdr + 16B nonce + 16B tag= */ +#define CMH_DS_EXPORT_OVERHEAD_PLAIN 16 /* 16B hdr only */ + +/** + * struct cmh_ioctl_ds_export - Export a datastore object to a wrapped blob + * @version: protocol version (CMH_MGMT_V1) + * @len: DMA buffer size; must be >=3D export blob size: + * wrapped: CMH_DS_EXPORT_OVERHEAD_WRAPPED + object_len + * plaintext: CMH_DS_EXPORT_OVERHEAD_PLAIN + object_len + * object_len is known from KEY_NEW or KEY_FIND. + * If too small, the eSW rejects the command (-EIO). + * @cid: caller ID of the object to export + * @wrap_key: wrapping key ref (CMH_REF_NONE =3D plaintext export) + * @data: user-space pointer to output buffer (at least @len bytes) + * @out_len: [out] actual blob bytes written on success + * @__reserved: must be zero + */ +struct cmh_ioctl_ds_export { + __u32 version; + __u32 len; /* buffer length (see sizing rule above) */ + __u64 cid; /* caller ID for response tagging */ + __u64 wrap_key; /* wrapping key ref (CMH_REF_NONE =3D plaintext) */ + __u64 data; /* user-space pointer to output buffer */ + __u32 out_len; /* [out] actual bytes written */ + __u32 __reserved; +}; + +struct cmh_ioctl_ds_import { + __u32 version; + __u32 len; /* blob length */ + __u64 wrap_key; /* wrapping key ref (CMH_REF_NONE =3D plaintext) */ + __u64 data; /* user-space pointer to import blob */ +}; + +/* Flags for cmh_ioctl_kic_hkdf1.flags / cmh_ioctl_kic_hkdf2.flags */ +#define CMH_KIC_FLAG_TEMP 0x01 /* store result in TEMP (not persistent DS)= */ + +/* + * KIC hardware base key references. + * + * Each CMH device has up to 8 hardware base keys provisioned in OTP/fuses. + * These values are passed in the base_key field of KIC ioctl structs. + * The key valid bitmask is visible via R_KIC_KEY_VALID (MMIO 0x100). + */ +#define CMH_KIC_KEY1 0x0000000100000001ULL +#define CMH_KIC_KEY2 0x0000000200000002ULL +#define CMH_KIC_KEY3 0x0000000300000003ULL +#define CMH_KIC_KEY4 0x0000000400000004ULL +#define CMH_KIC_KEY5 0x0000000500000005ULL +#define CMH_KIC_KEY6 0x0000000600000006ULL +#define CMH_KIC_KEY7 0x0000000700000007ULL +#define CMH_KIC_KEY8 0x0000000800000008ULL + +struct cmh_ioctl_kic_hkdf1 { + __u32 version; + __u32 key_len; /* output key length (e.g., 32) */ + __u64 base_key; /* KIC base key reference */ + __u64 cid; /* CID for the new DS entry (ignored if TEMP) */ + __u64 label; /* user-space pointer to label data */ + __u32 label_len; /* label length in bytes */ + __u32 flags; /* CMH_KIC_FLAG_* */ + __u64 ref; /* [out] derived key reference */ +}; + +struct cmh_ioctl_kic_hkdf2 { + __u32 version; + __u32 key_len; /* output key length (e.g., 32) */ + __u64 base_key; /* KIC base key reference */ + __u64 salt_key; /* salt key reference (CMH_REF_NONE =3D no salt) */ + __u64 cid; /* CID for the new DS entry (ignored if TEMP) */ + __u64 label; /* user-space pointer to label data */ + __u32 label_len; /* label length in bytes */ + __u32 flags; /* CMH_KIC_FLAG_* */ + __u64 ref; /* [out] derived key reference */ +}; + +struct cmh_ioctl_kic_aes_cmac_kdf { + __u32 version; + __u32 key_len; /* base & output key length (must be 32) */ + __u64 base_key; /* KIC base key or DS reference */ + __u64 cid; /* CID for the new DS entry (ignored if TEMP) */ + __u64 label; /* user-space pointer to label data */ + __u32 label_len; /* label length in bytes */ + __u32 flags; /* CMH_KIC_FLAG_* */ + __u64 ref; /* [out] derived key reference */ +}; + +#define KIC_DKEK_MAX_METADATA 64 /* max metadata length for DKEK */ + +struct cmh_ioctl_kic_dkek_derive { + __u32 version; + __u32 host_id; /* target host ID (0 =3D caller's own) */ + __u64 base_key; /* KIC base key reference */ + __u64 cid; /* CID for the new DS entry (ignored if TEMP) */ + __u64 metadata; /* user-space pointer to metadata */ + __u32 metadata_len; /* metadata length in bytes */ + __u32 flags; /* CMH_KIC_FLAG_* */ + __u64 ref; /* [out] derived KEK reference */ +}; + +/* -- PKE ioctl argument structures ----------- */ + +/* + * PKE multi-step protocol flows + * + * RSA encrypt/decrypt: + * 1. KEY_NEW(CMH_DS_RSA_PRIV_KEY) + KEY_WRITE -> priv_ref (or RSA_KEYGE= N -> priv_ref) + * 2. RSA_ENC(e, n, plaintext) -> ciphertext (public key =3D raw= e,n) + * 3. RSA_DEC(e, n, ciphertext, priv_ref) -> plaintext (or RSA_CRT_DEC) + * + * ECDSA sign: + * 1. EC_KEYGEN(curve) -> priv_ref (or KEY_NEW + KEY_= WRITE) + * 2. EC_PUBGEN(priv_ref) -> public_key (raw x||y returned) + * 3. ECDSA_SIGN(digest, priv_ref) -> signature + * SM2 sign uses the same path with curve=3DCMH_CURVE_SM2. + * + * ECDH shared secret: + * 1. EC_KEYGEN(curve) -> priv_ref (or KEY_NEW + KEY_= WRITE) + * 2. ECDH_KEYGEN(priv_ref) -> public_key_x (derive pub from p= riv) + * 3. Exchange public keys with peer + * 4. ECDH(peer_key_x, priv_ref) -> shared_secret (raw to host) + * + * EdDSA sign/verify: + * 1. EC_KEYGEN(CURVE_25519 or CURVE_448) -> priv_ref + * 2. EC_PUBGEN(priv_ref) -> public_key + * 3. EDDSA_SIGN(message, priv_ref) -> signature (pure EdDSA, not p= rehash) + * 4. EDDSA_VERIFY(message, signature, public_key_y) + * For Ed448 SCA: EDDSA_KEYGEN_SCA(priv_ref) -> sca_ref (2-share blinded= key) + * + * SM2 encryption (GM/T 0003.4): + * 1. EC_KEYGEN(CMH_CURVE_SM2) -> priv_ref (or KEY_NEW + KEY_= WRITE) + * 2. EC_PUBGEN(priv_ref) -> public_key + * 3. SM2_ENC_POINT(public_key) -> C1, enc_point (nonce_len=3D0: HW= ephemeral) + * 4. SM2_ENC_HASH(enc_point, message) -> ciphertext (C1||C3||C2) + * Decrypt: + * 5. SM2_DEC_POINT(C1, priv_ref) -> dec_point + * 6. SM2_DEC_HASH(ciphertext, dec_point) -> plaintext + * enc_point and dec_point are raw DMA buffers (64B each), not DS refs. + * + * SM2 key exchange (GM/T 0003.3): + * 1. EC_KEYGEN(CMH_CURVE_SM2) -> priv_ref (long-lived, persi= stent DS) + * 2. EC_PUBGEN(priv_ref) -> public_key + * 3. SM2_ID_DIGEST(id, public_key) -> ZA (SM3-based identit= y digest) + * 4. SM2_ECDH_KEYGEN(nonce) -> session_key, r (ephemeral scalar = r) + * - nonce_len=3D32: caller supplies r (deterministic) + * - nonce_len=3D0: HW generates r, writes it back to .nonce + * Exchange session_key with peer. + * 5. SM2_ECDH(r, priv_ref, peer_pub, peer_sess) -> shared_point + * - Must pass the same r from step 4 (nonce_len=3D32) + * - shared_point_ref=3D0: reads back raw shared_point, destroys DS s= lot + * - shared_point_ref=3D&ref: keeps DS slot alive, writes ref for ECD= H_HASH + * 6. SM2_ECDH_HASH(shared_point_ref, ZA_self, ZA_peer) -> shared_key (1= 6B) + * - shared_point_ref is a persistent DS reference from step 5 + * - The DS slot is consumed by the hub; caller should delete it afte= rward + * The nonce r is a raw 32-byte scalar in userspace memory between steps= 4-5. + * The shared_point is a persistent DS ref between steps 5-6. + * The long-lived private key (priv_ref) persists independently. + */ + +struct cmh_ioctl_pke_rsa_enc { + __u32 version; + __u32 bits; /* RSA key size in bits (1024-4096) */ + __u64 e; /* user-space pointer to public exponent */ + __u32 e_len; /* exponent length in bytes */ + __u32 __reserved; + __u64 n; /* user-space pointer to modulus */ + __u64 input; /* user-space pointer to input data */ + __u64 output; /* user-space pointer to output buffer */ +}; + +struct cmh_ioctl_pke_rsa_dec { + __u32 version; + __u32 bits; + __u64 e; /* public exponent */ + __u32 e_len; + __u32 __reserved; + __u64 n; /* modulus */ + __u64 input; /* ciphertext */ + __u64 output; /* plaintext output */ + __u64 key_ref; /* private key DS reference */ +}; + +struct cmh_ioctl_pke_rsa_crt_dec { + __u32 version; + __u32 bits; + __u64 e; + __u32 e_len; + __u32 __reserved; + __u64 n; + __u64 input; + __u64 output; + __u64 crt_ref; /* CRT key DS reference */ +}; + +struct cmh_ioctl_pke_rsa_keygen { + __u32 version; + __u32 bits; /* key size in bits */ + __u64 e; /* user-space pointer to public exponent */ + __u32 e_len; + __u32 flags; /* CMH_FLAG_* */ + __u64 n; /* [out] user-space pointer to modulus buffer */ + __u64 d_cid; /* CID for private key DS entry */ + __u64 d_ref; /* [out] private key reference */ + __u64 crt_cid; /* CID for CRT key DS entry (0 =3D skip CRT) */ + __u64 crt_ref; /* [out] CRT key reference */ +}; + +struct cmh_ioctl_pke_ecdsa_sign { + __u32 version; + __u32 curve; /* ABI curve ID (e.g. 0x03 =3D P-256) */ + __u64 digest; /* user-space pointer to hash digest */ + __u32 digest_len; /* digest length in bytes */ + __u32 __reserved; + __u64 signature; /* [out] user-space pointer to (r,s) */ + __u64 key_ref; /* private key DS reference */ +}; + +struct cmh_ioctl_pke_ecdh { + __u32 version; + __u32 curve; + __u64 peer_key_x; /* user-space pointer to peer public key X */ + __u64 key_ref; /* private key DS reference */ + __u32 flags; /* reserved, must be 0 */ + __u32 __reserved; + __u64 __reserved2; /* reserved, must be 0 */ + __u64 output; /* [out] raw shared secret */ +}; + +struct cmh_ioctl_pke_ecdh_keygen { + __u32 version; + __u32 curve; + __u64 key_ref; /* private key DS reference */ + __u64 public_key_x; /* [out] user-space pointer to public key X */ +}; + +struct cmh_ioctl_pke_eddsa_sign { + __u32 version; + __u32 curve; /* CURVE_25519 or CURVE_448 */ + __u64 digest; /* user-space ptr to message (not digest) */ + __u32 digest_len; + __u32 __reserved; + __u64 signature; /* [out] user-space pointer to signature */ + __u64 key_ref; /* private key DS reference */ +}; + +struct cmh_ioctl_pke_eddsa_verify { + __u32 version; + __u32 curve; + __u64 digest; + __u32 digest_len; + __u32 __reserved; + __u64 signature; + __u64 public_key_y; /* user-space pointer to public key Y */ +}; + +struct cmh_ioctl_pke_ec_keygen { + __u32 version; + __u32 curve; + __u32 flags; /* CMH_FLAG_* */ + __u32 __reserved; + __u64 cid; /* CID for the new key DS entry */ + __u64 ref; /* [out] private key reference */ +}; + +struct cmh_ioctl_pke_ec_pubgen { + __u32 version; + __u32 curve; + __u64 key_ref; /* private key DS reference */ + __u64 public_key; /* [out] user-space pointer to public key */ +}; + +struct cmh_ioctl_pke_eddsa_keygen_sca { + __u32 version; + __u32 curve; /* must be CURVE_448 */ + __u64 key_ref; /* input: normal Ed448 private key DS ref */ + __u64 cid; /* CID for the new SCA key DS entry */ + __u64 sca_ref; /* [out] SCA private key reference */ +}; + +/* + * ioctl numbers -- type 'J', sequential. + * 'C' conflicts with OSS sound, CAPI/ISDN, and COSA WAN drivers; + * 'J' is unregistered in Documentation/userspace-api/ioctl/ioctl-number.r= st. + */ + +#define CMH_MGMT_IOC_MAGIC 'J' + +#define CMH_IOCTL_KEY_NEW _IOWR(CMH_MGMT_IOC_MAGIC, 0x01, struct cmh_ioctl= _key_new) +#define CMH_IOCTL_KEY_WRITE _IOW(CMH_MGMT_IOC_MAGIC, 0x02, struct cmh_ioc= tl_key_write) +#define CMH_IOCTL_KEY_READ _IOWR(CMH_MGMT_IOC_MAGIC, 0x03, struct cmh_ioct= l_key_read) +#define CMH_IOCTL_KEY_FIND _IOWR(CMH_MGMT_IOC_MAGIC, 0x04, struct cmh_ioct= l_key_find) +#define CMH_IOCTL_KEY_GRANT _IOW(CMH_MGMT_IOC_MAGIC, 0x05, struct cmh_ioc= tl_key_grant) +#define CMH_IOCTL_KEY_DELETE _IOW(CMH_MGMT_IOC_MAGIC, 0x06, struct cmh_io= ctl_key_grant) +#define CMH_IOCTL_DS_EXPORT _IOWR(CMH_MGMT_IOC_MAGIC, 0x07, struct cmh_ioc= tl_ds_export) +#define CMH_IOCTL_DS_IMPORT _IOW(CMH_MGMT_IOC_MAGIC, 0x08, struct cmh_ioc= tl_ds_import) +#define CMH_IOCTL_KIC_HKDF1 _IOWR(CMH_MGMT_IOC_MAGIC, 0x09, struct cmh_ioc= tl_kic_hkdf1) +#define CMH_IOCTL_KIC_HKDF2 _IOWR(CMH_MGMT_IOC_MAGIC, 0x0A, struct cmh_ioc= tl_kic_hkdf2) +#define CMH_IOCTL_KEY_NEW_RANDOM _IOWR(CMH_MGMT_IOC_MAGIC, 0x0B, struct cm= h_ioctl_key_new) +#define CMH_IOCTL_KIC_AES_CMAC_KDF _IOWR(CMH_MGMT_IOC_MAGIC, 0x0C, \ + struct cmh_ioctl_kic_aes_cmac_kdf) +#define CMH_IOCTL_KIC_DKEK_DERIVE _IOWR(CMH_MGMT_IOC_MAGIC, 0x0D, \ + struct cmh_ioctl_kic_dkek_derive) +#define CMH_IOCTL_KEY_LIST _IOWR(CMH_MGMT_IOC_MAGIC, 0x0E, struct cmh_ioct= l_key_list) + +/* PKE operation ioctls */ +#define CMH_IOCTL_PKE_RSA_ENC _IOWR(CMH_MGMT_IOC_MAGIC, 0x10, \ + struct cmh_ioctl_pke_rsa_enc) +#define CMH_IOCTL_PKE_RSA_DEC _IOWR(CMH_MGMT_IOC_MAGIC, 0x11, \ + struct cmh_ioctl_pke_rsa_dec) +#define CMH_IOCTL_PKE_RSA_CRT_DEC _IOWR(CMH_MGMT_IOC_MAGIC, 0x12, \ + struct cmh_ioctl_pke_rsa_crt_dec) +#define CMH_IOCTL_PKE_RSA_KEYGEN _IOWR(CMH_MGMT_IOC_MAGIC, 0x13, \ + struct cmh_ioctl_pke_rsa_keygen) +#define CMH_IOCTL_PKE_ECDSA_SIGN _IOWR(CMH_MGMT_IOC_MAGIC, 0x14, \ + struct cmh_ioctl_pke_ecdsa_sign) +#define CMH_IOCTL_PKE_ECDH _IOWR(CMH_MGMT_IOC_MAGIC, 0x16, \ + struct cmh_ioctl_pke_ecdh) +#define CMH_IOCTL_PKE_ECDH_KEYGEN _IOWR(CMH_MGMT_IOC_MAGIC, 0x17, \ + struct cmh_ioctl_pke_ecdh_keygen) +#define CMH_IOCTL_PKE_EDDSA_SIGN _IOWR(CMH_MGMT_IOC_MAGIC, 0x18, \ + struct cmh_ioctl_pke_eddsa_sign) +#define CMH_IOCTL_PKE_EDDSA_VERIFY _IOW(CMH_MGMT_IOC_MAGIC, 0x19, \ + struct cmh_ioctl_pke_eddsa_verify) +#define CMH_IOCTL_PKE_EC_KEYGEN _IOWR(CMH_MGMT_IOC_MAGIC, 0x1A, \ + struct cmh_ioctl_pke_ec_keygen) +#define CMH_IOCTL_PKE_EC_PUBGEN _IOWR(CMH_MGMT_IOC_MAGIC, 0x1B, \ + struct cmh_ioctl_pke_ec_pubgen) +#define CMH_IOCTL_PKE_EDDSA_KEYGEN_SCA _IOWR(CMH_MGMT_IOC_MAGIC, 0x1C, \ + struct cmh_ioctl_pke_eddsa_keygen_sca) + +/* -- PQC ioctl argument structures ----------- */ + +/* + * PQC operation flags (bits [2:0]). + * PQC keygen ioctls accept CMH_FLAG_PT in bits [18:16] to explicitly + * set the DS key storage attribute when CMH_QSE_FLAG_DS_REF is set. + * CMH_FLAG_SCA and CMH_FLAG_XC are rejected -- QSE SCA protection uses + * polynomial masking (CMH_QSE_FLAG_MASKED), not 2-share storage, + * and the eSW dec/sign paths hardcode SYS_TYPE_FLAG_PT. + * If no CMH_FLAG_* bits are set, DS keys default to CMH_FLAG_PT. + */ +#define CMH_QSE_FLAG_MASKED _BITUL(0) /* use masked (SCA-resistant) HW com= mands */ +#define CMH_QSE_FLAG_DS_REF _BITUL(1) /* store key output in DS, return re= f */ +#define CMH_QSE_FLAG_HW_RNG _BITUL(2) /* use HW RNG for seed/randomness */ +#define CMH_QSE_FLAG_MASK (_BITUL(0) | _BITUL(1) | _BITUL(2)) + +/* -- SYS wrap header size -------------------- */ +/* sys_read prepends a 16-byte header even for plaintext reads */ +#define CMH_SYS_WRAP_HDR_SIZE 16 + +/* -- Seed / randomness lengths --------------- */ + +#define CMH_QSE_SEED_LEN 32 /* ML-KEM/ML-DSA seed size */ +#define CMH_QSE_SEED_LEN_MASKED 64 /* seed size for masked mode */ + +/* -- ML-DSA ExternalMu sentinel -------------- */ +/* Pass this as mlen to use 64-byte pre-hashed mu instead of raw message */ +#define CMH_ML_DSA_MLEN_EXTERNAL_MU 0xFFFFFFFFU + +/* -- ML-KEM size macros ---------------------- */ + +#define CMH_ML_KEM_EK_SIZE(k) (384U * (k) + 32U) +#define CMH_ML_KEM_DK_SIZE(k) (768U * (k) + 96U) +/* CT sizes: k=3D2 -> 768, k=3D3 -> 1088, k=3D4 -> 1568 */ +#define CMH_ML_KEM_CT_SIZE_512 768U +#define CMH_ML_KEM_CT_SIZE_768 1088U +#define CMH_ML_KEM_CT_SIZE_1024 1568U +#define CMH_ML_KEM_SS_LEN 32U + +/* -- ML-DSA size macros ---------------------- */ +/* Indexed by mode: [0]=3D44 (mode=3D2), [1]=3D65 (mode=3D3), [2]=3D87 (mo= de=3D5) */ + +#define CMH_ML_DSA_44_PK_SIZE 1312U +#define CMH_ML_DSA_44_SK_SIZE 2560U +#define CMH_ML_DSA_44_SIG_SIZE 2420U +#define CMH_ML_DSA_65_PK_SIZE 1952U +#define CMH_ML_DSA_65_SK_SIZE 4032U +#define CMH_ML_DSA_65_SIG_SIZE 3309U +#define CMH_ML_DSA_87_PK_SIZE 2592U +#define CMH_ML_DSA_87_SK_SIZE 4896U +#define CMH_ML_DSA_87_SIG_SIZE 4627U + +/* -- SLH-DSA parameter set IDs --------------- */ + +#define CMH_SLHDSA_SHAKE_128S 1U +#define CMH_SLHDSA_SHAKE_128F 2U +#define CMH_SLHDSA_SHAKE_192S 3U +#define CMH_SLHDSA_SHAKE_192F 4U +#define CMH_SLHDSA_SHAKE_256S 5U +#define CMH_SLHDSA_SHAKE_256F 6U +#define CMH_SLHDSA_SHA2_128S 7U +#define CMH_SLHDSA_SHA2_128F 8U +#define CMH_SLHDSA_SHA2_192S 9U +#define CMH_SLHDSA_SHA2_192F 10U +#define CMH_SLHDSA_SHA2_256S 11U +#define CMH_SLHDSA_SHA2_256F 12U +#define CMH_SLHDSA_PARAM_MAX 12U + +/* SLH-DSA prehash algorithm IDs */ +#define CMH_SLHDSA_PREHASH_SHA256 1U +#define CMH_SLHDSA_PREHASH_SHA512 2U +#define CMH_SLHDSA_PREHASH_SHAKE128 3U +#define CMH_SLHDSA_PREHASH_SHAKE256 4U +#define CMH_SLHDSA_PREHASH_MAX 4U + +/* SLH-DSA n-value table indexed by (param_set - 1) */ +#define CMH_SLHDSA_N_128 16U +#define CMH_SLHDSA_N_192 24U +#define CMH_SLHDSA_N_256 32U + +/* SLH-DSA key sizes: pk =3D 2*n, sk =3D 4*n, seed =3D 3*n */ +#define CMH_SLHDSA_PK_SIZE(n) (2U * (n)) +#define CMH_SLHDSA_SK_SIZE(n) (4U * (n)) +#define CMH_SLHDSA_SEED_SIZE(n) (3U * (n)) + +/* SLH-DSA signature sizes indexed by (param_set - 1) */ +#define CMH_SLHDSA_SIG_SIZE_SHAKE_128S 7856U +#define CMH_SLHDSA_SIG_SIZE_SHAKE_128F 17088U +#define CMH_SLHDSA_SIG_SIZE_SHAKE_192S 16224U +#define CMH_SLHDSA_SIG_SIZE_SHAKE_192F 35664U +#define CMH_SLHDSA_SIG_SIZE_SHAKE_256S 29792U +#define CMH_SLHDSA_SIG_SIZE_SHAKE_256F 49856U +#define CMH_SLHDSA_SIG_SIZE_SHA2_128S 7856U +#define CMH_SLHDSA_SIG_SIZE_SHA2_128F 17088U +#define CMH_SLHDSA_SIG_SIZE_SHA2_192S 16224U +#define CMH_SLHDSA_SIG_SIZE_SHA2_192F 35664U +#define CMH_SLHDSA_SIG_SIZE_SHA2_256S 29792U +#define CMH_SLHDSA_SIG_SIZE_SHA2_256F 49856U + +/* -- PKE curve IDs -------------- */ + +#define CMH_CURVE_P192 0x01U +#define CMH_CURVE_P224 0x02U +#define CMH_CURVE_P256 0x03U +#define CMH_CURVE_P384 0x04U +#define CMH_CURVE_P521 0x05U +#define CMH_CURVE_SECP256K1 0x07U +#define CMH_CURVE_BP192R1 0x11U +#define CMH_CURVE_BP224R1 0x12U +#define CMH_CURVE_BP256R1 0x13U +#define CMH_CURVE_BP320R1 0x14U +#define CMH_CURVE_BP384R1 0x15U +#define CMH_CURVE_BP512R1 0x16U +#define CMH_CURVE_SM2 0x18U +#define CMH_CURVE_25519 0x21U +#define CMH_CURVE_448 0x22U + +/* ML-KEM */ + +struct cmh_ioctl_ml_kem_keygen { + __u32 version; + __u32 k; /* security parameter: 2/3/4 */ + __u32 flags; /* CMH_QSE_FLAG_* */ + __u32 __reserved; + __u64 seed; /* user-space pointer to seed (or 0 for HW RNG) */ + __u64 z; /* user-space pointer to z (or 0 for HW RNG) */ + __u64 ek; /* [out] user-space pointer to encapsulation key */ + __u64 dk; /* [out] user-space pointer to decapsulation key + * or [out] DS ref if CMH_QSE_FLAG_DS_REF + */ + __u64 dk_cid; /* CID for DS entry (if DS_REF) */ + __u64 dk_ref; /* [out] dk DS reference (if DS_REF) */ +}; + +struct cmh_ioctl_ml_kem_enc { + __u32 version; + __u32 k; + __u32 flags; /* CMH_QSE_FLAG_* */ + __u32 __reserved; + __u64 coin; /* user-space pointer to random coin (or 0) */ + __u64 ek; /* user-space pointer to encapsulation key */ + __u64 ct; /* [out] user-space pointer to ciphertext */ + __u64 ss; /* [out] user-space pointer to shared secret */ + __u64 __reserved2[2]; /* reserved for future use */ +}; + +struct cmh_ioctl_ml_kem_dec { + __u32 version; + __u32 k; + __u32 flags; /* CMH_QSE_FLAG_* */ + __u32 __reserved; + __u64 ct; /* user-space pointer to ciphertext */ + __u64 dk; /* user-space pointer to dk or DS ref */ + __u64 ss; /* [out] user-space pointer to shared secret */ + __u64 __reserved2[2]; /* reserved for future use */ +}; + +/* ML-DSA */ + +struct cmh_ioctl_ml_dsa_keygen { + __u32 version; + __u32 mode; /* security parameter: 2/3/5 */ + __u32 flags; /* CMH_QSE_FLAG_* */ + __u32 __reserved; + __u64 seed; /* user-space pointer to seed (or 0 for HW RNG) */ + __u64 pk; /* [out] user-space pointer to public key */ + __u64 sk; /* [out] user-space pointer to secret key + * or [out] DS ref if CMH_QSE_FLAG_DS_REF + */ + __u64 sk_cid; /* CID for DS entry (if DS_REF) */ + __u64 sk_ref; /* [out] sk DS reference (if DS_REF) */ +}; + +struct cmh_ioctl_ml_dsa_sign { + __u32 version; + __u32 mode; + __u32 flags; /* CMH_QSE_FLAG_* */ + __u32 mlen; /* message length in bytes */ + __u64 m; /* user-space pointer to message */ + __u64 sk; /* user-space pointer to sk or DS ref */ + __u64 sig; /* [out] user-space pointer to signature */ + __u64 rnd; /* user-space pointer to randomness (or 0) */ +}; + +/* SLH-DSA */ + +struct cmh_ioctl_slhdsa_keygen { + __u32 version; + __u32 parameter_set; /* HCQ_SLHDSA_SHAKE_128S .. SHA2_256F */ + __u32 flags; /* CMH_QSE_FLAG_DS_REF */ + __u32 __reserved; + __u64 seed; /* user-space pointer to seed */ + __u64 pk; /* [out] user-space pointer to public key */ + __u64 sk; /* [out] user-space pointer to secret key + * or [out] DS ref if CMH_QSE_FLAG_DS_REF + */ + __u64 sk_cid; /* CID for DS entry (if DS_REF) */ + __u64 sk_ref; /* [out] sk DS reference (if DS_REF) */ +}; + +struct cmh_ioctl_slhdsa_sign { + __u32 version; + __u32 parameter_set; + __u32 msg_len; + __u32 ctx_len; + __u64 msg; /* user-space pointer to message */ + __u64 ctx; /* user-space pointer to context (or 0) */ + __u64 sk; /* DS ref for secret key */ + __u64 sig; /* [out] user-space pointer to signature */ + __u64 add_random; /* user-space pointer to addl. randomness (or 0) */ +}; + +struct cmh_ioctl_slhdsa_sign_prehash { + __u32 version; + __u32 parameter_set; + __u32 prehash_algo; /* CMH_SLHDSA_PREHASH_* */ + __u32 digest; /* 0 =3D raw msg (eSW hashes), 1 =3D pre-computed digest */ + __u32 msg_len; + __u32 ctx_len; + __u64 msg; /* user-space pointer to message/digest */ + __u64 ctx; /* user-space pointer to context (or 0) */ + __u64 sk; /* DS ref for secret key */ + __u64 sig; /* [out] user-space pointer to signature */ + __u64 add_random; /* user-space pointer to addl. randomness (or 0) */ +}; + +/* -- SM2 ioctl argument structures ----------- */ + +/* SM2 fixed key sizes (sm2p256v1 curve, 256-bit) */ +#define CMH_SM2_CLEN 32U /* coordinate length */ +#define CMH_SM2_PUBKEY_LEN 64U /* uncompressed (x||y) */ +#define CMH_SM2_POINT_LEN 64U /* EC point (x||y) */ +#define CMH_SM2_SHARED_KEY_LEN 16U /* ECDH shared key */ +#define CMH_SM2_DIGEST_LEN 32U /* SM3 digest (ZA) */ +/* + * SM2 enc_hash/dec_hash payload limit. + * + * The eSW PKE driver implements only a single-block GM/T 0003.4 KDF + * (one SM3 invocation, 32 bytes of key stream). Longer messages would + * silently produce incorrect ciphertext / plaintext, so the driver caps + * the payload at 32 bytes. See Documentation/ABI/testing/cmh-mgmt. + */ +#define CMH_SM2_MAX_MSG_LEN 32U /* encrypt/decrypt */ +#define CMH_SM2_MAX_ID_LEN 32U /* identity string */ +#define CMH_SM2_CT_OVERHEAD 96U /* C1(64) + C3(32) */ +#define CMH_SM2_MAX_CT_LEN 128U /* 96 + max_msg =3D 128 */ + +struct cmh_ioctl_sm2_ecdh_keygen { + __u32 version; + __u32 nonce_len; /* 0 =3D HW generates r (written back), 32 =3D caller pr= ovides r */ + __u64 nonce; /* [in/out] user-space pointer to nonce buffer (32B) */ + __u64 session_key; /* [out] user-space pointer to session key R=3Dr*G (64= B) */ +}; + +struct cmh_ioctl_sm2_ecdh { + __u32 version; + __u32 nonce_len; /* 0 =3D HW generates (written back), 32 =3D caller prov= ides */ + __u64 nonce; /* [in/out] user-space pointer to nonce r (32B) */ + __u64 peer_public_key; /* user-space pointer to peer pub key (64B) */ + __u64 peer_session_key; /* user-space pointer to peer session key (64B) */ + __u64 key_ref; /* private key DS reference */ + __u64 shared_point; /* [out] user-space pointer to shared point (64B) */ + __u64 shared_point_ref; /* [in/out] 0 =3D read-back mode; &ref =3D keep D= S, write ref */ +}; + +struct cmh_ioctl_sm2_dec_point { + __u32 version; + __u32 ciphertext_len; /* total ciphertext length (97..128) */ + __u64 ciphertext; /* user-space pointer to ciphertext (64B: C1) */ + __u64 dec_point; /* [out] user-space pointer to dec point (64B) */ + __u64 key_ref; /* private key DS reference */ +}; + +struct cmh_ioctl_sm2_enc_point { + __u32 version; + __u32 nonce_len; /* 0 =3D HW generates, 32 =3D caller provides */ + __u64 nonce; /* user-space pointer to nonce (or 0) */ + __u64 public_key; /* user-space pointer to public key (64B) */ + __u64 ciphertext; /* [out] user-space pointer to C1 (64B) */ + __u64 enc_point; /* [out] user-space pointer to enc point (64B) */ +}; + +struct cmh_ioctl_sm2_id_digest { + __u32 version; + __u32 id_len; /* identity length in bytes (<=3D32) */ + __u64 id; /* user-space pointer to identity string */ + __u64 public_key; /* user-space pointer to public key (64B) */ + __u64 digest; /* [out] user-space pointer to ZA digest (32B) */ +}; + +/* + * SM2 ECDH_HASH -- derive shared key from shared point + ZA digests. + * + * IMPORTANT: The digest fields use ABSOLUTE ordering per GM/T 0003.3, + * NOT relative own/peer ordering. Both parties must pass: + * peer_id_digest =3D Z_A (initiator's digest) -- hashed FIRST + * id_digest =3D Z_B (responder's digest) -- hashed SECOND + * The eSW computes: KDF(shared_point || peer_id_digest || id_digest). + */ +struct cmh_ioctl_sm2_ecdh_hash { + __u32 version; + __u32 __reserved; + __u64 peer_id_digest; /* ptr to Z_A -- initiator's digest (32B) */ + __u64 id_digest; /* ptr to Z_B -- responder's digest (32B) */ + __u64 shared_point_ref; /* DS reference from SM2_ECDH */ + __u64 shared_key; /* [out] ptr to shared key (16B) */ +}; + +struct cmh_ioctl_sm2_dec_hash { + __u32 version; + __u32 ciphertext_len; /* ciphertext length (97..128) */ + __u64 ciphertext; /* user-space pointer to full ciphertext */ + __u64 dec_point; /* user-space pointer to dec point (64B) */ + __u64 plaintext; /* [out] user-space pointer to plaintext */ +}; + +struct cmh_ioctl_sm2_enc_hash { + __u32 version; + __u32 message_len; /* message length (1..32) */ + __u64 message; /* user-space pointer to plaintext */ + __u64 enc_point; /* user-space pointer to enc point (64B) */ + __u64 ciphertext; /* [out] user-space pointer to ciphertext */ +}; + +/* PQC ioctl numbers */ +#define CMH_IOCTL_ML_KEM_KEYGEN _IOWR(CMH_MGMT_IOC_MAGIC, 0x20, \ + struct cmh_ioctl_ml_kem_keygen) +#define CMH_IOCTL_ML_KEM_ENC _IOWR(CMH_MGMT_IOC_MAGIC, 0x21, \ + struct cmh_ioctl_ml_kem_enc) +#define CMH_IOCTL_ML_KEM_DEC _IOWR(CMH_MGMT_IOC_MAGIC, 0x22, \ + struct cmh_ioctl_ml_kem_dec) +#define CMH_IOCTL_ML_DSA_KEYGEN _IOWR(CMH_MGMT_IOC_MAGIC, 0x23, \ + struct cmh_ioctl_ml_dsa_keygen) +#define CMH_IOCTL_ML_DSA_SIGN _IOWR(CMH_MGMT_IOC_MAGIC, 0x24, \ + struct cmh_ioctl_ml_dsa_sign) +#define CMH_IOCTL_SLHDSA_KEYGEN _IOWR(CMH_MGMT_IOC_MAGIC, 0x28, \ + struct cmh_ioctl_slhdsa_keygen) +#define CMH_IOCTL_SLHDSA_SIGN _IOWR(CMH_MGMT_IOC_MAGIC, 0x29, \ + struct cmh_ioctl_slhdsa_sign) +#define CMH_IOCTL_SLHDSA_SIGN_PREHASH _IOWR(CMH_MGMT_IOC_MAGIC, 0x2D, \ + struct cmh_ioctl_slhdsa_sign_prehash) + +/* SM2 operation ioctls */ +#define CMH_IOCTL_SM2_ECDH_KEYGEN _IOWR(CMH_MGMT_IOC_MAGIC, 0x30, \ + struct cmh_ioctl_sm2_ecdh_keygen) +#define CMH_IOCTL_SM2_ECDH _IOWR(CMH_MGMT_IOC_MAGIC, 0x31, \ + struct cmh_ioctl_sm2_ecdh) +#define CMH_IOCTL_SM2_DEC_POINT _IOWR(CMH_MGMT_IOC_MAGIC, 0x32, \ + struct cmh_ioctl_sm2_dec_point) +#define CMH_IOCTL_SM2_ENC_POINT _IOWR(CMH_MGMT_IOC_MAGIC, 0x33, \ + struct cmh_ioctl_sm2_enc_point) +#define CMH_IOCTL_SM2_ID_DIGEST _IOWR(CMH_MGMT_IOC_MAGIC, 0x34, \ + struct cmh_ioctl_sm2_id_digest) +#define CMH_IOCTL_SM2_ECDH_HASH _IOWR(CMH_MGMT_IOC_MAGIC, 0x35, \ + struct cmh_ioctl_sm2_ecdh_hash) +#define CMH_IOCTL_SM2_DEC_HASH _IOWR(CMH_MGMT_IOC_MAGIC, 0x36, \ + struct cmh_ioctl_sm2_dec_hash) +#define CMH_IOCTL_SM2_ENC_HASH _IOWR(CMH_MGMT_IOC_MAGIC, 0x37, \ + struct cmh_ioctl_sm2_enc_hash) + +/* + * EAC (Error and Alarm Controller) -- read and clear error registers. + * + * Returns a snapshot of all hardware error/safety/notification registers. + * The eSW atomically reads and clears the registers on each call, so + * successive reads show only new events. + */ +struct cmh_ioctl_eac_read { + __u32 version; /* must be CMH_MGMT_V1 */ + __u32 __reserved; + __u64 mailbox_notification; /* [out] MBX safety notification bitmask */ + __u32 hw_error; /* [out] HWC error bitmask */ + __u32 hw_nmi; /* [out] HWC NMI bitmask */ + __u32 hw_panic; /* [out] HWC panic bitmask */ + __u32 safety_fatal; /* [out] HWC fatal safety bitmask */ + __u32 safety_notification; /* [out] HWC safety notification bitmask */ + __u32 sw_info0; /* [out] eSW tracing info */ + __u32 sw_info1; /* [out] eSW tracing info */ + __u32 sram_bank_errors[4]; /* [out] correctable ECC error counts */ + __u32 __pad; /* explicit tail padding (prevent info leak) */ +}; + +/* + * DRBG CONFIG -- configure the hardware DRBG before first use. + * + * This is a management operation normally performed once at system + * start-up. Must be called before any hwrng reads or DRBG GENERATE + * operations. + */ +#define CMH_DRBG_RATIO_ONE 0 /* 1:1 entropy ratio */ +#define CMH_DRBG_RATIO_ONE_HALF 1 /* 1:2 */ +#define CMH_DRBG_RATIO_ONE_THIRD 2 /* 1:3 */ +#define CMH_DRBG_RATIO_ONE_FOURTH 3 /* 1:4 */ + +#define CMH_DRBG_STRENGTH_128 0x00 /* 128-bit security */ +#define CMH_DRBG_STRENGTH_256 0x10 /* 256-bit security */ + +struct cmh_ioctl_drbg_config { + __u32 version; /* must be CMH_MGMT_V1 */ + __u32 entropy_ratio; /* CMH_DRBG_RATIO_* */ + __u32 security_strength; /* CMH_DRBG_STRENGTH_* */ + __u32 __reserved; +}; + +/* EAC ioctl number */ +#define CMH_IOCTL_EAC_READ _IOWR(CMH_MGMT_IOC_MAGIC, 0x0F, \ + struct cmh_ioctl_eac_read) + +/* DRBG management ioctl number */ +#define CMH_IOCTL_DRBG_CONFIG _IOW(CMH_MGMT_IOC_MAGIC, 0x40, \ + struct cmh_ioctl_drbg_config) + +#endif /* _UAPI_CMH_MGMT_IOCTL_H */ --=20 2.43.7 From nobody Fri Sep 25 01:20:33 2026 Received: from CH4PR04CU002.outbound.protection.outlook.com (mail-northcentralusazon11023137.outbound.protection.outlook.com [40.107.201.137]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0E0774AAC5F; Thu, 17 Sep 2026 22:59:39 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=40.107.201.137 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789685983; cv=fail; b=K1v6CuMzaI9oOCxskBkFQo7xjHLHdOdHthkkM5IvMmfkpqhBI78QEoWdMeoUsmBtdCMpNG+UIsQwyVciGtwUZtc0VPavzCaeg16R1vNeGGys3SYYWv8fNOrCY08NJigLgrGhfvgw9gAIe2Q+9lf/76OsAZaJWz4ronkUxFjjWDI= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789685983; c=relaxed/simple; bh=/2vc/aMzPLXvGLoyijGjCfuQ/8pqUiLVI/E3PUfvCKY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=r5TFNnblsKXm/88bNHYdToxGWKTyrK11urPHj4xHxiuuskV9EVf942xJV7nQksiFdxtJstivSUWwCpb1/9PlorHJnKb7Od+oOaXROhA0qPIgGL0Sl7brjieu9A4ooE9SyZVr0+isO8GkLdMIr0WBDvYJAhL4Q3DRGL2FKFpsG74= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com; spf=fail smtp.mailfrom=rambus.com; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b=P9Bvp1eU; arc=fail smtp.client-ip=40.107.201.137 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=rambus.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b="P9Bvp1eU" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=vCn8x3QONbKPMjHMl5AYmcryHTJ+jP+NOwAnD5OrjB/O4xBeYqxYbJEJuJdqjqwtmJnMxU37vVbd/9Q4Vd+fq36WLHGkOL7l0pk8jTB0MyQyRQNdSgjwqn+xC/G0mR0Tf11q+YmnvK3wT+ZphHM7APlj6v94N68aEDJ2F3bLD8I/uT+Hle/CcsHlYs4ib+9cbR4DQ9uWVcA0FslPuamyVv8VHDJtWw1Jw9aL/GAG4V3EV1k7ze41R7orzbdzfnXW2U7KtkwCTgCzCd47TS/7PZLyS/TbZd7cLCz1tlXL16oDsACujcZfTtLOYvkqkDQut4T0L0wm6pngyYXtxazeEQ== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=bpkTV3CFzIOR2jA1eCCbMwkpyRFDDOnScx8EWM7JARQ=; b=CO9nJanEdDFAta/eJge7gZNwR2w82IDfAFZz7MOWOGh7TAkTE1JEmBRCaEDEMevxwwY+ofsFmF/2ClyQ8aRJKJdDpwcarmnnfpZSXbt43ZK+zXJgi/QkmVw/7PUMIHkvZ26+mOC5fuafXPZv1SffzJvery/uV8W4TOYtw+qI3b7CIza1O9BrEO980e2AKmogmejORPro0/hyAN2CW5p6hFlO8TDJTmmcZKMSWdWvZbJk1D0cXocvbHlTh41iF4Z9iTHSlgnKyF3dpmZT+hgdVyoXN8+P5dq8DHTWJ0ZyVNNUDIZsOoS1pAUhFRq8hY8eLLcR5MINx3lt4Nfb1IR3Ag== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 192.86.86.210) smtp.rcpttodomain=cryptography.com smtp.mailfrom=rambus.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=rambus.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=rambus.com; s=selector1; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=bpkTV3CFzIOR2jA1eCCbMwkpyRFDDOnScx8EWM7JARQ=; b=P9Bvp1eURcRsFYBj6pT5aKBTSb+lMw+yKdzI3/XV9EGuQDHxfDOCnTF1dcARwkPypaz2MCiSexkggvJvYBKsZL3dpkxQriJTnaQYTd32yrryBcxoGs9jBKcb2LIMn4iKi9fjFTQe7rQi9HFS85EgHBVKIDkky//Z9qCN7of3ChroYZt4GHfsgdLfyk9BsN4MXZYTJQD/SIkOn8QCIKq9a5k6Ci9rNAiVewG2guR0ERQRYf8HF1pm3XRZdtpkXfPzNiaWxvTgAWHvEdQ/ZDfM60Zppl7B+zv74TZDjWl2ozVnbzJgv4QHPirrzJM3hLcMRkllVHibLMYBXHa5v9+34A== Received: from BN0PR07CA0022.namprd07.prod.outlook.com (2603:10b6:408:141::14) by DS2PR04MB994244.namprd04.prod.outlook.com (2603:10b6:8:4a8::21) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.428.13; Thu, 17 Sep 2026 22:59:33 +0000 Received: from BN2PEPF0000A7FE.namprd02.prod.outlook.com (2603:10b6:408:141:cafe::90) by BN0PR07CA0022.outlook.office365.com (2603:10b6:408:141::14) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.428.12 via Frontend Transport; Thu, 17 Sep 2026 22:59:32 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 192.86.86.210) smtp.mailfrom=rambus.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=rambus.com; Received-SPF: Pass (protection.outlook.com: domain of rambus.com designates 192.86.86.210 as permitted sender) receiver=protection.outlook.com; client-ip=192.86.86.210; helo=hqxsv-psmtppxy02.rambus.com; pr=C Received: from hqxsv-psmtppxy02.rambus.com (192.86.86.210) by BN2PEPF0000A7FE.mail.protection.outlook.com (10.167.245.165) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.451.8 via Frontend Transport; Thu, 17 Sep 2026 22:59:32 +0000 Received: from hqxsv-cmdev3-aousherovitch.rambus.com (hqn-lb-int-float.rambus.com [10.12.20.20]) by hqxsv-psmtppxy02.rambus.com (Postfix) with ESMTP id E9BE5180175F; Thu, 17 Sep 2026 22:59:31 +0000 (UTC) From: Alex Ousherovitch To: Albert Ou , Alex Ousherovitch , Conor Dooley , "David S. Miller" , Herbert Xu , Jonathan Corbet , Krzysztof Kozlowski , Palmer Dabbelt , Paul Walmsley , Rob Herring , Saravanakrishnan Krishnamoorthy , Shuah Khan Cc: Alexandre Ghiti , devicetree@vger.kernel.org, Joel Wittenauer , linux-api@vger.kernel.org, linux-crypto@vger.kernel.org, linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-riscv@lists.infradead.org, Shuah Khan , Thi Nguyen Subject: [PATCH v5 04/19] crypto: cmh - add SHA-2/SHA-3/SHAKE ahash Date: Thu, 17 Sep 2026 15:59:13 -0700 Message-ID: <20260917225929.2494111-5-aousherovitch@rambus.com> X-Mailer: git-send-email 2.43.7 In-Reply-To: <20260917225929.2494111-1-aousherovitch@rambus.com> References: <20260917225929.2494111-1-aousherovitch@rambus.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: BN2PEPF0000A7FE:EE_|DS2PR04MB994244:EE_ X-MS-Office365-Filtering-Correlation-Id: dce906b5-89a8-424d-4841-08df150f5665 X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|376014|7416014|1800799024|82310400026|36860700016|23010399003|921020|3023799007|56012099006|11063799006|18002099003|22082099003|10067099003; X-Microsoft-Antispam-Message-Info: 8sYkFs8qxfcJ95iqNaQXfxcUopT/E1d99J0RWuLt3A9AJ/VoOkEHnpnYUqMCrskdRgPquajV3KO77Qef6/ckiTs0hGIIHde5DIpzuU6i45PaqIEkRQhlqTzhSbtq/pDL/bgHrNr1ssI9Egg0+S9HybqfrbOm4mLMoiij6ZSD0jO8XeRwMoWyZVcrfv0eD+rkba3F/zgSaj3sLU4K86gAukUDKOEcrUr0pjtMGNFCnqMvdUXFcyHprwAwFQg1yLtJc+wr1qk+td9QyTsAfYgtnJavMMEq2V6YEYDAxb/sxoB8Tzm0S6Nxxwbo4by1ZEhDmmTPHlEhH6m5+b1onbREm1DZKQjHtTcHDEarLPi2jYmVlpDN6xF9LK6YY80gOkUJdVYZzU99X72MPjRUPkU4UQPwADvbLDFDvvnxPb7AyLaNtZWMwYmrtBiBXTAffi4bhrw/1lVkJ2uKfuKU+8HxG4bj/8Xam6dwZS6dVcIZyevjXow8R75zy24vpjKMVb9wOIX6Np97Yq+3a7rp5s3qVKEuc25mgJiCYTbp1frsUPbU3a+8ZrXPiRarV+poyNWiNAj74uCyedA7SKLkCzBAOVg9l0y8pTYlVY/6xOgyCyWrwHWadcwObG/krXAanR1s68kXKJYiT5FXn3+kyh4vpJzx5L8r6pi27HnKXuvY6NU2gflk5MQ/FcexjeV/hT6GUwQC1uvlQm3glSXsvccYmzbOvZ2hfIoXlYS5O9jKkipSRsPpvaPF9LMPkyV1L/lb X-Forefront-Antispam-Report: CIP:192.86.86.210;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:hqxsv-psmtppxy02.rambus.com;PTR:InfoDomainNonexistent;CAT:NONE;SFS:(13230040)(376014)(7416014)(1800799024)(82310400026)(36860700016)(23010399003)(921020)(3023799007)(56012099006)(11063799006)(18002099003)(22082099003)(10067099003);DIR:OUT;SFP:1102; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: CGQhIHeicCYLDR5el7skCN2udspGLptV2/vvHf+QfbYIhhdeZOLphqDfG7mLQHiWe1I4pX1qw31zZzxA2MuR1b1SBNTVUMWAK3NK0PD11hkzLZQ5ZUjZPZs/8i0q6WakyupT0vb0NC2btFAFI7hIp3GziYtsP25USzPl54AprGauAbm2dWdQePev6DBtr1GEnQXCMPg5r5TtzNTIkt5hBuAesBSyvVkc7jXBq9qHx8QXobTpBqa9UB3+AJukCqO21gzxzfz9XQTtrl5jJ230lzwGF2t5i1Fj5CgKXjYqNv9rgX3bCE+2aINmYBuUd67vSjcYoU+ZmGfZb8PjKsKe/mLnncK3sjn21jrkhD5qMDU3CrUTi1A/C1sLjSwlw4+DH25Rm49IqUd7ZYmBp5tXZREi3q1nxaBnnKV7vRQGVgG7wsXYmtOupCWcQM8krewk X-OriginatorOrg: rambus.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 17 Sep 2026 22:59:32.5765 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: dce906b5-89a8-424d-4841-08df150f5665 X-MS-Exchange-CrossTenant-Id: bd0ba799-c2b9-413c-9c56-5d1731c4827c X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=bd0ba799-c2b9-413c-9c56-5d1731c4827c;Ip=[192.86.86.210];Helo=[hqxsv-psmtppxy02.rambus.com] X-MS-Exchange-CrossTenant-AuthSource: BN2PEPF0000A7FE.namprd02.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: DS2PR04MB994244 Content-Type: text/plain; charset="utf-8" Register ahash algorithms for SHA-224, SHA-256, SHA-384, SHA-512, SHA3-224, SHA3-256, SHA3-384, SHA3-512, SHAKE128, and SHAKE256 using the CMH hash core (core ID 0x02). The drivers set CRYPTO_AHASH_ALG_BLOCK_ONLY so the Crypto API buffers partial blocks; they support incremental update/finup and export/import for request cloning. Signed-off-by: Alex Ousherovitch Co-developed-by: Saravanakrishnan Krishnamoorthy Signed-off-by: Saravanakrishnan Krishnamoorthy --- drivers/crypto/cmh/Makefile | 3 +- drivers/crypto/cmh/cmh_hash.c | 812 ++++++++++++++++++++++++++ drivers/crypto/cmh/cmh_main.c | 9 + drivers/crypto/cmh/include/cmh_hash.h | 27 + 4 files changed, 850 insertions(+), 1 deletion(-) create mode 100644 drivers/crypto/cmh/cmh_hash.c create mode 100644 drivers/crypto/cmh/include/cmh_hash.h diff --git a/drivers/crypto/cmh/Makefile b/drivers/crypto/cmh/Makefile index bb7772555d6a..79c94d87e9ee 100644 --- a/drivers/crypto/cmh/Makefile +++ b/drivers/crypto/cmh/Makefile @@ -14,7 +14,8 @@ cmh-y :=3D \ cmh_dma.o \ cmh_sysfs.o \ cmh_key.o \ - cmh_sys.o + cmh_sys.o \ + cmh_hash.o =20 # Management ioctl device (/dev/cmh_mgmt): key lifecycle, PKE, PQC ioctls. cmh-$(CONFIG_CRYPTO_DEV_CMH_MGMT) +=3D \ diff --git a/drivers/crypto/cmh/cmh_hash.c b/drivers/crypto/cmh/cmh_hash.c new file mode 100644 index 000000000000..1acc8bdab352 --- /dev/null +++ b/drivers/crypto/cmh/cmh_hash.c @@ -0,0 +1,812 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Kernel Crypto API Hash Driver + * + * Registers asynchronous hash (ahash) algorithms with the Linux crypto + * subsystem. Implements SHA-2 (224/256/384/512), SHA-3 + * (224/256/384/512), and SHAKE (128/256) families using the CMH Hash + * Core (HC). + * + * Incremental HW update model. The driver sets + * CRYPTO_AHASH_ALG_BLOCK_ONLY, so the Crypto API buffers partial + * blocks: .update() is only handed whole-block-aligned data and returns + * any sub-block remainder for the API to hold until the next call. + * + * .init() -> software-only: zero per-request context + * .update() -> INIT [+ RESTORE] + UPDATE(full blocks) + SAVE + FLUSH; + * return -EINPROGRESS, completing with the leftover byte + * count the API must buffer (0 when block-aligned) + * .finup() -> INIT [+ RESTORE] [+ UPDATE(residual)] + FINAL + FLUSH + * (also serves .final via the API: nbytes =3D=3D 0) + * .digest() -> INIT + UPDATE + FINAL + FLUSH (single-shot via finup) + * .export()/.import() -> software-only: copy the HC context + * checkpoint; the API appends its own partial-block buffer + * + * The FLUSH after each .update() releases the HC core, so no lockout. + * Two hash sessions interleave fine on the same MBX -- each saves its + * own state via SAVE and restores via RESTORE on the next call. + * + * These are sg-only drivers (no CRYPTO_ALG_REQ_VIRT): BLOCK_ONLY buffer + * prepending assumes scatterlists. Export/import carry only HW state, + * enabling crypto API transform clone for all plain-hash algorithms. + */ + +#include +#include +#include +#include +#include +#include +#include +#include + +#include "cmh_hash.h" +#include "cmh_vcq.h" +#include "cmh_txn.h" +#include "cmh_dma.h" + +/* Algorithm Table */ + +struct cmh_hash_alg_info { + u32 hc_algo; /* HC_ALGO_* (SHA2, SHA3, SHAKE) */ + u32 digest_size; /* bytes */ + u32 block_size; /* cra_blocksize for Linux crypto API */ + const char *alg_name; /* Linux crypto name: "sha256" */ + const char *drv_name; /* driver name: "rambus-cmh-sha256" */ +}; + +static const struct cmh_hash_alg_info cmh_hash_algs_info[] =3D { + /* SHA-2 family */ + { + .hc_algo =3D HC_ALGO_SHA2_224, + .digest_size =3D CMH_SHA224_DIGEST_SIZE, + .block_size =3D 64, + .alg_name =3D "sha224", + .drv_name =3D "rambus-cmh-sha224", + }, + { + .hc_algo =3D HC_ALGO_SHA2_256, + .digest_size =3D CMH_SHA256_DIGEST_SIZE, + .block_size =3D 64, + .alg_name =3D "sha256", + .drv_name =3D "rambus-cmh-sha256", + }, + { + .hc_algo =3D HC_ALGO_SHA2_384, + .digest_size =3D CMH_SHA384_DIGEST_SIZE, + .block_size =3D 128, + .alg_name =3D "sha384", + .drv_name =3D "rambus-cmh-sha384", + }, + { + .hc_algo =3D HC_ALGO_SHA2_512, + .digest_size =3D CMH_SHA512_DIGEST_SIZE, + .block_size =3D 128, + .alg_name =3D "sha512", + .drv_name =3D "rambus-cmh-sha512", + }, + /* SHA-3 family */ + { + .hc_algo =3D HC_ALGO_SHA3_224, + .digest_size =3D CMH_SHA3_224_DIGEST_SIZE, + .block_size =3D 144, /* rate =3D 1600/8 - 2*224/8 =3D 144 */ + .alg_name =3D "sha3-224", + .drv_name =3D "rambus-cmh-sha3-224", + }, + { + .hc_algo =3D HC_ALGO_SHA3_256, + .digest_size =3D CMH_SHA3_256_DIGEST_SIZE, + .block_size =3D 136, /* rate =3D 1600/8 - 2*256/8 =3D 136 */ + .alg_name =3D "sha3-256", + .drv_name =3D "rambus-cmh-sha3-256", + }, + { + .hc_algo =3D HC_ALGO_SHA3_384, + .digest_size =3D CMH_SHA3_384_DIGEST_SIZE, + .block_size =3D 104, /* rate =3D 1600/8 - 2*384/8 =3D 104 */ + .alg_name =3D "sha3-384", + .drv_name =3D "rambus-cmh-sha3-384", + }, + { + .hc_algo =3D HC_ALGO_SHA3_512, + .digest_size =3D CMH_SHA3_512_DIGEST_SIZE, + .block_size =3D 72, /* rate =3D 1600/8 - 2*512/8 =3D 72 */ + .alg_name =3D "sha3-512", + .drv_name =3D "rambus-cmh-sha3-512", + }, + /* + * SHAKE (XOF) family -- fixed-output ahash registration. + * + * cra_blocksize =3D 1: SHAKE is a sponge/XOF, not Merkle-Damgaard. + * With BLOCK_ONLY this means the API never holds anything back and + * every byte reaches .update(); the HC core absorbs any sub-rate + * remainder into its saved context across SAVE/RESTORE, so unaligned + * intermediate absorbs are fine. + */ + { + .hc_algo =3D HC_ALGO_SHAKE128, + .digest_size =3D CMH_SHAKE128_DIGEST_SIZE, + .block_size =3D 1, /* XOF: no meaningful block for crypto API */ + .alg_name =3D "shake128", + .drv_name =3D "rambus-cmh-shake128", + }, + { + .hc_algo =3D HC_ALGO_SHAKE256, + .digest_size =3D CMH_SHAKE256_DIGEST_SIZE, + .block_size =3D 1, /* XOF: no meaningful block for crypto API */ + .alg_name =3D "shake256", + .drv_name =3D "rambus-cmh-shake256", + }, +}; + +#define CMH_HASH_ALG_COUNT ARRAY_SIZE(cmh_hash_algs_info) + +/* Per-Request State */ + +/* + * Exported hash state -- serialised by .export(), deserialised by + * .import(). This is what statesize advertises to the crypto + * subsystem; the API appends its own partial-block buffer on top. + */ +struct cmh_hash_export_state { + u8 checkpoint[HC_CONTEXT_SIZE]; /* HC context from last SAVE */ + u32 hw_started; /* non-zero if checkpoint valid */ +}; + +/* + * Maximum payload commands any hash transaction can produce: + * INIT + RESTORE + UPDATE + SAVE/FINAL + FLUSH =3D 5 + * Worst-case packed output (stride=3D7, 1 payload per VCQ): + * 5 VCQs x 2 entries =3D 10 + */ +#define CMH_HASH_MAX_PAYLOAD 5 +#define CMH_HASH_MAX_PACKED (CMH_HASH_MAX_PAYLOAD * 2) + +/* + * Stored in ahash_request_ctx(). Tracks the algorithm, an HC context + * checkpoint from the last SAVE, and DMA state for the current in-flight + * async operation. Partial-block buffering is handled by the Crypto API + * (CRYPTO_AHASH_ALG_BLOCK_ONLY), not here. + * + * The checkpoint is embedded inline rather than heap-allocated because + * the kernel ahash API has no per-request destructor. If a request is + * abandoned without a final op (e.g. transform freed early), a heap + * checkpoint would leak unconditionally. + * + * Mapping the inline checkpoint for DMA is safe even on non-coherent + * platforms: it is only ever mapped DMA_TO_DEVICE, and its bytes are + * frozen from dma_map_single() until the matching unmap (it is written + * solely in the completion, after the unmap). The device therefore + * always reads the value written back at map time, regardless of CPU + * writes to adjacent fields (packed[]) sharing the same cacheline -- + * those writes never alter the checkpoint bytes, and a TO_DEVICE unmap + * performs no cache invalidate. The shared-cacheline hazard applies + * only to FROM_DEVICE / BIDIRECTIONAL mappings, and all such buffers + * here (save_buf, digest_buf) are separately kmalloc'd. + */ +struct cmh_hash_reqctx { + const struct cmh_hash_alg_info *info; + int error; + u32 hw_started; /* non-zero after first HW submission */ + u32 has_checkpoint; /* non-zero if checkpoint[] valid */ + u32 update_remainder; /* sub-block bytes the API must re-buffer */ + /* DMA state for current async operation */ + dma_addr_t ckpt_dma; /* RESTORE input */ + dma_addr_t save_dma; /* SAVE output (update only) */ + dma_addr_t data_dma; /* UPDATE input */ + dma_addr_t digest_dma; /* FINAL output (final/digest only) */ + u8 *save_buf; /* SAVE output buffer */ + u8 *data_buf; /* linearised data for DMA */ + u32 data_len; /* bytes in data_buf */ + u8 *digest_buf; /* digest output buffer */ + u8 checkpoint[HC_CONTEXT_SIZE]; /* HC context from last SAVE */ + struct vcq_cmd packed[CMH_HASH_MAX_PACKED]; +}; + +/* VCQ Builders (HC-specific; shared builders in cmh_hc_abi.h / cmh_vcq.h)= */ + +/* Add an HC_CMD_UPDATE entry */ +static void vcq_add_hc_update(struct vcq_cmd *slot, u32 core_id, u64 input= _phys, u32 len) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, HC_CMD_UPDATE); + slot->hwc.hc.cmd_update.input =3D input_phys; + slot->hwc.hc.cmd_update.inlen =3D len; +} + +/* Add an HC_CMD_SAVE entry */ +static void vcq_add_hc_save(struct vcq_cmd *slot, u32 core_id, u64 output_= phys, u32 outlen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, HC_CMD_SAVE); + slot->hwc.hc.cmd_save.output =3D output_phys; + slot->hwc.hc.cmd_save.outlen =3D outlen; +} + +/* Add an HC_CMD_RESTORE entry */ +static void vcq_add_hc_restore(struct vcq_cmd *slot, u32 core_id, u64 inpu= t_phys, u32 inlen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, HC_CMD_RESTORE); + slot->hwc.hc.cmd_restore.input =3D input_phys; + slot->hwc.hc.cmd_restore.inlen =3D inlen; +} + +/* Request Context Cleanup */ + +static void cmh_hash_free_reqctx(struct cmh_hash_reqctx *rctx) +{ + rctx->has_checkpoint =3D 0; +} + +/* VCQ Packing + Submit */ + +/* ahash Operations */ + +/* + * Wrapper struct: embeds ahash_alg + a pointer to our alg_info table + * entry so we can recover it in the tfm callbacks. + */ +struct cmh_hash_alg_drv { + struct ahash_alg alg; + const struct cmh_hash_alg_info *info; +}; + +/* + * Find the cmh_hash_alg_info from the crypto_ahash (embedded in our + * registered template). We stash the info pointer in the algorithm's + * driver-private area at registration time (see cmh_hash_register). + */ +static const struct cmh_hash_alg_info * +cmh_hash_get_info(struct crypto_ahash *tfm) +{ + struct ahash_alg *alg =3D crypto_ahash_alg(tfm); + + return container_of(alg, struct cmh_hash_alg_drv, alg)->info; +} + +static int cmh_hash_init(struct ahash_request *req) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_hash_reqctx *rctx =3D ahash_request_ctx(req); + + memset(rctx, 0, sizeof(*rctx)); + rctx->info =3D cmh_hash_get_info(tfm); + return 0; +} + +/* + * Update completion -- called from threaded IRQ after SAVE completes. + * Takes ownership of save_buf as the new checkpoint. + */ +static void cmh_hash_update_complete(void *data, int error) +{ + struct ahash_request *req =3D data; + struct cmh_hash_reqctx *rctx =3D ahash_request_ctx(req); + + if (error =3D=3D -EINPROGRESS) { + cmh_complete(&req->base, error); + return; + } + + /* Unmap DMA buffers */ + if (rctx->has_checkpoint) + cmh_dma_unmap_single(rctx->ckpt_dma, HC_CONTEXT_SIZE, + DMA_TO_DEVICE); + cmh_dma_unmap_single(rctx->save_dma, HC_CONTEXT_SIZE, + DMA_FROM_DEVICE); + cmh_dma_unmap_single(rctx->data_dma, rctx->data_len, + DMA_TO_DEVICE); + + if (!error) { + memcpy(rctx->checkpoint, rctx->save_buf, HC_CONTEXT_SIZE); + rctx->has_checkpoint =3D 1; + kfree(rctx->save_buf); + rctx->save_buf =3D NULL; + rctx->hw_started =3D 1; + /* Hand the API the sub-block remainder it must re-buffer. */ + error =3D rctx->update_remainder; + } else { + kfree(rctx->save_buf); + rctx->save_buf =3D NULL; + rctx->error =3D error; + } + + kfree(rctx->data_buf); + rctx->data_buf =3D NULL; + rctx->data_len =3D 0; + + cmh_complete(&req->base, error); +} + +/* + * .update -- submit whole blocks to HW. + * + * With CRYPTO_AHASH_ALG_BLOCK_ONLY the Crypto API prepends any bytes it + * held back from earlier calls, so req->src carries at least one full + * block. We hash the block-aligned prefix as: + * INIT [+ RESTORE] + UPDATE(full blocks) + SAVE + FLUSH + * and return the sub-block remainder for the API to re-buffer (reported + * through the async completion). For XOFs (block_size 1) there is never + * a remainder; the HC core absorbs sub-rate data across SAVE/RESTORE. + */ +static int cmh_hash_update(struct ahash_request *req) +{ + struct cmh_hash_reqctx *rctx =3D ahash_request_ctx(req); + const struct cmh_hash_alg_info *info =3D rctx->info; + struct vcq_cmd cmds[CMH_HASH_MAX_PAYLOAD]; + struct core_dispatch d; + u32 block_size =3D info->block_size; + u32 full_len; + u32 idx; + int ret; + gfp_t gfp; + + if (rctx->error) + return rctx->error; + + if (!req->nbytes) + return 0; + + /* + * block_size is not always a power of two (SHA-3 rates: 144/136/ + * 104/72), so use modulo -- round_down() would corrupt the split. + */ + rctx->update_remainder =3D req->nbytes % block_size; + full_len =3D req->nbytes - rctx->update_remainder; + + gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + + /* + * Reject a single update whose linearisation would exceed the largest + * kmalloc: return a permanent -EMSGSIZE ("message too long") rather + * than a transient -ENOMEM the client would keep retrying. + */ + if (full_len > KMALLOC_MAX_SIZE) + return -EMSGSIZE; + + /* + * Linearise the block-aligned prefix from the scatterlist. + * __GFP_NOWARN keeps a borderline-large (but sub-cap) request from + * splatting the page allocator if it still cannot be satisfied. + */ + rctx->data_buf =3D kmalloc(full_len, gfp | __GFP_NOWARN); + if (!rctx->data_buf) + return -ENOMEM; + + scatterwalk_map_and_copy(rctx->data_buf, req->src, 0, full_len, 0); + + /* Allocate SAVE output buffer */ + rctx->save_buf =3D kzalloc(HC_CONTEXT_SIZE, gfp); + if (!rctx->save_buf) { + ret =3D -ENOMEM; + goto err_free; + } + + /* DMA map data, save output, and checkpoint */ + rctx->data_dma =3D cmh_dma_map_single(rctx->data_buf, full_len, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->data_dma)) { + ret =3D -ENOMEM; + goto err_free; + } + + rctx->save_dma =3D cmh_dma_map_single(rctx->save_buf, HC_CONTEXT_SIZE, + DMA_FROM_DEVICE); + if (cmh_dma_map_error(rctx->save_dma)) { + ret =3D -ENOMEM; + goto err_unmap_data; + } + + rctx->ckpt_dma =3D DMA_MAPPING_ERROR; + if (rctx->has_checkpoint) { + rctx->ckpt_dma =3D cmh_dma_map_single(rctx->checkpoint, + HC_CONTEXT_SIZE, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->ckpt_dma)) { + ret =3D -ENOMEM; + goto err_unmap_save; + } + } + + rctx->data_len =3D full_len; + + /* Build VCQ: INIT [+ RESTORE] + UPDATE + SAVE + FLUSH */ + d =3D cmh_core_select_instance(CMH_CORE_HC); + idx =3D 0; + + vcq_add_hc_init(&cmds[idx++], d.core_id, info->hc_algo); + + if (rctx->has_checkpoint) + vcq_add_hc_restore(&cmds[idx++], d.core_id, + (u64)rctx->ckpt_dma, HC_CONTEXT_SIZE); + + vcq_add_hc_update(&cmds[idx++], d.core_id, + (u64)rctx->data_dma, full_len); + + vcq_add_hc_save(&cmds[idx++], d.core_id, + (u64)rctx->save_dma, HC_CONTEXT_SIZE); + + vcq_add_flush(&cmds[idx++], d.core_id); + + ret =3D cmh_vcq_pack_and_submit_async(cmds, idx, rctx->packed, + CMH_HASH_MAX_PACKED, + d.mbx_idx, + cmh_hash_update_complete, req, + !!(req->base.flags & + CRYPTO_TFM_REQ_MAY_BACKLOG), + cmh_tm_async_timeout_jiffies()); + if (ret && ret !=3D -EBUSY) + goto err_unmap_ckpt; + + if (ret =3D=3D -EBUSY) + return -EBUSY; + return -EINPROGRESS; + +err_unmap_ckpt: + if (rctx->has_checkpoint) + cmh_dma_unmap_single(rctx->ckpt_dma, HC_CONTEXT_SIZE, + DMA_TO_DEVICE); +err_unmap_save: + cmh_dma_unmap_single(rctx->save_dma, HC_CONTEXT_SIZE, + DMA_FROM_DEVICE); +err_unmap_data: + cmh_dma_unmap_single(rctx->data_dma, full_len, DMA_TO_DEVICE); +err_free: + kfree(rctx->save_buf); + rctx->save_buf =3D NULL; + kfree(rctx->data_buf); + rctx->data_buf =3D NULL; + rctx->data_len =3D 0; + return ret; +} + +/* + * Final completion -- unmap all DMA, copy digest, signal done. + */ +static void cmh_hash_final_complete(void *data, int error) +{ + struct ahash_request *req =3D data; + struct cmh_hash_reqctx *rctx =3D ahash_request_ctx(req); + + if (error =3D=3D -EINPROGRESS) { + cmh_complete(&req->base, error); + return; + } + + if (rctx->has_checkpoint) + cmh_dma_unmap_single(rctx->ckpt_dma, HC_CONTEXT_SIZE, + DMA_TO_DEVICE); + if (rctx->data_buf) + cmh_dma_unmap_single(rctx->data_dma, rctx->data_len, + DMA_TO_DEVICE); + cmh_dma_unmap_single(rctx->digest_dma, rctx->info->digest_size, + DMA_FROM_DEVICE); + + if (!error) + memcpy(req->result, rctx->digest_buf, + rctx->info->digest_size); + + kfree(rctx->digest_buf); + rctx->digest_buf =3D NULL; + kfree(rctx->data_buf); + rctx->data_buf =3D NULL; + cmh_hash_free_reqctx(rctx); + cmh_complete(&req->base, error); +} + +/* + * Submit the final VCQ transaction: + * INIT [+ RESTORE] [+ UPDATE(residual)] + FINAL + FLUSH + * + * @data_buf: linearised residual data, or NULL for empty-hash. + * Ownership transferred -- callback frees it. + * @data_len: bytes in data_buf. + */ +static int cmh_hash_submit_final(struct ahash_request *req, + u8 *data_buf, u32 data_len) +{ + struct cmh_hash_reqctx *rctx =3D ahash_request_ctx(req); + const struct cmh_hash_alg_info *info =3D rctx->info; + struct vcq_cmd cmds[CMH_HASH_MAX_PAYLOAD]; + struct core_dispatch d; + u32 idx; + int ret; + gfp_t gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + + rctx->data_buf =3D data_buf; + rctx->data_len =3D data_len; + + /* Allocate digest output buffer */ + rctx->digest_buf =3D kzalloc(info->digest_size, gfp); + if (!rctx->digest_buf) { + ret =3D -ENOMEM; + goto err_free_data; + } + + rctx->digest_dma =3D cmh_dma_map_single(rctx->digest_buf, + info->digest_size, + DMA_FROM_DEVICE); + if (cmh_dma_map_error(rctx->digest_dma)) { + ret =3D -ENOMEM; + goto err_free_digest; + } + + /* Map residual data for UPDATE */ + rctx->data_dma =3D DMA_MAPPING_ERROR; + if (data_buf && data_len > 0) { + rctx->data_dma =3D cmh_dma_map_single(data_buf, data_len, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->data_dma)) { + ret =3D -ENOMEM; + goto err_unmap_digest; + } + } + + /* Map checkpoint for RESTORE */ + rctx->ckpt_dma =3D DMA_MAPPING_ERROR; + if (rctx->has_checkpoint) { + rctx->ckpt_dma =3D cmh_dma_map_single(rctx->checkpoint, + HC_CONTEXT_SIZE, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->ckpt_dma)) { + ret =3D -ENOMEM; + goto err_unmap_data; + } + } + + /* Build VCQ: INIT [+ RESTORE] [+ UPDATE] + FINAL + FLUSH */ + d =3D cmh_core_select_instance(CMH_CORE_HC); + idx =3D 0; + + vcq_add_hc_init(&cmds[idx++], d.core_id, info->hc_algo); + + if (rctx->has_checkpoint) + vcq_add_hc_restore(&cmds[idx++], d.core_id, + (u64)rctx->ckpt_dma, HC_CONTEXT_SIZE); + + if (data_buf && data_len > 0) + vcq_add_hc_update(&cmds[idx++], d.core_id, + (u64)rctx->data_dma, data_len); + + vcq_add_hc_final(&cmds[idx++], d.core_id, + (u64)rctx->digest_dma, info->digest_size); + + vcq_add_flush(&cmds[idx++], d.core_id); + + ret =3D cmh_vcq_pack_and_submit_async(cmds, idx, rctx->packed, + CMH_HASH_MAX_PACKED, + d.mbx_idx, + cmh_hash_final_complete, req, + !!(req->base.flags & + CRYPTO_TFM_REQ_MAY_BACKLOG), + cmh_tm_async_timeout_jiffies()); + if (ret =3D=3D -EBUSY) + return -EBUSY; + if (ret) + goto err_unmap_ckpt; + + return -EINPROGRESS; + +err_unmap_ckpt: + if (rctx->has_checkpoint) + cmh_dma_unmap_single(rctx->ckpt_dma, HC_CONTEXT_SIZE, + DMA_TO_DEVICE); +err_unmap_data: + if (data_buf && data_len > 0) + cmh_dma_unmap_single(rctx->data_dma, data_len, + DMA_TO_DEVICE); +err_unmap_digest: + cmh_dma_unmap_single(rctx->digest_dma, info->digest_size, + DMA_FROM_DEVICE); +err_free_digest: + kfree(rctx->digest_buf); + rctx->digest_buf =3D NULL; +err_free_data: + kfree(data_buf); + rctx->data_buf =3D NULL; + /* + * Preserve the HC checkpoint on failure: a synchronous rejection is + * retryable, and for a terminal error the inline checkpoint is freed + * with the request context, so it never leaks. It is cleared only in + * the completion after a successful final(). + */ + return ret; +} + +static int cmh_hash_finup(struct ahash_request *req); + +/* + * One-shot digest -- delegates to init + finup so that all data is + * linearised and mapped through cmh_dma_map_single(), which is the + * only DMA mapping path aware of all supported DMA backends. + */ +static int cmh_hash_digest(struct ahash_request *req) +{ + int ret; + + ret =3D cmh_hash_init(req); + if (ret) + return ret; + return cmh_hash_finup(req); +} + +/* + * .finup -- hash any remaining data and finalise in one transaction. + * + * With BLOCK_ONLY the Crypto API prepends the bytes it held back, so + * req->src already carries the full tail; linearise it and submit + * INIT [+ RESTORE] [+ UPDATE(residual)] + FINAL + FLUSH. This also + * serves .final (the API calls finup with nbytes =3D=3D 0) and avoids + * ahash_def_finup(), which would clone via export/import. + */ +static int cmh_hash_finup(struct ahash_request *req) +{ + struct cmh_hash_reqctx *rctx =3D ahash_request_ctx(req); + u32 data_len =3D req->nbytes; + u8 *data_buf =3D NULL; + gfp_t gfp; + + if (rctx->error) + return rctx->error; + + if (data_len =3D=3D 0) + return cmh_hash_submit_final(req, NULL, 0); + + /* Reject an oversized linearisation with a permanent -EMSGSIZE. */ + if (data_len > KMALLOC_MAX_SIZE) + return -EMSGSIZE; + + gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + + data_buf =3D kmalloc(data_len, gfp | __GFP_NOWARN); + if (!data_buf) + return -ENOMEM; + + scatterwalk_map_and_copy(data_buf, req->src, 0, data_len, 0); + + return cmh_hash_submit_final(req, data_buf, data_len); +} + +/* + * Export core -- purely software. + * + * Serialise the HC checkpoint (if any). The Crypto API appends its own + * partial-block buffer to the exported state; this callback carries only + * HW state. No HW interaction needed because the incremental model + * keeps the checkpoint up-to-date after each .update(). + */ +static int cmh_hash_export(struct ahash_request *req, void *out) +{ + struct cmh_hash_reqctx *rctx =3D ahash_request_ctx(req); + struct cmh_hash_export_state *state =3D out; + + /* + * Zero the whole exported state first: the struct may carry padding, + * so without this the padding would leak kernel memory to user space + * through the ahash export (e.g. algif_hash). + */ + memset(state, 0, sizeof(*state)); + + if (rctx->hw_started && rctx->has_checkpoint) + memcpy(state->checkpoint, rctx->checkpoint, HC_CONTEXT_SIZE); + + state->hw_started =3D rctx->hw_started; + + return 0; +} + +/* + * Import core -- purely software. + * + * Restore the HC checkpoint from a previously exported state. The + * Crypto API restores its own partial-block buffer separately. The + * next .update() or final op will RESTORE the checkpoint into HW. + */ +static int cmh_hash_import(struct ahash_request *req, const void *in) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_hash_reqctx *rctx =3D ahash_request_ctx(req); + const struct cmh_hash_export_state *state =3D in; + + memset(rctx, 0, sizeof(*rctx)); + rctx->info =3D cmh_hash_get_info(tfm); + + rctx->hw_started =3D state->hw_started; + + if (state->hw_started) { + memcpy(rctx->checkpoint, state->checkpoint, HC_CONTEXT_SIZE); + rctx->has_checkpoint =3D 1; + } + + return 0; +} + +/* Registration */ + +static struct cmh_hash_alg_drv cmh_hash_drvs[CMH_HASH_ALG_COUNT]; + +/** + * cmh_hash_register() - Register SHA-256/384/512/3-256/3-384/3-512 hash a= lgorithms + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_hash_register(void) +{ + unsigned int i; + int ret; + + if (!cmh_core_present(CMH_CORE_HC)) + return 0; + + for (i =3D 0; i < CMH_HASH_ALG_COUNT; i++) { + const struct cmh_hash_alg_info *info =3D &cmh_hash_algs_info[i]; + struct cmh_hash_alg_drv *drv =3D &cmh_hash_drvs[i]; + struct ahash_alg *alg =3D &drv->alg; + + drv->info =3D info; + + alg->init =3D cmh_hash_init; + alg->update =3D cmh_hash_update; + alg->finup =3D cmh_hash_finup; + alg->digest =3D cmh_hash_digest; + alg->export =3D cmh_hash_export; + alg->import =3D cmh_hash_import; + + alg->halg.digestsize =3D info->digest_size; + alg->halg.statesize =3D sizeof(struct cmh_hash_export_state); + + strscpy(alg->halg.base.cra_name, info->alg_name, + CRYPTO_MAX_ALG_NAME); + strscpy(alg->halg.base.cra_driver_name, info->drv_name, + CRYPTO_MAX_ALG_NAME); + alg->halg.base.cra_priority =3D 300; + alg->halg.base.cra_flags =3D CRYPTO_ALG_KERN_DRIVER_ONLY | + CRYPTO_ALG_NO_FALLBACK | + CRYPTO_ALG_ASYNC | + CRYPTO_AHASH_ALG_BLOCK_ONLY; + alg->halg.base.cra_blocksize =3D info->block_size; + alg->halg.base.cra_ctxsize =3D 0; + alg->halg.base.cra_reqsize =3D sizeof(struct cmh_hash_reqctx); + alg->halg.base.cra_module =3D THIS_MODULE; + + ret =3D crypto_register_ahash(alg); + if (ret) { + dev_err(cmh_dev(), "hash: failed to register %s (rc=3D%d)\n", + info->drv_name, ret); + /* Unregister any already-registered algorithms */ + while (i--) + crypto_unregister_ahash(&cmh_hash_drvs[i].alg); + return ret; + } + + dev_dbg(cmh_dev(), "hash: registered %s (priority 300)\n", + info->drv_name); + } + + return 0; +} + +/** + * cmh_hash_unregister() - Unregister SHA hash algorithms from the crypto = framework + */ +void cmh_hash_unregister(void) +{ + unsigned int i; + + if (!cmh_core_present(CMH_CORE_HC)) + return; + + for (i =3D 0; i < CMH_HASH_ALG_COUNT; i++) { + crypto_unregister_ahash(&cmh_hash_drvs[i].alg); + dev_dbg(cmh_dev(), "hash: unregistered %s\n", + cmh_hash_algs_info[i].drv_name); + } +} diff --git a/drivers/crypto/cmh/cmh_main.c b/drivers/crypto/cmh/cmh_main.c index d7753c4630ba..50218f19ad6f 100644 --- a/drivers/crypto/cmh/cmh_main.c +++ b/drivers/crypto/cmh/cmh_main.c @@ -31,6 +31,7 @@ #include "cmh_mqi.h" #include "cmh_txn.h" #include "cmh_rh.h" +#include "cmh_hash.h" #include "cmh_mgmt.h" #include "cmh_registers.h" #include "cmh_debugfs.h" @@ -202,6 +203,11 @@ static int cmh_probe(struct platform_device *pdev) if (ret) goto err_rh_init; =20 + /* Register hash algorithms with the kernel crypto API */ + ret =3D cmh_hash_register(); + if (ret) + goto err_hash_register; + /* Register key management device (/dev/cmh_mgmt) */ ret =3D cmh_mgmt_register(); if (ret) @@ -212,6 +218,8 @@ static int cmh_probe(struct platform_device *pdev) return 0; =20 err_mgmt_register: + cmh_hash_unregister(); +err_hash_register: cmh_rh_cleanup(cfg); err_rh_init: cmh_tm_cleanup(); @@ -238,6 +246,7 @@ static void cmh_remove(struct platform_device *pdev) cfg =3D &dev->config; =20 cmh_mgmt_unregister(); + cmh_hash_unregister(); cmh_rh_cleanup(cfg); cmh_tm_cleanup(); cmh_mqi_cleanup(cfg); diff --git a/drivers/crypto/cmh/include/cmh_hash.h b/drivers/crypto/cmh/inc= lude/cmh_hash.h new file mode 100644 index 000000000000..198e557be280 --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_hash.h @@ -0,0 +1,27 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Kernel Crypto API Hash Driver + * + * Registers ahash algorithms (SHA-2, SHA-3, and SHAKE families) with the + * Linux crypto subsystem. These are CRYPTO_AHASH_ALG_BLOCK_ONLY drivers, + * so the Crypto API buffers partial blocks and hands the driver only + * whole-block-aligned data: + * + * .init() -> software-only: zero per-request context + * .update() -> INIT [+ RESTORE] + UPDATE(full blocks) + SAVE + FLUSH + * (also serves final: the API calls finup with nbytes =3D= =3D 0) + * .digest() -> INIT + UPDATE + FINAL + FLUSH (single-shot) + * .export() -> software-only: copy the HC checkpoint + * .import() -> software-only: restore the HC checkpoint + */ + +#ifndef CMH_HASH_H +#define CMH_HASH_H + +#include "cmh_config.h" + +int cmh_hash_register(void); +void cmh_hash_unregister(void); + +#endif /* CMH_HASH_H */ --=20 2.43.7 From nobody Fri Sep 25 01:20:34 2026 Received: from CO1PR03CU002.outbound.protection.outlook.com (mail-westus2azon11020074.outbound.protection.outlook.com [52.101.46.74]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BD5C343DEA2; Thu, 17 Sep 2026 22:59:41 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.46.74 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789685985; cv=fail; b=HqPs9lCoXcKKrfHw8LeF4jVVingmMjFAb4QMlB741NqaDG1O/L1W+T2PQEAUAWZzSH1BnUz7yWAhJCYey6jymVpmp4dnv5zBbKfb3P0XNye3KXVGx917+xsPjGfP0tt91ccDY9Dfj3V5KT3yHCL1Y6oFX5qo7PWxFLsfzwRlhRU= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789685985; c=relaxed/simple; bh=tGz3nA1IkJ/SbNyGqO+JMDRvMrB1e04pevIkLGWaS7c=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=MJN8XWugpqL5t2H/qX047UuKHJr/A/z1NW3B5Jmlaqh3NGWvRrgXNaZo4//DbyDevs7NnoqARgQqKPUxx3QMl65hZZGTEQnqxKIUG+ii95wnpH2/9vno50W6I57BX6BPk43rb3QU2eZn37Qac1P20fEpbK3fhZW1hboYtP4G/iw= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com; spf=fail smtp.mailfrom=rambus.com; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b=c65R40Cw; arc=fail smtp.client-ip=52.101.46.74 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=rambus.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b="c65R40Cw" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=rDGjJU8VBm20D28FjpMuaLpp+8gzZf8qHg7AbszfiZr4u10PqZP7JxBmc4rJsu8mtHWcAA1peAr3NDU3etAtUA7mceSEnWGwxoy9Ug4s38WEFfbp/MTlWloVdwSnNJMuGNoKiA+ImYilSgss9BCE87TddnBv0n4XgL7JJX9WfjVQy7g7/1Kya2v5XFXFLlDr5R6HRrtiHYvqqa8dr+Aqi9+eIawP6RdkZ+KeqRXtj5LjCjLMuNwubFPZyHHQ0yGnNmwANvUftTbH6ybTB8lmE7evTu2CT8UP3u/vG6U45PyvuVRLNgeHqWkSXHf7ixRc7LT74/rX0C/+knv0FPl7Sg== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=s1m9vDA9rGtFlC3YUIVbDELEb1F5sbVpEg50LGXOiLw=; b=tbz6ZT/wTOHCpG1sch3nIp+nwnb2Ao4b7c17HUDZeTXiRyNnosLWF/Cz4pW6HG2cVUN3m0KBoyS/V940hplng2rsgenFs8CpAzc401Q4rFXibS+T/UQk/qV7ei9jLX4HavNxg8eiXoScDlSjInAD9CAQ8M02AjQONJ5rQgPXdAiS0XIV7sct5cENxZN4BlsIwomLOqfdGYP5W0zdhRONmSaC5U0GUS/cFk+x12UsoRH5GLAwyCMSU/lZfVCQAok521m8KFmoUvR6BMP5/t1ltnnLD4CYm5+InzpkNiyJHrsZB9k+ajaXF17O39vzyM5mK90HZe1bhZ3m4qqJW26mSA== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 192.86.86.210) smtp.rcpttodomain=cryptography.com smtp.mailfrom=rambus.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=rambus.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=rambus.com; s=selector1; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=s1m9vDA9rGtFlC3YUIVbDELEb1F5sbVpEg50LGXOiLw=; b=c65R40Cwpn1K1gOkuMFVj+PpWjGNWJJMauLu0OnwR+I7bIv3MrQGvYfTOJx8lNrAuVA5320Obel9hOR2FAeRbV0V4oGv4kdYIGQ3M3ysmTqeSUUl4KyknsCTXV1hkLN6zHuW0Vlxp6CsBunJ4TBQcB3c/pkdAHgwDPkkcpJZg1a60ZfPf5FcFXnmGR8HjXp+84YOQNn07c3FyK7urX+8q9LZAL+GRZc820H+GmdQ1LIC4VV5jNynXBwWsXJ4ZUpXGhxFqIX3/rGKuqKQCxJLP9lEipeDcsGuteHkbIqqR3+7VU5gqN4smIYgW+z4dEVVbCKdsla9lBiUrsolQR0Inw== Received: from CH5P221CA0023.NAMP221.PROD.OUTLOOK.COM (2603:10b6:610:1f2::25) by SJ0PR04MB7757.namprd04.prod.outlook.com (2603:10b6:a03:3af::24) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.428.13; Thu, 17 Sep 2026 22:59:33 +0000 Received: from BN2PEPF0000A801.namprd02.prod.outlook.com (2603:10b6:610:1f2:cafe::2d) by CH5P221CA0023.outlook.office365.com (2603:10b6:610:1f2::25) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.428.12 via Frontend Transport; Thu, 17 Sep 2026 22:59:33 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 192.86.86.210) smtp.mailfrom=rambus.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=rambus.com; Received-SPF: Pass (protection.outlook.com: domain of rambus.com designates 192.86.86.210 as permitted sender) receiver=protection.outlook.com; client-ip=192.86.86.210; helo=hqxsv-psmtppxy02.rambus.com; pr=C Received: from hqxsv-psmtppxy02.rambus.com (192.86.86.210) by BN2PEPF0000A801.mail.protection.outlook.com (10.167.245.170) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.451.8 via Frontend Transport; Thu, 17 Sep 2026 22:59:33 +0000 Received: from hqxsv-cmdev3-aousherovitch.rambus.com (hqn-lb-int-float.rambus.com [10.12.20.20]) by hqxsv-psmtppxy02.rambus.com (Postfix) with ESMTP id F3E411801760; Thu, 17 Sep 2026 22:59:31 +0000 (UTC) From: Alex Ousherovitch To: Albert Ou , Alex Ousherovitch , Conor Dooley , "David S. Miller" , Herbert Xu , Jonathan Corbet , Krzysztof Kozlowski , Palmer Dabbelt , Paul Walmsley , Rob Herring , Saravanakrishnan Krishnamoorthy , Shuah Khan Cc: Alexandre Ghiti , devicetree@vger.kernel.org, Joel Wittenauer , linux-api@vger.kernel.org, linux-crypto@vger.kernel.org, linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-riscv@lists.infradead.org, Shuah Khan , Thi Nguyen Subject: [PATCH v5 05/19] crypto: cmh - add HMAC ahash Date: Thu, 17 Sep 2026 15:59:14 -0700 Message-ID: <20260917225929.2494111-6-aousherovitch@rambus.com> X-Mailer: git-send-email 2.43.7 In-Reply-To: <20260917225929.2494111-1-aousherovitch@rambus.com> References: <20260917225929.2494111-1-aousherovitch@rambus.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: BN2PEPF0000A801:EE_|SJ0PR04MB7757:EE_ X-MS-Office365-Filtering-Correlation-Id: 02cc95f3-7163-49df-b935-08df150f56cc X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|82310400026|36860700016|23010399003|1800799024|7416014|376014|56012099006|11063799006|5023799004|3023799007|10067099003|6133799003|921020|22082099003|18002099003; X-Microsoft-Antispam-Message-Info: U6mKRGFR453fDzb2KI6d2bT1uA3vpZjaNVe8+zLHAaHTuonBLuFr6EtjF5F3yZMk2tPPtDuXr6iJWmnaldplA+U8ahXLNf2qxUC2W9PrTUvxXYEdBp+IvYhhG4Lgwy1gYcohcNPJnOeJr/oPDmSkHrCGtr7rGnafMzad2+bjjYeoKSEcwfmcoj84tX0iuGZHCNqv9FBFMrwRIXUstdenvIqKwl8XXHyWWNugfPZlGacPTgDWM2XJKQkiCYPiA0ixT2RrI8clDXjf3+gMa5AfBBPLH2rczNqmZio6JE88VHKtRuvIYcB0iZKsHWkoc2mszkqbB6YaFfeEnmTEe1FsAff0pCm7n39DHwaeIdgNdIY96fcOvk6ExGdNrNW9gNGUQCkwWsSG29JciSyDNZ/ScdELD6vn+Q0ZU42yLXU2ipG9Tpev95gePDCXPyzWnY0OkN6fnajyc2npF90n/bQep89w2JiZibmFZwDv8Dnq1BjfdlfyUatSBogiHhz2byd/au1g9jpkBYUgitvRxcKjvmeQdvIFQKOOihdmBfq87m7jgLDyHeT/rFrViLglb+UxIU0a0jBdWiSDfJM+9Sj9CnHAMVDC76dXHFfwKIY/h/PJxRZhrF0i7bF1DpXZWAwB8aWMyqXHDC4qRbto62kt7Z9iUgmIiPlZ708lYlfPMeV1hEixJBJECj0tRBy2+p9yNiBlAiagunpuWyVPyGwQVCa0wVb0lg2QwW7KAowmrHTKBmMHjh6PLi66PEZzInWJ X-Forefront-Antispam-Report: CIP:192.86.86.210;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:hqxsv-psmtppxy02.rambus.com;PTR:InfoDomainNonexistent;CAT:NONE;SFS:(13230040)(82310400026)(36860700016)(23010399003)(1800799024)(7416014)(376014)(56012099006)(11063799006)(5023799004)(3023799007)(10067099003)(6133799003)(921020)(22082099003)(18002099003);DIR:OUT;SFP:1102; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: r89V1pKlXgfDf6mVh/MXLrl0DFWmV7dUcf5cDRnVwev6QBKHweqhX0VvRJh6fCFhKM4qUm5Rd3QzhqM30k8aYbX22W8M61Eb1J25A5Ygf0ipgIbTwQkILKQ2NLBmHg88gIER+I5h1U4g+MdqQWturdUYcq77SL4XEty8IbSzn8MsD0wHU7RGfi/N4LAJRRbOSdcNzWNmuxcPgdtcbUj2Xqk7mYMnHeRhf1Is9gDNor0Jlj6B5c5JzBKGUlPvoQdSt8jDtSYZCzDRAD0bD+i1MmnBSaT0dh1FLEGe+xwBSmJa1xK7larwJNPAj23PA1jYwbS+gJF697711Gs1HVLmxnfVhPXhoAKolZACon0M2/IRSXyoa2q4QMpKzc48vuNk5jHSNMxSFqKZyOnoLOfVC1KJ/KxahcxJSPNnMtciB+DOSmRzHtN8WTr3e2HKdwtW X-OriginatorOrg: rambus.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 17 Sep 2026 22:59:33.2535 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: 02cc95f3-7163-49df-b935-08df150f56cc X-MS-Exchange-CrossTenant-Id: bd0ba799-c2b9-413c-9c56-5d1731c4827c X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=bd0ba799-c2b9-413c-9c56-5d1731c4827c;Ip=[192.86.86.210];Helo=[hqxsv-psmtppxy02.rambus.com] X-MS-Exchange-CrossTenant-AuthSource: BN2PEPF0000A801.namprd02.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: SJ0PR04MB7757 Content-Type: text/plain; charset="utf-8" Register ahash algorithms for HMAC-SHA-224, HMAC-SHA-256, HMAC-SHA-384, HMAC-SHA-512, HMAC-SHA3-224, HMAC-SHA3-256, HMAC-SHA3-384, and HMAC-SHA3-512 using the CMH hash core. HMAC is deliberately not CRYPTO_AHASH_ALG_BLOCK_ONLY, unlike the plain hashes: the hardware has no keyed-MAC save/restore, so a MAC cannot be computed block-incrementally. The driver accumulates the request input and submits it in a single HW transaction at finalize. Above a 64KB hardware window it switches transparently to the generic software HMAC, so there is no input-size limit. Signed-off-by: Alex Ousherovitch Co-developed-by: Saravanakrishnan Krishnamoorthy Signed-off-by: Saravanakrishnan Krishnamoorthy --- drivers/crypto/cmh/Makefile | 3 +- drivers/crypto/cmh/cmh_hmac.c | 908 ++++++++++++++++++++++++++ drivers/crypto/cmh/cmh_main.c | 9 + drivers/crypto/cmh/include/cmh_hmac.h | 16 + 4 files changed, 935 insertions(+), 1 deletion(-) create mode 100644 drivers/crypto/cmh/cmh_hmac.c create mode 100644 drivers/crypto/cmh/include/cmh_hmac.h diff --git a/drivers/crypto/cmh/Makefile b/drivers/crypto/cmh/Makefile index 79c94d87e9ee..acd1827cc084 100644 --- a/drivers/crypto/cmh/Makefile +++ b/drivers/crypto/cmh/Makefile @@ -15,7 +15,8 @@ cmh-y :=3D \ cmh_sysfs.o \ cmh_key.o \ cmh_sys.o \ - cmh_hash.o + cmh_hash.o \ + cmh_hmac.o =20 # Management ioctl device (/dev/cmh_mgmt): key lifecycle, PKE, PQC ioctls. cmh-$(CONFIG_CRYPTO_DEV_CMH_MGMT) +=3D \ diff --git a/drivers/crypto/cmh/cmh_hmac.c b/drivers/crypto/cmh/cmh_hmac.c new file mode 100644 index 000000000000..7d984bc745e1 --- /dev/null +++ b/drivers/crypto/cmh/cmh_hmac.c @@ -0,0 +1,908 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Kernel Crypto API HMAC Driver + * + * Registers HMAC ahash algorithms with the Linux crypto subsystem. + * Supports HMAC-SHA-2 (224/256/384/512) and HMAC-SHA-3 (224/256/384/512) + * using the CMH Hash Core (HC) via HC_CMD_HMAC. + * + * Uses the same self-contained transaction model as cmh_hash.c: + * .setkey() -> store raw key bytes + * .init() -> software-only: initialize per-request context + * .update() -> software-only: copy SG data into per-call chunk + * .final() -> [SYS_CMD_WRITE] + HC_CMD_HMAC + [GATHER] + FINAL + FLUSH + * + * Raw-key atomicity: SYS_CMD_WRITE to SYS_REF_TEMP is packed into + * the same VCQ as HC_CMD_HMAC (see cmh_key.h for details). + * + * ahash .export()/.import() (state cloning): the HW hash core does NOT + * support save/restore of intermediate HMAC state, so the driver + * accumulates input in kernel memory and serialises that buffer for + * the common (bounded) case. When the accumulated input exceeds the + * HW cap (HMAC_MAX_DATA, 64 KB) or the flat export window, the request + * transparently switches to a generic software HMAC fallback that the + * driver allocates and keys itself: buffered chunks are replayed into it = and + * all further input streams through it, so arbitrary-length hashing and + * transform clone both remain conformant with O(1) driver memory. + */ + +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "cmh_hmac.h" +#include "cmh_vcq.h" +#include "cmh_hc_abi.h" +#include "cmh_sys_abi.h" +#include "cmh_sys.h" +#include "cmh_txn.h" +#include "cmh_dma.h" +#include "cmh_key.h" + +/* + * Maximum data that can be accumulated across .update() calls. + * HMAC save/restore is intentionally unsupported (see file header), + * so all data must be buffered in kernel memory and submitted + * atomically in .final(). This cap prevents unbounded allocation. + */ +#define HMAC_MAX_DATA (64 * 1024) + +/* Algorithm Table */ + +struct cmh_hmac_alg_info { + u32 hc_algo; /* HC_ALGO_* */ + u32 digest_size; /* bytes */ + u32 block_size; /* cra_blocksize */ + const char *alg_name; /* Linux crypto name: "hmac(sha256)" */ + const char *drv_name; /* driver name: "rambus-cmh-hmac-sha256" */ +}; + +static const struct cmh_hmac_alg_info cmh_hmac_algs_info[] =3D { + /* HMAC-SHA-2 family */ + { + .hc_algo =3D HC_ALGO_SHA2_224, + .digest_size =3D CMH_SHA224_DIGEST_SIZE, + .block_size =3D 64, + .alg_name =3D "hmac(sha224)", + .drv_name =3D "rambus-cmh-hmac-sha224", + }, + { + .hc_algo =3D HC_ALGO_SHA2_256, + .digest_size =3D CMH_SHA256_DIGEST_SIZE, + .block_size =3D 64, + .alg_name =3D "hmac(sha256)", + .drv_name =3D "rambus-cmh-hmac-sha256", + }, + { + .hc_algo =3D HC_ALGO_SHA2_384, + .digest_size =3D CMH_SHA384_DIGEST_SIZE, + .block_size =3D 128, + .alg_name =3D "hmac(sha384)", + .drv_name =3D "rambus-cmh-hmac-sha384", + }, + { + .hc_algo =3D HC_ALGO_SHA2_512, + .digest_size =3D CMH_SHA512_DIGEST_SIZE, + .block_size =3D 128, + .alg_name =3D "hmac(sha512)", + .drv_name =3D "rambus-cmh-hmac-sha512", + }, + /* HMAC-SHA-3 family */ + { + .hc_algo =3D HC_ALGO_SHA3_224, + .digest_size =3D CMH_SHA3_224_DIGEST_SIZE, + .block_size =3D 144, + .alg_name =3D "hmac(sha3-224)", + .drv_name =3D "rambus-cmh-hmac-sha3-224", + }, + { + .hc_algo =3D HC_ALGO_SHA3_256, + .digest_size =3D CMH_SHA3_256_DIGEST_SIZE, + .block_size =3D 136, + .alg_name =3D "hmac(sha3-256)", + .drv_name =3D "rambus-cmh-hmac-sha3-256", + }, + { + .hc_algo =3D HC_ALGO_SHA3_384, + .digest_size =3D CMH_SHA3_384_DIGEST_SIZE, + .block_size =3D 104, + .alg_name =3D "hmac(sha3-384)", + .drv_name =3D "rambus-cmh-hmac-sha3-384", + }, + { + .hc_algo =3D HC_ALGO_SHA3_512, + .digest_size =3D CMH_SHA3_512_DIGEST_SIZE, + .block_size =3D 72, + .alg_name =3D "hmac(sha3-512)", + .drv_name =3D "rambus-cmh-hmac-sha3-512", + }, +}; + +#define CMH_HMAC_ALG_COUNT ARRAY_SIZE(cmh_hmac_algs_info) + +/* Per-Request State */ + +struct cmh_hmac_chunk { + struct list_head list; + struct list_head tfm_node; /* per-tfm orphan tracking */ + u32 len; + u8 data[]; +}; + +/* + * Maximum payload commands any HMAC transaction can produce: + * [SYS_CMD_WRITE] + HC_CMD_HMAC + [GATHER] + FINAL + FLUSH =3D 5 + * Worst-case packed output (stride=3D7, 1 payload per VCQ): + * 5 VCQs x 2 entries =3D 10 + */ +#define CMH_HMAC_MAX_PAYLOAD 5 +#define CMH_HMAC_MAX_PACKED (CMH_HMAC_MAX_PAYLOAD * 2) + +struct cmh_hmac_reqctx { + const struct cmh_hmac_alg_info *info; + int error; + struct list_head chunks; + u32 num_chunks; + u32 total_len; + bool switched; /* handed off to SW fallback = */ + /* DMA state for async final */ + dma_addr_t digest_dma; + dma_addr_t key_dma; + u8 *digest_buf; + struct cmh_sg_map *sgm; + u32 keylen; + struct vcq_cmd packed[CMH_HMAC_MAX_PACKED]; +}; + +/* + * Flat state for export/import, tagged by the leading @format byte: + * + * CMH_HMAC_FMT_RAW -- the accumulated input bytes, verbatim. Used + * while the request is still on the HW-buffered path and fits the + * flat window (total_len <=3D CMH_HMAC_EXPORT_MAX). + * CMH_HMAC_FMT_FB -- the software fallback's own exported state. + * Used once the request has switched to the fallback (oversized + * input, or an export past the flat window), so export/import + * (transform clone) works at any input length. + */ +#define CMH_HMAC_FMT_RAW 0 +#define CMH_HMAC_FMT_FB 1 + +struct cmh_hmac_export_state { + u8 format; + u8 __pad[3]; + u32 total_len; + u8 data[]; +}; + +/* + * The crypto subsystem pre-allocates statesize bytes per request. + * CMH_HMAC_STATE_SIZE (4096) sizes both the CMH_HMAC_FMT_RAW window + * (CMH_HMAC_EXPORT_MAX accumulated bytes) and the CMH_HMAC_FMT_FB + * software state (the generic fallback's much smaller statesize). A + * RAW export past CMH_HMAC_EXPORT_MAX transparently switches to the + * fallback and emits CMH_HMAC_FMT_FB instead, so export/import is not + * capped. + */ +#define CMH_HMAC_STATE_SIZE 4096 +#define CMH_HMAC_EXPORT_MAX (CMH_HMAC_STATE_SIZE - sizeof(struct cmh_hmac_= export_state)) + +/* Per-Transform State (carries key across requests) */ + +struct cmh_hmac_tfm_ctx { + struct cmh_key_ctx key; + struct crypto_ahash *fb; /* generic SW fallback (oversized ops) */ + spinlock_t chunk_lock; /* protects all_chunks + tfm_buffered */ + struct list_head all_chunks; /* orphan-safe chunk tracking */ + size_t tfm_buffered; /* bytes on all_chunks; DoS cap */ +}; + +/* + * Per-transform cap on total bytes buffered across all_chunks. Bounds + * memory an AF_ALG client can pin via repeated open/update/abandon of + * request sockets (the crypto API has no per-request destructor). + */ +#define CMH_HMAC_TFM_MAX_BUFFERED (16 * 1024 * 1024) + +/* VCQ Builders (HMAC-specific; shared builders in cmh_hc_abi.h / cmh_vcq.= h) */ + +/* Add an HC_CMD_HMAC entry */ +static void vcq_add_hc_hmac(struct vcq_cmd *slot, u32 core_id, u64 key_ref, + u32 keylen, u32 algo) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, HC_CMD_HMAC); + slot->hwc.hc.cmd_hmac.key =3D key_ref; + slot->hwc.hc.cmd_hmac.keylen =3D keylen; + slot->hwc.hc.cmd_hmac.algo =3D algo; +} + +/* Request Context Cleanup */ + +static void cmh_hmac_free_chunks(struct cmh_hmac_reqctx *rctx, + struct cmh_hmac_tfm_ctx *tctx) +{ + struct cmh_hmac_chunk *chunk, *tmp; + + spin_lock_bh(&tctx->chunk_lock); + list_for_each_entry_safe(chunk, tmp, &rctx->chunks, list) { + list_del(&chunk->list); + list_del(&chunk->tfm_node); + tctx->tfm_buffered -=3D chunk->len; + kfree_sensitive(chunk); + } + spin_unlock_bh(&tctx->chunk_lock); + rctx->num_chunks =3D 0; + rctx->total_len =3D 0; +} + +/* + * Build a DMA-mapped CMH eSW scatter-gather chain from accumulated chunks. + */ +static struct cmh_sg_map * +cmh_hmac_build_sg(struct cmh_hmac_reqctx *rctx, gfp_t gfp) +{ + struct cmh_dma_buf *bufs; + struct cmh_hmac_chunk *chunk; + struct cmh_sg_map *sgm; + u32 i; + + bufs =3D kcalloc(rctx->num_chunks, sizeof(*bufs), gfp); + if (!bufs) + return NULL; + + i =3D 0; + list_for_each_entry(chunk, &rctx->chunks, list) { + bufs[i].data =3D chunk->data; + bufs[i].len =3D chunk->len; + i++; + } + + sgm =3D cmh_dma_build_sg(bufs, rctx->num_chunks, gfp); + kfree(bufs); + return sgm; +} + +/* VCQ Packing + Submit */ + +/* ahash Operations */ + +struct cmh_hmac_alg_drv { + struct ahash_alg alg; + const struct cmh_hmac_alg_info *info; +}; + +static const struct cmh_hmac_alg_info * +cmh_hmac_get_info(struct crypto_ahash *tfm) +{ + struct ahash_alg *alg =3D crypto_ahash_alg(tfm); + + return container_of(alg, struct cmh_hmac_alg_drv, alg)->info; +} + +/* Software-fallback helpers (arbitrary-length + transform-clone support) = */ + +/* + * The fallback ahash_request lives immediately after the reqctx. + * cmh_hmac_cra_init() reserves crypto_ahash_reqsize(fb) bytes for it and + * PTR_ALIGN keeps it aligned for the fallback's own request context. + */ +static struct ahash_request *cmh_hmac_fb_req(struct cmh_hmac_reqctx *rctx) +{ + return PTR_ALIGN((void *)(rctx + 1), crypto_tfm_ctx_alignment()); +} + +/* + * Feed @len bytes of the linear buffer @data to the fallback request. + * The core allocated the fallback as a virt-capable transform, so a + * virtual address can be handed to it directly. The fallback is + * synchronous (shash-backed), so crypto_ahash_update() completes inline. + */ +static int cmh_hmac_fb_update_virt(struct ahash_request *fb_req, + const u8 *data, u32 len) +{ + ahash_request_set_virt(fb_req, data, NULL, len); + return crypto_ahash_update(fb_req); +} + +/* + * Switch a request from the HW-buffered path to the software fallback: + * initialise the fallback request, replay every accumulated chunk + * through it, then drop the chunks (their bytes now live in the + * fallback's running state). Afterwards the request is O(1) in memory + * and no longer input-capped. The fallback transform was keyed by + * cmh_hmac_setkey() when the caller installed the MAC key. + */ +static int cmh_hmac_switch_to_fb(struct ahash_request *req) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_hmac_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cmh_hmac_reqctx *rctx =3D ahash_request_ctx(req); + struct ahash_request *fb_req =3D cmh_hmac_fb_req(rctx); + struct cmh_hmac_chunk *chunk; + int ret; + + ahash_request_set_tfm(fb_req, tctx->fb); + ahash_request_set_callback(fb_req, 0, NULL, NULL); + + ret =3D crypto_ahash_init(fb_req); + if (ret) + return ret; + + list_for_each_entry(chunk, &rctx->chunks, list) { + ret =3D cmh_hmac_fb_update_virt(fb_req, chunk->data, chunk->len); + if (ret) + return ret; + } + + cmh_hmac_free_chunks(rctx, tctx); + rctx->switched =3D true; + return 0; +} + +/* + * Forward the current update() payload to the fallback and remember any + * error so a later final()/update() reports it. @req may carry either a + * virtual buffer or a scatterlist. + */ +static int cmh_hmac_fb_forward(struct ahash_request *req, + struct cmh_hmac_reqctx *rctx) +{ + struct ahash_request *fb_req =3D cmh_hmac_fb_req(rctx); + int ret; + + if (req->base.flags & CRYPTO_AHASH_REQ_VIRT) { + ret =3D cmh_hmac_fb_update_virt(fb_req, req->svirt, req->nbytes); + } else { + ahash_request_set_crypt(fb_req, req->src, NULL, req->nbytes); + ret =3D crypto_ahash_update(fb_req); + } + if (ret) + rctx->error =3D ret; + return ret; +} + +static int cmh_hmac_setkey(struct crypto_ahash *tfm, const u8 *key, + unsigned int keylen) +{ + struct cmh_hmac_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + int ret; + + ret =3D cmh_key_setkey_raw(&tctx->key, key, keylen, CORE_ID_HC); + if (ret) + return ret; + + /* Keep the software fallback keyed in lock-step for oversized ops. */ + return crypto_ahash_setkey(tctx->fb, key, keylen); +} + +static int cmh_hmac_init(struct ahash_request *req) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_hmac_reqctx *rctx =3D ahash_request_ctx(req); + + rctx->info =3D cmh_hmac_get_info(tfm); + rctx->error =3D 0; + INIT_LIST_HEAD(&rctx->chunks); + rctx->num_chunks =3D 0; + rctx->total_len =3D 0; + rctx->switched =3D false; + + return 0; +} + +static int cmh_hmac_update(struct ahash_request *req) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_hmac_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cmh_hmac_reqctx *rctx =3D ahash_request_ctx(req); + struct cmh_hmac_chunk *chunk; + int nents; + + if (rctx->error) + return rctx->error; + + if (!req->nbytes) + return 0; + + /* Already handed off to the fallback: forward directly (O(1) mem). */ + if (rctx->switched) + return cmh_hmac_fb_forward(req, rctx); + + /* + * Exceeding the HW input cap: switch to the software fallback + * (replaying the buffered chunks) rather than failing, then + * forward this update. + */ + if (req->nbytes > HMAC_MAX_DATA - rctx->total_len) { + rctx->error =3D cmh_hmac_switch_to_fb(req); + if (rctx->error) + goto err_free_chunks; + return cmh_hmac_fb_forward(req, rctx); + } + + chunk =3D kmalloc(sizeof(*chunk) + req->nbytes, + req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC); + if (!chunk) { + rctx->error =3D -ENOMEM; + goto err_free_chunks; + } + + chunk->len =3D req->nbytes; + if (req->base.flags & CRYPTO_AHASH_REQ_VIRT) { + memcpy(chunk->data, req->svirt, req->nbytes); + } else { + nents =3D sg_nents_for_len(req->src, req->nbytes); + if (nents < 0 || + sg_copy_to_buffer(req->src, nents, + chunk->data, req->nbytes) !=3D req->nbytes) { + kfree_sensitive(chunk); + rctx->error =3D -EINVAL; + goto err_free_chunks; + } + } + + spin_lock_bh(&tctx->chunk_lock); + if (tctx->tfm_buffered + chunk->len > CMH_HMAC_TFM_MAX_BUFFERED) { + spin_unlock_bh(&tctx->chunk_lock); + kfree_sensitive(chunk); + rctx->error =3D -ENOMEM; + goto err_free_chunks; + } + list_add_tail(&chunk->list, &rctx->chunks); + list_add_tail(&chunk->tfm_node, &tctx->all_chunks); + tctx->tfm_buffered +=3D chunk->len; + spin_unlock_bh(&tctx->chunk_lock); + rctx->num_chunks++; + rctx->total_len +=3D req->nbytes; + + return 0; + +err_free_chunks: + /* + * Terminal error -- free all previously accumulated chunks. + * The crypto API hash path does not call .final() + * on error, and hash_sock_destruct has no per-request + * destructor, so chunks would be orphaned otherwise. + */ + cmh_hmac_free_chunks(rctx, tctx); + return rctx->error; +} + +static void cmh_hmac_complete(void *data, int error) +{ + struct ahash_request *req =3D data; + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_hmac_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cmh_hmac_reqctx *rctx =3D ahash_request_ctx(req); + + if (error =3D=3D -EINPROGRESS) { + cmh_complete(&req->base, error); + return; + } + + cmh_dma_unmap_single(rctx->digest_dma, rctx->info->digest_size, + DMA_FROM_DEVICE); + + if (!error) + memcpy(req->result, rctx->digest_buf, + rctx->info->digest_size); + + kfree(rctx->digest_buf); + rctx->digest_buf =3D NULL; + cmh_dma_free_sg(rctx->sgm); + rctx->sgm =3D NULL; + cmh_hmac_free_chunks(rctx, tctx); + cmh_complete(&req->base, error); +} + +static int cmh_hmac_final(struct ahash_request *req) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_hmac_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cmh_hmac_reqctx *rctx =3D ahash_request_ctx(req); + const struct cmh_hmac_alg_info *info =3D rctx->info; + struct vcq_cmd cmds[CMH_HMAC_MAX_PAYLOAD]; + struct cmh_sg_map *sgm =3D NULL; + dma_addr_t digest_dma =3D DMA_MAPPING_ERROR, key_dma =3D DMA_MAPPING_ERRO= R; + u8 *digest_buf; + u64 key_ref; + u32 keylen; + struct core_dispatch d; + s32 target_mbx; + u32 core_id; + u32 idx; + int ret; + gfp_t gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + + if (rctx->error) { + ret =3D rctx->error; + goto out_free; + } + + /* Switched to the software fallback: complete there (synchronous). */ + if (rctx->switched) { + struct ahash_request *fb_req =3D cmh_hmac_fb_req(rctx); + + ahash_request_set_crypt(fb_req, NULL, req->result, 0); + return crypto_ahash_final(fb_req); + } + + if (tctx->key.mode =3D=3D CMH_KEY_NONE) { + ret =3D -ENOKEY; + goto out_free; + } + + if (rctx->num_chunks > 0) { + sgm =3D cmh_hmac_build_sg(rctx, gfp); + if (!sgm) { + ret =3D -ENOMEM; + goto out_free; + } + } + + digest_buf =3D kzalloc(info->digest_size, gfp); + if (!digest_buf) { + ret =3D -ENOMEM; + goto out_free_sg; + } + digest_dma =3D cmh_dma_map_single(digest_buf, info->digest_size, + DMA_FROM_DEVICE); + if (cmh_dma_map_error(digest_dma)) { + ret =3D -ENOMEM; + goto out_free_digest; + } + + /* Resolve key reference */ + idx =3D 0; + + /* + * Raw key: pack SYS_CMD_WRITE(SYS_REF_TEMP) into the + * same VCQ so the key write + HMAC are atomic. + */ + key_dma =3D tctx->key.raw.dma; + vcq_add_sys_write(&cmds[idx++], SYS_REF_TEMP, (u64)key_dma, + SYS_REF_NONE, tctx->key.raw.len, + tctx->key.raw.sys_type); + key_ref =3D SYS_REF_TEMP; + keylen =3D tctx->key.raw.len; + d =3D cmh_core_select_instance(CMH_CORE_HC); + + target_mbx =3D d.mbx_idx; + + core_id =3D d.core_id; + + vcq_add_hc_hmac(&cmds[idx++], core_id, key_ref, keylen, info->hc_algo); + + if (sgm) + vcq_add_hc_gather(&cmds[idx++], core_id, (u64)sgm->items_dma, + HC_CMD_UPDATE); + + vcq_add_hc_final(&cmds[idx++], core_id, (u64)digest_dma, info->digest_siz= e); + vcq_add_flush(&cmds[idx++], core_id); + + rctx->digest_buf =3D digest_buf; + rctx->digest_dma =3D digest_dma; + rctx->sgm =3D sgm; + + ret =3D cmh_vcq_pack_and_submit_async(cmds, idx, rctx->packed, + CMH_HMAC_MAX_PACKED, + target_mbx, + cmh_hmac_complete, req, + !!(req->base.flags & + CRYPTO_TFM_REQ_MAY_BACKLOG), + cmh_tm_async_timeout_jiffies()); + if (ret =3D=3D -EBUSY) + return -EBUSY; + if (ret) { + /* + * Synchronous rejection (e.g. -EAGAIN: CMQ full, no backlog). + * Free only the per-submit transients and keep the accumulated + * chunks intact so the caller can retry the identical final(). + * If no retry comes, cra_exit reclaims the orphaned chunks; the + * per-tfm buffered-byte cap bounds how much stays pinned. + */ + cmh_dma_unmap_single(digest_dma, info->digest_size, + DMA_FROM_DEVICE); + kfree(digest_buf); + rctx->digest_buf =3D NULL; + cmh_dma_free_sg(sgm); + rctx->sgm =3D NULL; + return ret; + } + + return -EINPROGRESS; + +out_free_digest: + kfree(digest_buf); + +out_free_sg: + cmh_dma_free_sg(sgm); + +out_free: + cmh_hmac_free_chunks(rctx, tctx); + return ret; +} + +static int cmh_hmac_finup(struct ahash_request *req) +{ + int ret; + + ret =3D cmh_hmac_update(req); + if (ret) + return ret; + + return cmh_hmac_final(req); +} + +static int cmh_hmac_digest(struct ahash_request *req) +{ + int ret; + + ret =3D cmh_hmac_init(req); + if (ret) + return ret; + + return cmh_hmac_finup(req); +} + +/* + * ahash .export()/.import(): serialize/deserialize the software + * accumulation buffer. No HW state is involved. + */ + +static int cmh_hmac_export(struct ahash_request *req, void *out) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_hmac_reqctx *rctx =3D ahash_request_ctx(req); + struct cmh_hmac_export_state *state =3D out; + struct cmh_hmac_chunk *chunk; + u32 offset =3D 0; + int ret; + + /* + * If more data is buffered than the flat window holds, switch to + * the software fallback so a bounded, fixed-size state can be + * exported -- making export/import (clone) work at any length. + */ + if (!rctx->switched && rctx->total_len > CMH_HMAC_EXPORT_MAX) { + ret =3D cmh_hmac_switch_to_fb(req); + if (ret) + return ret; + } + + /* Zero the whole state buffer so no kernel memory leaks out. */ + memset(state, 0, crypto_ahash_statesize(tfm)); + + if (rctx->switched) { + state->format =3D CMH_HMAC_FMT_FB; + return crypto_ahash_export(cmh_hmac_fb_req(rctx), state->data); + } + + state->format =3D CMH_HMAC_FMT_RAW; + state->total_len =3D rctx->total_len; + list_for_each_entry(chunk, &rctx->chunks, list) { + memcpy(state->data + offset, chunk->data, chunk->len); + offset +=3D chunk->len; + } + return 0; +} + +static int cmh_hmac_import(struct ahash_request *req, const void *in) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_hmac_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cmh_hmac_reqctx *rctx =3D ahash_request_ctx(req); + const struct cmh_hmac_export_state *state =3D in; + struct cmh_hmac_chunk *chunk; + + /* + * Do NOT call free_chunks() here: the crypto API does not + * guarantee the request context is in a valid state before + * import(), so the list pointers may be stale or invalid. + * Re-initialize from scratch instead. Any pre-existing chunks + * are tracked on tctx->all_chunks and freed in cra_exit. + */ + rctx->info =3D cmh_hmac_get_info(tfm); + rctx->error =3D 0; + INIT_LIST_HEAD(&rctx->chunks); + rctx->num_chunks =3D 0; + rctx->total_len =3D 0; + rctx->switched =3D false; + + /* Fallback-format state: replay it into a fallback request. */ + if (state->format =3D=3D CMH_HMAC_FMT_FB) { + struct ahash_request *fb_req =3D cmh_hmac_fb_req(rctx); + int ret; + + ahash_request_set_tfm(fb_req, tctx->fb); + ahash_request_set_callback(fb_req, 0, NULL, NULL); + ret =3D crypto_ahash_import(fb_req, state->data); + if (ret) + return ret; + rctx->switched =3D true; + return 0; + } + + if (state->format !=3D CMH_HMAC_FMT_RAW) + return -EINVAL; + + if (state->total_len > CMH_HMAC_EXPORT_MAX) + return -EINVAL; + + if (state->total_len) { + chunk =3D kmalloc(sizeof(*chunk) + state->total_len, + req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC); + if (!chunk) + return -ENOMEM; + chunk->len =3D state->total_len; + memcpy(chunk->data, state->data, state->total_len); + spin_lock_bh(&tctx->chunk_lock); + list_add_tail(&chunk->list, &rctx->chunks); + list_add_tail(&chunk->tfm_node, &tctx->all_chunks); + tctx->tfm_buffered +=3D chunk->len; + spin_unlock_bh(&tctx->chunk_lock); + rctx->num_chunks =3D 1; + rctx->total_len =3D state->total_len; + } + return 0; +} + +/* Transform init/exit (cra_init/cra_exit) */ + +static int cmh_hmac_cra_init(struct crypto_tfm *tfm) +{ + struct crypto_ahash *ahash =3D __crypto_ahash_cast(tfm); + struct cmh_hmac_tfm_ctx *tctx =3D crypto_tfm_ctx(tfm); + struct crypto_ahash *fb; + + memset(tctx, 0, sizeof(*tctx)); + tctx->key.mode =3D CMH_KEY_NONE; + spin_lock_init(&tctx->chunk_lock); + INIT_LIST_HEAD(&tctx->all_chunks); + + /* + * Allocate the generic software fallback used when the HW input cap + * is exceeded or an oversized clone is exported. Masking out + * CRYPTO_ALG_ASYNC excludes this (async) driver, so the allocator + * picks the generic hmac; its request is embedded after the reqctx. + */ + fb =3D crypto_alloc_ahash(crypto_ahash_alg_name(ahash), 0, + CRYPTO_ALG_ASYNC); + if (IS_ERR(fb)) + return PTR_ERR(fb); + tctx->fb =3D fb; + + /* + * The FB-format export copies the fallback's state into the flat + * window (state->data, CMH_HMAC_EXPORT_MAX bytes). If the fallback's + * statesize exceeds that, crypto_ahash_export() would overrun the + * driver's export buffer -- refuse to instantiate instead. + */ + if (crypto_ahash_statesize(fb) > CMH_HMAC_EXPORT_MAX) { + crypto_free_ahash(fb); + tctx->fb =3D NULL; + return -EINVAL; + } + + crypto_ahash_set_reqsize(ahash, + sizeof(struct cmh_hmac_reqctx) + + crypto_tfm_ctx_alignment() + + sizeof(struct ahash_request) + + crypto_ahash_reqsize(fb)); + return 0; +} + +static void cmh_hmac_cra_exit(struct crypto_tfm *tfm) +{ + struct cmh_hmac_tfm_ctx *tctx =3D crypto_tfm_ctx(tfm); + struct cmh_hmac_chunk *chunk, *tmp; + + /* Free any orphaned chunks (e.g. testmgr export/reimport poison) */ + spin_lock_bh(&tctx->chunk_lock); + list_for_each_entry_safe(chunk, tmp, &tctx->all_chunks, tfm_node) { + list_del(&chunk->tfm_node); + tctx->tfm_buffered -=3D chunk->len; + kfree_sensitive(chunk); + } + spin_unlock_bh(&tctx->chunk_lock); + + if (tctx->fb) + crypto_free_ahash(tctx->fb); + cmh_key_destroy(&tctx->key); +} + +/* Registration */ + +static struct cmh_hmac_alg_drv cmh_hmac_drvs[CMH_HMAC_ALG_COUNT]; + +/** + * cmh_hmac_register() - Register HMAC-SHA hash algorithms with the crypto= framework + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_hmac_register(void) +{ + unsigned int i; + int ret; + + if (!cmh_core_present(CMH_CORE_HC)) + return 0; + + for (i =3D 0; i < CMH_HMAC_ALG_COUNT; i++) { + const struct cmh_hmac_alg_info *info =3D &cmh_hmac_algs_info[i]; + struct cmh_hmac_alg_drv *drv =3D &cmh_hmac_drvs[i]; + struct ahash_alg *alg =3D &drv->alg; + + drv->info =3D info; + + alg->init =3D cmh_hmac_init; + alg->update =3D cmh_hmac_update; + alg->final =3D cmh_hmac_final; + alg->finup =3D cmh_hmac_finup; + alg->digest =3D cmh_hmac_digest; + alg->export =3D cmh_hmac_export; + alg->import =3D cmh_hmac_import; + alg->setkey =3D cmh_hmac_setkey; + + alg->halg.digestsize =3D info->digest_size; + alg->halg.statesize =3D CMH_HMAC_STATE_SIZE; + + strscpy(alg->halg.base.cra_name, info->alg_name, + CRYPTO_MAX_ALG_NAME); + strscpy(alg->halg.base.cra_driver_name, info->drv_name, + CRYPTO_MAX_ALG_NAME); + alg->halg.base.cra_priority =3D 300; + alg->halg.base.cra_flags =3D CRYPTO_ALG_KERN_DRIVER_ONLY | + CRYPTO_ALG_NO_FALLBACK | + CRYPTO_ALG_ASYNC | + CRYPTO_ALG_REQ_VIRT; + alg->halg.base.cra_blocksize =3D info->block_size; + alg->halg.base.cra_ctxsize =3D sizeof(struct cmh_hmac_tfm_ctx); + alg->halg.base.cra_init =3D cmh_hmac_cra_init; + alg->halg.base.cra_exit =3D cmh_hmac_cra_exit; + alg->halg.base.cra_module =3D THIS_MODULE; + + ret =3D crypto_register_ahash(alg); + if (ret) { + dev_err(cmh_dev(), "hmac: failed to register %s (rc=3D%d)\n", + info->drv_name, ret); + while (i--) + crypto_unregister_ahash(&cmh_hmac_drvs[i].alg); + return ret; + } + + dev_dbg(cmh_dev(), "hmac: registered %s (priority 300)\n", + info->drv_name); + } + + return 0; +} + +/** + * cmh_hmac_unregister() - Unregister HMAC-SHA hash algorithms from the cr= ypto framework + */ +void cmh_hmac_unregister(void) +{ + unsigned int i; + + if (!cmh_core_present(CMH_CORE_HC)) + return; + + for (i =3D 0; i < CMH_HMAC_ALG_COUNT; i++) { + crypto_unregister_ahash(&cmh_hmac_drvs[i].alg); + dev_dbg(cmh_dev(), "hmac: unregistered %s\n", + cmh_hmac_algs_info[i].drv_name); + } +} diff --git a/drivers/crypto/cmh/cmh_main.c b/drivers/crypto/cmh/cmh_main.c index 50218f19ad6f..589363cc4b28 100644 --- a/drivers/crypto/cmh/cmh_main.c +++ b/drivers/crypto/cmh/cmh_main.c @@ -32,6 +32,7 @@ #include "cmh_txn.h" #include "cmh_rh.h" #include "cmh_hash.h" +#include "cmh_hmac.h" #include "cmh_mgmt.h" #include "cmh_registers.h" #include "cmh_debugfs.h" @@ -208,6 +209,11 @@ static int cmh_probe(struct platform_device *pdev) if (ret) goto err_hash_register; =20 + /* Register HMAC hash algorithms */ + ret =3D cmh_hmac_register(); + if (ret) + goto err_hmac_register; + /* Register key management device (/dev/cmh_mgmt) */ ret =3D cmh_mgmt_register(); if (ret) @@ -218,6 +224,8 @@ static int cmh_probe(struct platform_device *pdev) return 0; =20 err_mgmt_register: + cmh_hmac_unregister(); +err_hmac_register: cmh_hash_unregister(); err_hash_register: cmh_rh_cleanup(cfg); @@ -246,6 +254,7 @@ static void cmh_remove(struct platform_device *pdev) cfg =3D &dev->config; =20 cmh_mgmt_unregister(); + cmh_hmac_unregister(); cmh_hash_unregister(); cmh_rh_cleanup(cfg); cmh_tm_cleanup(); diff --git a/drivers/crypto/cmh/include/cmh_hmac.h b/drivers/crypto/cmh/inc= lude/cmh_hmac.h new file mode 100644 index 000000000000..fb1a11fb76eb --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_hmac.h @@ -0,0 +1,16 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Kernel Crypto API HMAC Driver + * + * Registers HMAC ahash algorithms (HMAC-SHA-2, HMAC-SHA-3) with the + * Linux crypto subsystem using HC_CMD_HMAC. + */ + +#ifndef CMH_HMAC_H +#define CMH_HMAC_H + +int cmh_hmac_register(void); +void cmh_hmac_unregister(void); + +#endif /* CMH_HMAC_H */ --=20 2.43.7 From nobody Fri Sep 25 01:20:34 2026 Received: from BL2PR02CU003.outbound.protection.outlook.com (mail-eastusazon11021110.outbound.protection.outlook.com [52.101.52.110]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D00734B1CF3; Thu, 17 Sep 2026 22:59:43 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.52.110 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789685991; cv=fail; b=BUfyY5xK1I6DJFDWkRsPKV+4RcbW7c+QWm+X11df7NI9bsu6DkG8XfBJY5vkyCK+xP13nMWvM5OyC1lA4baS232pcFMecLiHSjpovUXsBcHjQ1JN7JBf0bFbWpmfUUb/20MdC26UrbVvP8PC81v2UbAq7OxI1JOck56GHdRflGs= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789685991; c=relaxed/simple; bh=7YHKUA+DGe1uuqaNhk8GA+c1L59kcm3dlfLIgZnfeLI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=uY8uUNDJ9a2bsmJDRV68OWSpHTfcm3DcH+k399uSQA0UNtiubMFqjxXn+0rPZAypohJdxFvL/HzPjLHY+8pej9GmHOxPW1ByndPLCt7Yu77jdxs09a/9SMeFUyDh4w33o6IXBnvssa/L8bXUA7fOW6aOQ6q3ehBbhkBkAPJAUF0= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com; spf=fail smtp.mailfrom=rambus.com; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b=nYEfYa/z; arc=fail smtp.client-ip=52.101.52.110 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=rambus.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b="nYEfYa/z" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=UpkK5RdNgE/FyW32UGd0fBo7ZYifsiZ7oVJY7oPjly7s6g9fEv6VFRaylHoW8TGl1XFtvuDreANY3Wa+OZYFEkE35Vms3fjOHNrwKsdhWIy6jN47WH7tkEFBC6rZQwVZt531AXK4WQRxH4bVGEDfLF4oljhLyRuIwp4NeEaysSjVvVu39hcH/ee8BCj36jK+o9pz4eZF90uUr4PiToG2m1pjiuSBWdvm9JX3px8o+Q19GfWTNNqQ3ZSeMJxTIdz0vZM1YhoW6E++CzRrFAqMQr+DmF15P4ipe2/R/Hk6eEPhrydRrbI9q6tR3K+ZCftfwGcj8j0ubl4eq0lW64d1rw== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=nCD5w+Gpy7DCGtPstZab3NCNkV4UwKYMvaiiGog4Odk=; b=oFEDuiufxpCeVTBSxFEc6OosktmLRqMN1eui68kN63vsIK/wL2bUVMJqM7kzXWnTKXaY9fhttWrge/fe7vMy7t08etxBjC6rZWF1xD4ndjP5YtsX+AMC/BRKUAU5rNY9NtkiGGA6hPKpr39WxSGKQRWaEfHLLl3LBSjyKpTvEUWfjokqLpallU0dRkozT8KMY1YbLf9pdGh0PlejRD+6kqhhv+p7MLRURUPE/Qhbff9QFAd5EiOm10E8N4oGMmTFD0uRg5B0e+/FkafQLD3XUSYkeZYrTnMH5nctlXBC4uzaI/zSEBHId8kcwD3eJwAIUtrhTS/BkzX6gqbBLzwm+g== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 192.86.86.210) smtp.rcpttodomain=cryptography.com smtp.mailfrom=rambus.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=rambus.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=rambus.com; s=selector1; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=nCD5w+Gpy7DCGtPstZab3NCNkV4UwKYMvaiiGog4Odk=; b=nYEfYa/zmpgdjNLFm3zp8+9gNM1ctVXUZaj64FHkUznXYtnUktTktDl+TihD7jiHbqPzKsX+NZOooEblk+OryQUxix75MO+Af/wrQZoyHGHE8qQGMM7mxtd2kHyVvSuZGT4qQ1TJt6RcMuMI96zFNiQjko22J/odx4nPZQhKGX5b71tm9qtZKgjRqYe3kZt7Yi1xGsMGz1WVjnKGVOAzjEKWVP3BnH00p9/IvXRx5XpM97AfvjpMkEKZ/uNU26htxWIR/fpzOGCXLoKinRy+joKmU+APZ6x7SV1bapDl6+kMlozfEweH7NGHVboCeH9gtHVQjKf4LocXxdfh72H63w== Received: from MN2PR14CA0023.namprd14.prod.outlook.com (2603:10b6:208:23e::28) by SJ0PR04MB7647.namprd04.prod.outlook.com (2603:10b6:a03:32f::8) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.428.13; Thu, 17 Sep 2026 22:59:34 +0000 Received: from BL6PEPF00020E63.namprd04.prod.outlook.com (2603:10b6:208:23e:cafe::14) by MN2PR14CA0023.outlook.office365.com (2603:10b6:208:23e::28) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.428.12 via Frontend Transport; Thu, 17 Sep 2026 22:59:33 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 192.86.86.210) smtp.mailfrom=rambus.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=rambus.com; Received-SPF: Pass (protection.outlook.com: domain of rambus.com designates 192.86.86.210 as permitted sender) receiver=protection.outlook.com; client-ip=192.86.86.210; helo=hqxsv-psmtppxy02.rambus.com; pr=C Received: from hqxsv-psmtppxy02.rambus.com (192.86.86.210) by BL6PEPF00020E63.mail.protection.outlook.com (10.167.249.24) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.451.8 via Frontend Transport; Thu, 17 Sep 2026 22:59:33 +0000 Received: from hqxsv-cmdev3-aousherovitch.rambus.com (hqn-lb-int-float.rambus.com [10.12.20.20]) by hqxsv-psmtppxy02.rambus.com (Postfix) with ESMTP id 0ABBB1801761; Thu, 17 Sep 2026 22:59:32 +0000 (UTC) From: Alex Ousherovitch To: Albert Ou , Alex Ousherovitch , Conor Dooley , "David S. Miller" , Herbert Xu , Jonathan Corbet , Krzysztof Kozlowski , Palmer Dabbelt , Paul Walmsley , Rob Herring , Saravanakrishnan Krishnamoorthy , Shuah Khan Cc: Alexandre Ghiti , devicetree@vger.kernel.org, Joel Wittenauer , linux-api@vger.kernel.org, linux-crypto@vger.kernel.org, linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-riscv@lists.infradead.org, Shuah Khan , Thi Nguyen Subject: [PATCH v5 06/19] crypto: cmh - add CSHAKE/KMAC ahash Date: Thu, 17 Sep 2026 15:59:15 -0700 Message-ID: <20260917225929.2494111-7-aousherovitch@rambus.com> X-Mailer: git-send-email 2.43.7 In-Reply-To: <20260917225929.2494111-1-aousherovitch@rambus.com> References: <20260917225929.2494111-1-aousherovitch@rambus.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: BL6PEPF00020E63:EE_|SJ0PR04MB7647:EE_ X-MS-Office365-Filtering-Correlation-Id: 75e3572e-76cd-487c-032d-08df150f56d2 X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|1800799024|23010399003|376014|7416014|82310400026|36860700016|921020|22082099003|18002099003|56012099006|3023799007|11063799006|5023799004|10067099003|6133799003; X-Microsoft-Antispam-Message-Info: qd5N125H+fhMJSjH026eciV9Y+33nSxK3IK3dFzYqi/wzoO0N7y9lg1RFj/v8XrDmX21s3Zwiqyildz2Laj4ZF5HFoCadlHqCpOo49oiL/luFjClTn7vebf8uw2QNd+1giJcbE0aVHBdfbAnsgpmyxj+9y01k26iFITi8d+3n5NmHoVJIVzL23ScD3cfReWvrybDWTQ7Dd2BLZOtJ7oJDtih7XtiP5s5i11XZ90wXniIaVN6JiGw5KNCxEWXf5tm11zPgoIMi7uMo2d6aPmOa6naXt25irEwYVoD2oO3btefKVRbF6s3/EuqEPtskFI3vORFozlZnLHacoHIg8/zpxkreuqmWD0zQEbSplP3JDyKb4hTExQiDWlg2UlNLofIrQWNwyPyW7gubfoBT0RcS25YlnPFAKYu8CFL2GyhZpR36ybTPqdgRKF1i8uG/CsLgtGe5OffDRiRQM98YU5N1EZ8fthGN68H7jhQ0S4jI9KiXK5tJbPhlXy5qi53P83ZiGgO70MBNA5hDwE2Bag7ir56NHgUBGbeldCjUlM51KbBDkShC+SNQq4rA5jzeE7W4n9zRi1W+YS1cAlNzl0J13iVZxOMvLhh8z4XFXsm5mTrxn7HlNMFGN65/8qzX92jtiiVekl939YUHY/pfhAx5tq5OaJvh5Ux28zEy1x5okQeijyb6oEboHmjfUzoR0GwjIc/46SZdYdz2usYrjUxZEsvNV8GMAHa5T3nw9LTJVy995ouMSde9rpuBE5/ekQp X-Forefront-Antispam-Report: CIP:192.86.86.210;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:hqxsv-psmtppxy02.rambus.com;PTR:InfoDomainNonexistent;CAT:NONE;SFS:(13230040)(1800799024)(23010399003)(376014)(7416014)(82310400026)(36860700016)(921020)(22082099003)(18002099003)(56012099006)(3023799007)(11063799006)(5023799004)(10067099003)(6133799003);DIR:OUT;SFP:1102; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: VaV7KVYmWPDu4+mH5HkxUK7xB71af4S1r0vGL7M7nGeKrVzYy4kT5DRNwmJQOtfCiyZDAHyfkmyyRfdpT1GBOVVWzr1VVh+MvR9L2VEHhLQfPfQnEzY+GyUuitTaHkE0NwB41Jcip4zwq7I5SBeh+Xs7t9pgUw1e4YfD6E3oK9Ne66WD2/ud9eg7iuMQix67ZLaC6aNqXBe3KgROd1R6i2CH7lLaextPT9sonIpmfwSue38y6KTsHka1h/leFxjrbvBj4gm6qK5FS239g7qOxuDFSCZo78LYciIsgj1vVgmjX8AL1Dmwy5rlgd/VcIwrPycoaNENBWTRp2PhDkT38F7q6s3Bdg3JaH9YQqgaoMYP4N6cf0eTw43DmkWq/qRXYxR7n3wmtLy89OebK1Tu7ids+sULDu03ji9wAepoZ9tHXksimzNIdgWs4LliJXut X-OriginatorOrg: rambus.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 17 Sep 2026 22:59:33.2858 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: 75e3572e-76cd-487c-032d-08df150f56d2 X-MS-Exchange-CrossTenant-Id: bd0ba799-c2b9-413c-9c56-5d1731c4827c X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=bd0ba799-c2b9-413c-9c56-5d1731c4827c;Ip=[192.86.86.210];Helo=[hqxsv-psmtppxy02.rambus.com] X-MS-Exchange-CrossTenant-AuthSource: BL6PEPF00020E63.namprd04.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: SJ0PR04MB7647 Content-Type: text/plain; charset="utf-8" Register ahash algorithms for cSHAKE128, cSHAKE256, KMAC128, and KMAC256 using the CMH hash core. cSHAKE is unkeyed, so like SHAKE it is CRYPTO_AHASH_ALG_BLOCK_ONLY and supports incremental update and export/import. KMAC is keyed: the hardware has no keyed-MAC save/restore, so (like HMAC) it is not block-only and accumulates the input for a single finalize-time submission. KMAC has no generic software equivalent, so its input is capped at the 64KB hardware window. Signed-off-by: Alex Ousherovitch Co-developed-by: Saravanakrishnan Krishnamoorthy Signed-off-by: Saravanakrishnan Krishnamoorthy --- drivers/crypto/cmh/Makefile | 4 +- drivers/crypto/cmh/cmh_cshake.c | 873 ++++++++++++++++++++++++ drivers/crypto/cmh/cmh_kmac.c | 690 +++++++++++++++++++ drivers/crypto/cmh/cmh_main.c | 18 + drivers/crypto/cmh/include/cmh_cshake.h | 16 + drivers/crypto/cmh/include/cmh_kmac.h | 16 + 6 files changed, 1616 insertions(+), 1 deletion(-) create mode 100644 drivers/crypto/cmh/cmh_cshake.c create mode 100644 drivers/crypto/cmh/cmh_kmac.c create mode 100644 drivers/crypto/cmh/include/cmh_cshake.h create mode 100644 drivers/crypto/cmh/include/cmh_kmac.h diff --git a/drivers/crypto/cmh/Makefile b/drivers/crypto/cmh/Makefile index acd1827cc084..3ab154744658 100644 --- a/drivers/crypto/cmh/Makefile +++ b/drivers/crypto/cmh/Makefile @@ -16,7 +16,9 @@ cmh-y :=3D \ cmh_key.o \ cmh_sys.o \ cmh_hash.o \ - cmh_hmac.o + cmh_hmac.o \ + cmh_cshake.o \ + cmh_kmac.o =20 # Management ioctl device (/dev/cmh_mgmt): key lifecycle, PKE, PQC ioctls. cmh-$(CONFIG_CRYPTO_DEV_CMH_MGMT) +=3D \ diff --git a/drivers/crypto/cmh/cmh_cshake.c b/drivers/crypto/cmh/cmh_cshak= e.c new file mode 100644 index 000000000000..56a6453a0e4f --- /dev/null +++ b/drivers/crypto/cmh/cmh_cshake.c @@ -0,0 +1,873 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Kernel Crypto API CSHAKE Driver + * + * Registers cSHAKE-128 and cSHAKE-256 as ahash algorithms using the + * CMH Hash Core (HC) via HC_CMD_CSHAKE. + * + * CSHAKE (NIST SP 800-185) extends SHAKE with two domain separation + * parameters: function name N and customization string S. When both + * are empty, cSHAKE reduces to plain SHAKE -- the driver falls back to + * HC_CMD_INIT in that case (per SP 800-185 S6.2). + * + * N and S are set via .setkey() using a self-describing binary header + * (matching the upstream authenc precedent): + * + * struct cshake_cfg { __be32 n_len; __be32 s_len; }; + * setkey blob: cshake_cfg || N[n_len] || S[s_len] + * + * If .setkey() is never called, the driver defaults to plain SHAKE + * (N=3D"" S=3D""). .setkey() is per-tfm, not per-request. + * + * N is embedded inline in the HC_CMD_CSHAKE struct (max 36 bytes). + * S is passed as VCQ inline data following the command slot (multi-span). + * + * Uses the same streaming transaction model as cmh_hash.c. cSHAKE is a + * sponge XOF, so cra_blocksize is 1 (the Keccak rate 168/136 exceeds + * MAX_ALGAPI_BLOCKSIZE). With CRYPTO_AHASH_ALG_BLOCK_ONLY and blocksize + * 1 the Crypto API holds nothing back: .update() absorbs the whole + * chunk and the HC core buffers any sub-rate remainder in its SAVEd + * context across SAVE/RESTORE, so cSHAKE hashes and clones at any length + * with bounded kernel memory. + * .init() -> software-only + * .update() -> [first: CSHAKE(+S) | resume: CSHAKE(+S)/INIT + RESTORE] + * + UPDATE(all data) + SAVE + FLUSH + * .finup() -> [first: CSHAKE(+S) | resume: CSHAKE(+S)/INIT + RESTORE] + * [+ UPDATE(residual)] + FINAL + FLUSH (also serves .final) + * .export()/.import() -> serialise the HC checkpoint only; + * the API appends its own partial-block buffer + * + * The cSHAKE prefix (function name N, customization string S) is + * absorbed once by HC_CMD_CSHAKE on the first submission (bytepad-ded to + * a full rate block, so message absorption stays rate-aligned) and is + * carried thereafter in the SAVEd sponge state; resume re-establishes + * the mode (CSHAKE for keyed N/S, INIT for plain SHAKE) before + * HC_CMD_RESTORE. The HC core supports HC_CMD_SAVE / HC_CMD_RESTORE for + * SHAKE/cSHAKE (outlen stays 0, unlike KMAC), which is what enables both + * streaming and transform cloning. This is an sg-only driver (no + * CRYPTO_ALG_REQ_VIRT): BLOCK_ONLY buffer prepending assumes scatterlists. + * + * .setkey() here configures public domain-separation parameters (N, S), + * not a secret key. + */ + +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "cmh_cshake.h" +#include "cmh_vcq.h" +#include "cmh_hc_abi.h" +#include "cmh_txn.h" +#include "cmh_dma.h" + +/* Algorithm Table */ + +struct cmh_cshake_alg_info { + u32 hc_algo; + u32 digest_size; + u32 block_size; /* cra_blocksize (XOF: 1) */ + const char *alg_name; + const char *drv_name; +}; + +static const struct cmh_cshake_alg_info cmh_cshake_algs_info[] =3D { + { + .hc_algo =3D HC_ALGO_SHAKE128, + .digest_size =3D CMH_SHAKE128_DIGEST_SIZE, + .block_size =3D 1, /* XOF */ + .alg_name =3D "cshake128", + .drv_name =3D "rambus-cmh-cshake128", + }, + { + .hc_algo =3D HC_ALGO_SHAKE256, + .digest_size =3D CMH_SHAKE256_DIGEST_SIZE, + .block_size =3D 1, /* XOF */ + .alg_name =3D "cshake256", + .drv_name =3D "rambus-cmh-cshake256", + }, +}; + +#define CMH_CSHAKE_ALG_COUNT ARRAY_SIZE(cmh_cshake_algs_info) + +/* Per-Request State */ + +/* + * Max payload slots for a streaming cSHAKE transaction. Worst case is a + * resume with the largest customization string: + * CSHAKE (1) + inline S (4) + RESTORE (1) + UPDATE (1) + SAVE/FINAL (1) + * + FLUSH (1) =3D 9 + */ +#define CMH_CSHAKE_MAX_PAYLOAD 9 +#define CMH_CSHAKE_MAX_PACKED (CMH_CSHAKE_MAX_PAYLOAD * 2) + +/* Fail the build if a large customization string would overflow cmds[]. */ +static_assert(CMH_CSHAKE_MAX_PAYLOAD >=3D + 1 + DIV_ROUND_UP(HC_CSHAKE_MAX_CUSTOMLEN, + sizeof(struct vcq_cmd)) + 4); + +/* + * Exported request state (statesize): the HC checkpoint from the last + * SAVE. The Crypto API appends its own partial-block buffer. Mirrors + * the unkeyed hash driver so cSHAKE streams and clones at any length. + */ +struct cmh_cshake_export_state { + u8 checkpoint[HC_CONTEXT_SIZE]; + u32 hw_started; +}; + +/* + * Stored in ahash_request_ctx(). The checkpoint is embedded inline + * (not heap): the kernel ahash API has no per-request destructor, so an + * abandoned request must not leak. + * + * Mapping the inline checkpoint for DMA is safe even on non-coherent + * platforms: it is only ever mapped DMA_TO_DEVICE and its bytes are + * frozen from dma_map_single() until the matching unmap (it is written + * solely in the completion, after the unmap). CPU writes to adjacent + * fields sharing a cacheline never alter the checkpoint bytes, and a + * TO_DEVICE unmap performs no cache invalidate; the shared-cacheline + * hazard applies only to FROM_DEVICE buffers, which are kmalloc'd. + */ +struct cmh_cshake_reqctx { + const struct cmh_cshake_alg_info *info; + int error; + u32 hw_started; /* non-zero after first HW submission */ + u32 has_checkpoint; /* non-zero if checkpoint[] valid */ + u32 update_remainder; /* sub-block bytes the API must re-buffer */ + /* DMA state for the current async operation */ + dma_addr_t ckpt_dma; /* RESTORE input */ + dma_addr_t save_dma; /* SAVE output (update only) */ + dma_addr_t data_dma; /* UPDATE input */ + dma_addr_t digest_dma; /* FINAL output (final/digest only) */ + u8 *save_buf; + u8 *data_buf; + u32 data_len; + u8 *digest_buf; + u8 checkpoint[HC_CONTEXT_SIZE]; /* HC context from last SAVE */ + struct vcq_cmd packed[CMH_CSHAKE_MAX_PACKED]; +}; + +/* Per-Transform State (carries N and S across requests) */ + +struct cmh_cshake_tfm_ctx { + u8 *func_name; /* N (function name), NULL if empty */ + u32 func_name_len; + u8 *custom; /* S (customization string), NULL if empty */ + u32 custom_len; +}; + +/* VCQ Builders */ + +/* VCQ Builders (cSHAKE-specific; shared builders in cmh_hc_abi.h / cmh_vc= q.h) */ + +static void vcq_add_hc_save(struct vcq_cmd *slot, u32 core_id, + u64 output_phys, u32 outlen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, HC_CMD_SAVE); + slot->hwc.hc.cmd_save.output =3D output_phys; + slot->hwc.hc.cmd_save.outlen =3D outlen; +} + +static void vcq_add_hc_restore(struct vcq_cmd *slot, u32 core_id, + u64 input_phys, u32 inlen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, HC_CMD_RESTORE); + slot->hwc.hc.cmd_restore.input =3D input_phys; + slot->hwc.hc.cmd_restore.inlen =3D inlen; +} + +static void vcq_add_hc_cshake(struct vcq_cmd *slot, u32 core_id, u32 algo, + const u8 *name, u32 namelen, + u32 customlen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, HC_CMD_CSHAKE); + slot->hwc.hc.cmd_cshake.custom =3D 0; /* inline -- CMH eSW reads from ne= xt slot(s) */ + slot->hwc.hc.cmd_cshake.customlen =3D customlen; + slot->hwc.hc.cmd_cshake.algo =3D algo; + slot->hwc.hc.cmd_cshake.namelen =3D namelen; + if (namelen > 0 && name) + memcpy(slot->hwc.hc.cmd_cshake.name, name, + min_t(u32, namelen, HC_CSHAKE_MAX_NAMELEN)); +} + +/* Add an HC_CMD_UPDATE entry */ +static void vcq_add_hc_update(struct vcq_cmd *slot, u32 core_id, u64 input= _phys, u32 len) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, HC_CMD_UPDATE); + slot->hwc.hc.cmd_update.input =3D input_phys; + slot->hwc.hc.cmd_update.inlen =3D len; +} + +/* + * Emit the HC prologue that (re-)establishes the sponge mode before an + * UPDATE and the trailing SAVE/FINAL: + * cSHAKE (N/S set): CSHAKE [+ inline S] + * plain SHAKE: INIT + * followed by RESTORE when resuming a saved context. The mode command + * is re-issued on resume too: the eSW RESTORE rejects an algo_mode + * mismatch, and cSHAKE uses a distinct mode from INIT. RESTORE then + * overwrites the freshly re-absorbed prefix with the saved sponge state. + * Returns the next free slot index. + */ +static u32 cmh_cshake_emit_prologue(struct vcq_cmd *cmds, u32 idx, + const struct core_dispatch *d, + const struct cmh_cshake_reqctx *rctx, + const struct cmh_cshake_tfm_ctx *tctx, + dma_addr_t ckpt_dma) +{ + if (tctx->func_name_len > 0 || tctx->custom_len > 0) { + u32 span; + + vcq_add_hc_cshake(&cmds[idx], d->core_id, rctx->info->hc_algo, + tctx->func_name, tctx->func_name_len, + tctx->custom_len); + span =3D vcq_add_inline_data(&cmds[idx], tctx->custom, + tctx->custom_len); + idx +=3D span; + } else { + vcq_add_hc_init(&cmds[idx++], d->core_id, rctx->info->hc_algo); + } + + if (rctx->has_checkpoint) + vcq_add_hc_restore(&cmds[idx++], d->core_id, (u64)ckpt_dma, + HC_CONTEXT_SIZE); + + return idx; +} + +/* Reset per-request state (no heap held between operations). */ +static void cmh_cshake_free_reqctx(struct cmh_cshake_reqctx *rctx) +{ + rctx->has_checkpoint =3D 0; +} + +/* VCQ Packing + Submit */ + +/* ahash Operations */ + +struct cmh_cshake_alg_drv { + struct ahash_alg alg; + const struct cmh_cshake_alg_info *info; +}; + +static const struct cmh_cshake_alg_info * +cmh_cshake_get_info(struct crypto_ahash *tfm) +{ + struct ahash_alg *alg =3D crypto_ahash_alg(tfm); + + return container_of(alg, struct cmh_cshake_alg_drv, alg)->info; +} + +/* + * .setkey() -- parse N and S from the self-describing cshake_cfg header. + * + * Blob format: cshake_cfg { __be32 n_len; __be32 s_len; } || N || S + * If never called, the driver defaults to plain SHAKE (N=3D"" S=3D""). + */ +struct cshake_cfg { + __be32 n_len; + __be32 s_len; +}; + +static int cmh_cshake_setkey(struct crypto_ahash *tfm, const u8 *key, + unsigned int keylen) +{ + struct cmh_cshake_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cshake_cfg cfg; + u32 n_len, s_len; + const u8 *ptr; + + if (keylen < sizeof(cfg)) + return -EINVAL; + + memcpy(&cfg, key, sizeof(cfg)); + n_len =3D be32_to_cpu(cfg.n_len); + s_len =3D be32_to_cpu(cfg.s_len); + + if (keylen !=3D sizeof(cfg) + n_len + s_len) + return -EINVAL; + + if (n_len > HC_CSHAKE_MAX_NAMELEN) + return -EINVAL; + + if (s_len > HC_CSHAKE_MAX_CUSTOMLEN) + return -EINVAL; + + /* + * Free previous N and S. Unlocked against a concurrent .update()/ + * .final() that reads them: the crypto API does not issue setkey + * concurrently with an in-flight request on the same tfm, and the + * only frontend that allowed that race (AF_ALG setsockopt vs I/O) + * is not built on this kernel. + */ + kfree(tctx->func_name); + kfree(tctx->custom); + tctx->func_name =3D NULL; + tctx->func_name_len =3D 0; + tctx->custom =3D NULL; + tctx->custom_len =3D 0; + + ptr =3D key + sizeof(cfg); + + if (n_len > 0) { + tctx->func_name =3D kmemdup(ptr, n_len, GFP_KERNEL); + if (!tctx->func_name) + return -ENOMEM; + tctx->func_name_len =3D n_len; + ptr +=3D n_len; + } + + if (s_len > 0) { + tctx->custom =3D kmemdup(ptr, s_len, GFP_KERNEL); + if (!tctx->custom) { + kfree(tctx->func_name); + tctx->func_name =3D NULL; + tctx->func_name_len =3D 0; + return -ENOMEM; + } + tctx->custom_len =3D s_len; + } + + return 0; +} + +static int cmh_cshake_init(struct ahash_request *req) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_cshake_reqctx *rctx =3D ahash_request_ctx(req); + + memset(rctx, 0, sizeof(*rctx)); + rctx->info =3D cmh_cshake_get_info(tfm); + + return 0; +} + +/* + * Update completion -- runs from the threaded IRQ after SAVE. Takes the + * SAVEd context as the new checkpoint. + */ +static void cmh_cshake_update_complete(void *data, int error) +{ + struct ahash_request *req =3D data; + struct cmh_cshake_reqctx *rctx =3D ahash_request_ctx(req); + + if (error =3D=3D -EINPROGRESS) { + cmh_complete(&req->base, error); + return; + } + + if (rctx->has_checkpoint) + cmh_dma_unmap_single(rctx->ckpt_dma, HC_CONTEXT_SIZE, + DMA_TO_DEVICE); + cmh_dma_unmap_single(rctx->save_dma, HC_CONTEXT_SIZE, + DMA_FROM_DEVICE); + cmh_dma_unmap_single(rctx->data_dma, rctx->data_len, + DMA_TO_DEVICE); + + if (!error) { + memcpy(rctx->checkpoint, rctx->save_buf, HC_CONTEXT_SIZE); + rctx->has_checkpoint =3D 1; + rctx->hw_started =3D 1; + /* Hand the API the sub-block remainder it must re-buffer. */ + error =3D rctx->update_remainder; + } else { + rctx->error =3D error; + } + + kfree(rctx->save_buf); + rctx->save_buf =3D NULL; + kfree(rctx->data_buf); + rctx->data_buf =3D NULL; + rctx->data_len =3D 0; + + cmh_complete(&req->base, error); +} + +/* + * .update -- absorb the update data in hardware. + * + * cra_blocksize is 1, so the Crypto API hands over the whole update; it + * is submitted as: + * [first: CSHAKE(+S) | resume: CSHAKE(+S)/INIT + RESTORE] + UPDATE + + * SAVE + FLUSH + * The HC core buffers any sub-rate remainder into its SAVEd context, so + * nothing is held back to the API (the completion returns 0). + */ +static int cmh_cshake_update(struct ahash_request *req) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_cshake_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cmh_cshake_reqctx *rctx =3D ahash_request_ctx(req); + struct vcq_cmd cmds[CMH_CSHAKE_MAX_PAYLOAD]; + struct core_dispatch d; + u32 full_len; + u32 idx; + int ret; + gfp_t gfp; + + if (rctx->error) + return rctx->error; + + if (!req->nbytes) + return 0; + + /* XOF (blocksize 1): absorb everything; HC buffers the sub-rate tail. */ + full_len =3D req->nbytes; + rctx->update_remainder =3D 0; + + gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + + /* + * Reject a single update whose linearisation would exceed the largest + * kmalloc: return a permanent -EMSGSIZE ("message too long") rather + * than a transient -ENOMEM the client would keep retrying. + */ + if (full_len > KMALLOC_MAX_SIZE) + return -EMSGSIZE; + + rctx->data_buf =3D kmalloc(full_len, gfp | __GFP_NOWARN); + if (!rctx->data_buf) + return -ENOMEM; + + scatterwalk_map_and_copy(rctx->data_buf, req->src, 0, full_len, 0); + + rctx->save_buf =3D kzalloc(HC_CONTEXT_SIZE, gfp); + if (!rctx->save_buf) { + ret =3D -ENOMEM; + goto err_free; + } + + rctx->data_dma =3D cmh_dma_map_single(rctx->data_buf, full_len, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->data_dma)) { + ret =3D -ENOMEM; + goto err_free; + } + + rctx->save_dma =3D cmh_dma_map_single(rctx->save_buf, HC_CONTEXT_SIZE, + DMA_FROM_DEVICE); + if (cmh_dma_map_error(rctx->save_dma)) { + ret =3D -ENOMEM; + goto err_unmap_data; + } + + rctx->ckpt_dma =3D DMA_MAPPING_ERROR; + if (rctx->has_checkpoint) { + rctx->ckpt_dma =3D cmh_dma_map_single(rctx->checkpoint, + HC_CONTEXT_SIZE, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->ckpt_dma)) { + ret =3D -ENOMEM; + goto err_unmap_save; + } + } + + rctx->data_len =3D full_len; + + d =3D cmh_core_select_instance(CMH_CORE_HC); + idx =3D cmh_cshake_emit_prologue(cmds, 0, &d, rctx, tctx, + rctx->ckpt_dma); + + vcq_add_hc_update(&cmds[idx++], d.core_id, + (u64)rctx->data_dma, full_len); + vcq_add_hc_save(&cmds[idx++], d.core_id, + (u64)rctx->save_dma, HC_CONTEXT_SIZE); + vcq_add_flush(&cmds[idx++], d.core_id); + + ret =3D cmh_vcq_pack_and_submit_async(cmds, idx, rctx->packed, + CMH_CSHAKE_MAX_PACKED, + d.mbx_idx, + cmh_cshake_update_complete, req, + !!(req->base.flags & + CRYPTO_TFM_REQ_MAY_BACKLOG), + cmh_tm_async_timeout_jiffies()); + if (ret && ret !=3D -EBUSY) + goto err_unmap_ckpt; + + if (ret =3D=3D -EBUSY) + return -EBUSY; + return -EINPROGRESS; + +err_unmap_ckpt: + if (rctx->has_checkpoint) + cmh_dma_unmap_single(rctx->ckpt_dma, HC_CONTEXT_SIZE, + DMA_TO_DEVICE); +err_unmap_save: + cmh_dma_unmap_single(rctx->save_dma, HC_CONTEXT_SIZE, + DMA_FROM_DEVICE); +err_unmap_data: + cmh_dma_unmap_single(rctx->data_dma, full_len, DMA_TO_DEVICE); +err_free: + kfree(rctx->save_buf); + rctx->save_buf =3D NULL; + kfree(rctx->data_buf); + rctx->data_buf =3D NULL; + rctx->data_len =3D 0; + return ret; +} + +/* + * Final completion -- unmap DMA, copy digest, signal done. + */ +static void cmh_cshake_final_complete(void *data, int error) +{ + struct ahash_request *req =3D data; + struct cmh_cshake_reqctx *rctx =3D ahash_request_ctx(req); + + if (error =3D=3D -EINPROGRESS) { + cmh_complete(&req->base, error); + return; + } + + if (rctx->has_checkpoint) + cmh_dma_unmap_single(rctx->ckpt_dma, HC_CONTEXT_SIZE, + DMA_TO_DEVICE); + if (rctx->data_buf) + cmh_dma_unmap_single(rctx->data_dma, rctx->data_len, + DMA_TO_DEVICE); + cmh_dma_unmap_single(rctx->digest_dma, rctx->info->digest_size, + DMA_FROM_DEVICE); + + if (!error) + memcpy(req->result, rctx->digest_buf, + rctx->info->digest_size); + + kfree(rctx->digest_buf); + rctx->digest_buf =3D NULL; + kfree(rctx->data_buf); + rctx->data_buf =3D NULL; + cmh_cshake_free_reqctx(rctx); + cmh_complete(&req->base, error); +} + +/* + * Submit the final transaction: + * [first: CSHAKE(+S) | resume: INIT + RESTORE] [+ UPDATE(residual)] + * + FINAL + FLUSH + * + * @data_buf: linearised residual bytes, or NULL for empty input. + * Ownership transferred -- the callback frees it. + */ +static int cmh_cshake_submit_final(struct ahash_request *req, + u8 *data_buf, u32 data_len) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_cshake_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cmh_cshake_reqctx *rctx =3D ahash_request_ctx(req); + const struct cmh_cshake_alg_info *info =3D rctx->info; + struct vcq_cmd cmds[CMH_CSHAKE_MAX_PAYLOAD]; + struct core_dispatch d; + u32 idx; + int ret; + gfp_t gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + + rctx->data_buf =3D data_buf; + rctx->data_len =3D data_len; + + rctx->digest_buf =3D kzalloc(info->digest_size, gfp); + if (!rctx->digest_buf) { + ret =3D -ENOMEM; + goto err_free_data; + } + + rctx->digest_dma =3D cmh_dma_map_single(rctx->digest_buf, + info->digest_size, + DMA_FROM_DEVICE); + if (cmh_dma_map_error(rctx->digest_dma)) { + ret =3D -ENOMEM; + goto err_free_digest; + } + + rctx->data_dma =3D DMA_MAPPING_ERROR; + if (data_buf && data_len > 0) { + rctx->data_dma =3D cmh_dma_map_single(data_buf, data_len, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->data_dma)) { + ret =3D -ENOMEM; + goto err_unmap_digest; + } + } + + rctx->ckpt_dma =3D DMA_MAPPING_ERROR; + if (rctx->has_checkpoint) { + rctx->ckpt_dma =3D cmh_dma_map_single(rctx->checkpoint, + HC_CONTEXT_SIZE, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->ckpt_dma)) { + ret =3D -ENOMEM; + goto err_unmap_data; + } + } + + d =3D cmh_core_select_instance(CMH_CORE_HC); + idx =3D cmh_cshake_emit_prologue(cmds, 0, &d, rctx, tctx, + rctx->ckpt_dma); + + if (data_buf && data_len > 0) + vcq_add_hc_update(&cmds[idx++], d.core_id, + (u64)rctx->data_dma, data_len); + + vcq_add_hc_final(&cmds[idx++], d.core_id, + (u64)rctx->digest_dma, info->digest_size); + vcq_add_flush(&cmds[idx++], d.core_id); + + ret =3D cmh_vcq_pack_and_submit_async(cmds, idx, rctx->packed, + CMH_CSHAKE_MAX_PACKED, + d.mbx_idx, + cmh_cshake_final_complete, req, + !!(req->base.flags & + CRYPTO_TFM_REQ_MAY_BACKLOG), + cmh_tm_async_timeout_jiffies()); + if (ret =3D=3D -EBUSY) + return -EBUSY; + if (ret) + goto err_unmap_ckpt; + + return -EINPROGRESS; + +err_unmap_ckpt: + if (rctx->has_checkpoint) + cmh_dma_unmap_single(rctx->ckpt_dma, HC_CONTEXT_SIZE, + DMA_TO_DEVICE); +err_unmap_data: + if (data_buf && data_len > 0) + cmh_dma_unmap_single(rctx->data_dma, data_len, + DMA_TO_DEVICE); +err_unmap_digest: + cmh_dma_unmap_single(rctx->digest_dma, info->digest_size, + DMA_FROM_DEVICE); +err_free_digest: + kfree(rctx->digest_buf); + rctx->digest_buf =3D NULL; +err_free_data: + kfree(data_buf); + rctx->data_buf =3D NULL; + /* + * Preserve the HC checkpoint on failure: a synchronous rejection is + * retryable, and for a terminal error the inline checkpoint is freed + * with the request context, so it never leaks. It is cleared only in + * the completion after a successful final(). + */ + return ret; +} + +/* + * .finup -- hash any remaining data and finalise in one transaction. + * With BLOCK_ONLY the Crypto API prepends the bytes it held back, so + * req->src already carries the full tail. Also serves .final (nbytes + * =3D=3D 0). Avoids ahash_def_finup(), which would clone via export/impo= rt. + */ +static int cmh_cshake_finup(struct ahash_request *req) +{ + struct cmh_cshake_reqctx *rctx =3D ahash_request_ctx(req); + u32 data_len =3D req->nbytes; + u8 *data_buf =3D NULL; + gfp_t gfp; + + if (rctx->error) + return rctx->error; + + if (data_len =3D=3D 0) + return cmh_cshake_submit_final(req, NULL, 0); + + /* Reject an oversized linearisation with a permanent -EMSGSIZE. */ + if (data_len > KMALLOC_MAX_SIZE) + return -EMSGSIZE; + + gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + + data_buf =3D kmalloc(data_len, gfp | __GFP_NOWARN); + if (!data_buf) + return -ENOMEM; + + scatterwalk_map_and_copy(data_buf, req->src, 0, data_len, 0); + + return cmh_cshake_submit_final(req, data_buf, data_len); +} + +static int cmh_cshake_digest(struct ahash_request *req) +{ + int ret; + + ret =3D cmh_cshake_init(req); + if (ret) + return ret; + + return cmh_cshake_finup(req); +} + +/* + * Export core -- purely software. Serialise the HC checkpoint (if any); + * the Crypto API appends its own partial-block buffer. The streaming + * .update() keeps the checkpoint current after each block. + */ +static int cmh_cshake_export(struct ahash_request *req, void *out) +{ + struct cmh_cshake_reqctx *rctx =3D ahash_request_ctx(req); + struct cmh_cshake_export_state *state =3D out; + + /* + * Zero the whole state first: the struct may carry padding, so this + * avoids leaking kernel memory through the ahash export. + */ + memset(state, 0, sizeof(*state)); + + if (rctx->hw_started) + memcpy(state->checkpoint, rctx->checkpoint, HC_CONTEXT_SIZE); + + state->hw_started =3D rctx->hw_started; + + return 0; +} + +/* + * Import core -- purely software. Restore the HC checkpoint; the next + * .update()/final op RESTOREs it into HW. The Crypto API restores its + * own partial-block buffer separately. + */ +static int cmh_cshake_import(struct ahash_request *req, const void *in) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_cshake_reqctx *rctx =3D ahash_request_ctx(req); + const struct cmh_cshake_export_state *state =3D in; + + memset(rctx, 0, sizeof(*rctx)); + rctx->info =3D cmh_cshake_get_info(tfm); + + rctx->hw_started =3D state->hw_started; + + if (state->hw_started) { + memcpy(rctx->checkpoint, state->checkpoint, HC_CONTEXT_SIZE); + rctx->has_checkpoint =3D 1; + } + + return 0; +} + +/* Transform init/exit */ + +static int cmh_cshake_cra_init(struct crypto_tfm *tfm) +{ + struct cmh_cshake_tfm_ctx *tctx =3D crypto_tfm_ctx(tfm); + + tctx->func_name =3D NULL; + tctx->func_name_len =3D 0; + tctx->custom =3D NULL; + tctx->custom_len =3D 0; + return 0; +} + +static void cmh_cshake_cra_exit(struct crypto_tfm *tfm) +{ + struct cmh_cshake_tfm_ctx *tctx =3D crypto_tfm_ctx(tfm); + + kfree(tctx->func_name); + kfree(tctx->custom); + tctx->func_name =3D NULL; + tctx->custom =3D NULL; +} + +/* Registration */ + +static struct cmh_cshake_alg_drv cmh_cshake_drvs[CMH_CSHAKE_ALG_COUNT]; + +/** + * cmh_cshake_register() - Register cSHAKE-128/256 hash algorithms with th= e crypto framework + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_cshake_register(void) +{ + unsigned int i; + int ret; + + if (!cmh_core_present(CMH_CORE_HC)) + return 0; + + for (i =3D 0; i < CMH_CSHAKE_ALG_COUNT; i++) { + const struct cmh_cshake_alg_info *info =3D + &cmh_cshake_algs_info[i]; + struct cmh_cshake_alg_drv *drv =3D &cmh_cshake_drvs[i]; + struct ahash_alg *alg =3D &drv->alg; + + drv->info =3D info; + + alg->init =3D cmh_cshake_init; + alg->update =3D cmh_cshake_update; + alg->finup =3D cmh_cshake_finup; + alg->digest =3D cmh_cshake_digest; + alg->export =3D cmh_cshake_export; + alg->import =3D cmh_cshake_import; + alg->setkey =3D cmh_cshake_setkey; + + alg->halg.digestsize =3D info->digest_size; + alg->halg.statesize =3D sizeof(struct cmh_cshake_export_state); + + strscpy(alg->halg.base.cra_name, info->alg_name, + CRYPTO_MAX_ALG_NAME); + strscpy(alg->halg.base.cra_driver_name, info->drv_name, + CRYPTO_MAX_ALG_NAME); + alg->halg.base.cra_priority =3D 300; + alg->halg.base.cra_flags =3D CRYPTO_ALG_KERN_DRIVER_ONLY | + CRYPTO_ALG_NO_FALLBACK | + CRYPTO_ALG_ASYNC | + CRYPTO_ALG_OPTIONAL_KEY | + CRYPTO_AHASH_ALG_BLOCK_ONLY; + alg->halg.base.cra_blocksize =3D info->block_size; /* XOF: 1 */ + alg->halg.base.cra_ctxsize =3D sizeof(struct cmh_cshake_tfm_ctx); + alg->halg.base.cra_reqsize =3D sizeof(struct cmh_cshake_reqctx); + alg->halg.base.cra_init =3D cmh_cshake_cra_init; + alg->halg.base.cra_exit =3D cmh_cshake_cra_exit; + alg->halg.base.cra_module =3D THIS_MODULE; + + ret =3D crypto_register_ahash(alg); + if (ret) { + dev_err(cmh_dev(), "cshake: failed to register %s (rc=3D%d)\n", + info->drv_name, ret); + while (i--) + crypto_unregister_ahash(&cmh_cshake_drvs[i].alg); + return ret; + } + + dev_dbg(cmh_dev(), "cshake: registered %s (priority 300)\n", + info->drv_name); + } + + return 0; +} + +/** + * cmh_cshake_unregister() - Unregister cSHAKE hash algorithms from the cr= ypto framework + */ +void cmh_cshake_unregister(void) +{ + unsigned int i; + + if (!cmh_core_present(CMH_CORE_HC)) + return; + + for (i =3D 0; i < CMH_CSHAKE_ALG_COUNT; i++) { + crypto_unregister_ahash(&cmh_cshake_drvs[i].alg); + dev_dbg(cmh_dev(), "cshake: unregistered %s\n", + cmh_cshake_algs_info[i].drv_name); + } +} diff --git a/drivers/crypto/cmh/cmh_kmac.c b/drivers/crypto/cmh/cmh_kmac.c new file mode 100644 index 000000000000..97ff36af13a0 --- /dev/null +++ b/drivers/crypto/cmh/cmh_kmac.c @@ -0,0 +1,690 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Kernel Crypto API KMAC Driver + * + * Registers KMAC-128 and KMAC-256 as keyed ahash algorithms using the + * CMH Hash Core (HC) via HC_CMD_KMAC. + * + * KMAC (NIST SP 800-185) is a keyed variant of cSHAKE. The function + * name N is always "KMAC" (hardcoded by the CMH eSW). The user sets: + * - A key via .setkey() (raw bytes + optional S) + * - An optional customization string S via the setkey blob + * + * setkey blob format: + * struct kmac_key_param { __be32 keylen; __be32 s_len; }; + * blob: kmac_key_param || key[keylen] || S[s_len] + * + * Uses the same self-contained transaction model as cmh_hmac.c: + * .setkey() -> store raw key (+ S) + * .init() -> software-only + * .update() -> software-only (accumulate chunks) + * .final() -> [SYS_CMD_WRITE] + HC_CMD_KMAC [+ inline S] + + * [GATHER] + FINAL + FLUSH + */ + +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "cmh_kmac.h" +#include "cmh_vcq.h" +#include "cmh_hc_abi.h" +#include "cmh_sys_abi.h" +#include "cmh_sys.h" +#include "cmh_txn.h" +#include "cmh_dma.h" +#include "cmh_key.h" + +/* + * Maximum data that can be accumulated across .update() calls. + * The CMH eSW rejects HC_CMD_SAVE when ctx->outlen !=3D 0, which is + * always the case for KMAC (eip59_hc_kmac() sets ctx->outlen for + * right_encode(outlen) at finalization). All data must be buffered + * in kernel memory and submitted atomically in .final(). + * + * The CMH eSW does not serialize outlen into the external save + * context, so HC_CMD_SAVE fails for KMAC mode. + */ +#define KMAC_MAX_DATA (64 * 1024) + +/* Algorithm Table */ + +struct cmh_kmac_alg_info { + u32 hc_algo; + u32 digest_size; + const char *alg_name; + const char *drv_name; +}; + +static const struct cmh_kmac_alg_info cmh_kmac_algs_info[] =3D { + { + .hc_algo =3D HC_ALGO_SHAKE128, + .digest_size =3D CMH_SHAKE128_DIGEST_SIZE, + .alg_name =3D "kmac128", + .drv_name =3D "rambus-cmh-kmac128", + }, + { + .hc_algo =3D HC_ALGO_SHAKE256, + .digest_size =3D CMH_SHAKE256_DIGEST_SIZE, + .alg_name =3D "kmac256", + .drv_name =3D "rambus-cmh-kmac256", + }, +}; + +#define CMH_KMAC_ALG_COUNT ARRAY_SIZE(cmh_kmac_algs_info) + +/* Per-Request State */ + +struct cmh_kmac_chunk { + struct list_head list; + struct list_head tfm_node; /* per-tfm orphan tracking */ + u32 len; + u8 data[]; +}; + +/* + * Max payload slots for KMAC: + * SYS_CMD_WRITE (1) + KMAC (1) + inline S (3 max) + GATHER (1) + + * FINAL (1) + FLUSH (1) =3D 8 + */ +#define CMH_KMAC_MAX_PAYLOAD 9 +#define CMH_KMAC_MAX_PACKED (CMH_KMAC_MAX_PAYLOAD * 2) + +struct cmh_kmac_reqctx { + const struct cmh_kmac_alg_info *info; + int error; + struct list_head chunks; + u32 num_chunks; + u32 total_len; + /* DMA state for async final */ + dma_addr_t digest_dma; + dma_addr_t key_dma; + u8 *digest_buf; + struct cmh_sg_map *sgm; + u32 keylen; + struct vcq_cmd packed[CMH_KMAC_MAX_PACKED]; +}; + +/* Per-Transform State (carries key + S across requests) */ + +struct cmh_kmac_tfm_ctx { + struct cmh_key_ctx key; + u8 *custom; /* S (customization string), NULL if empty */ + u32 custom_len; + /* + * protects all_chunks. Only ever taken from process/softirq context + * (ahash .update/.final and the softirq/threaded-irq completion), + * never from hardirq, so spin_lock_bh() is the correct lock class. + */ + spinlock_t chunk_lock; + struct list_head all_chunks; /* orphan-safe chunk tracking */ + size_t tfm_buffered; /* bytes on all_chunks; DoS cap */ +}; + +/* + * Per-transform cap on total bytes buffered across all_chunks. Bounds + * memory an AF_ALG client can pin via repeated open/update/abandon of + * request sockets (the crypto API has no per-request destructor). + */ +#define CMH_KMAC_TFM_MAX_BUFFERED (16 * 1024 * 1024) + +/* VCQ Builders (KMAC-specific; shared builders in cmh_hc_abi.h / cmh_vcq.= h) */ + +static void vcq_add_hc_kmac(struct vcq_cmd *slot, u32 core_id, u64 key_ref= , u32 keylen, + u32 customlen, u32 algo, u32 outlen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, HC_CMD_KMAC); + slot->hwc.hc.cmd_kmac.key =3D key_ref; + slot->hwc.hc.cmd_kmac.custom =3D 0; /* inline */ + slot->hwc.hc.cmd_kmac.keylen =3D keylen; + slot->hwc.hc.cmd_kmac.customlen =3D customlen; + slot->hwc.hc.cmd_kmac.algo =3D algo; + slot->hwc.hc.cmd_kmac.outlen =3D outlen; +} + +/* Request Context Cleanup */ + +static void cmh_kmac_free_chunks(struct cmh_kmac_reqctx *rctx, + struct cmh_kmac_tfm_ctx *tctx) +{ + struct cmh_kmac_chunk *chunk, *tmp; + + spin_lock_bh(&tctx->chunk_lock); + list_for_each_entry_safe(chunk, tmp, &rctx->chunks, list) { + list_del(&chunk->list); + list_del(&chunk->tfm_node); + tctx->tfm_buffered -=3D chunk->len; + kfree(chunk); + } + spin_unlock_bh(&tctx->chunk_lock); + rctx->num_chunks =3D 0; + rctx->total_len =3D 0; +} + +static struct cmh_sg_map * +cmh_kmac_build_sg(struct cmh_kmac_reqctx *rctx, gfp_t gfp) +{ + struct cmh_dma_buf *bufs; + struct cmh_kmac_chunk *chunk; + struct cmh_sg_map *sgm; + u32 i; + + bufs =3D kcalloc(rctx->num_chunks, sizeof(*bufs), gfp); + if (!bufs) + return NULL; + + i =3D 0; + list_for_each_entry(chunk, &rctx->chunks, list) { + bufs[i].data =3D chunk->data; + bufs[i].len =3D chunk->len; + i++; + } + + sgm =3D cmh_dma_build_sg(bufs, rctx->num_chunks, gfp); + kfree(bufs); + return sgm; +} + +/* VCQ Packing + Submit */ + +/* ahash Operations */ + +struct cmh_kmac_alg_drv { + struct ahash_alg alg; + const struct cmh_kmac_alg_info *info; +}; + +static const struct cmh_kmac_alg_info * +cmh_kmac_get_info(struct crypto_ahash *tfm) +{ + struct ahash_alg *alg =3D crypto_ahash_alg(tfm); + + return container_of(alg, struct cmh_kmac_alg_drv, alg)->info; +} + +/* + * setkey blob for KMAC (raw key path): + * struct kmac_key_param { __be32 keylen; __be32 s_len; }; + * blob: kmac_key_param || key[keylen] || S[s_len] + */ +struct kmac_key_param { + __be32 keylen; + __be32 s_len; +}; + +static int cmh_kmac_setkey(struct crypto_ahash *tfm, const u8 *key, + unsigned int keylen) +{ + struct cmh_kmac_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + /* raw key bytes with optional S */ + { + struct kmac_key_param hdr; + u32 raw_keylen, s_len; + const u8 *ptr; + + if (keylen < sizeof(hdr)) + return -EINVAL; + + memcpy(&hdr, key, sizeof(hdr)); + raw_keylen =3D be32_to_cpu(hdr.keylen); + s_len =3D be32_to_cpu(hdr.s_len); + + if (keylen !=3D sizeof(hdr) + raw_keylen + s_len) + return -EINVAL; + + if (raw_keylen =3D=3D 0) + return -EINVAL; + + if (s_len > HC_CSHAKE_MAX_CUSTOMLEN) + return -EINVAL; + + ptr =3D key + sizeof(hdr); + + /* Store raw key */ + { + int ret =3D cmh_key_setkey_raw(&tctx->key, ptr, + raw_keylen, CORE_ID_HC); + if (ret) + return ret; + } + ptr +=3D raw_keylen; + + /* + * Store S. The raw key (above) and S are replaced without a + * lock against a concurrent .final() reader; safe because the + * crypto API does not issue setkey during an in-flight request + * on the same tfm, and AF_ALG (the only racing frontend) is + * not built on this kernel. + */ + kfree(tctx->custom); + tctx->custom =3D NULL; + tctx->custom_len =3D 0; + + if (s_len > 0) { + tctx->custom =3D kmemdup(ptr, s_len, GFP_KERNEL); + if (!tctx->custom) { + cmh_key_destroy(&tctx->key); + return -ENOMEM; + } + tctx->custom_len =3D s_len; + } + + return 0; + } +} + +static int cmh_kmac_init(struct ahash_request *req) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_kmac_reqctx *rctx =3D ahash_request_ctx(req); + + rctx->info =3D cmh_kmac_get_info(tfm); + rctx->error =3D 0; + INIT_LIST_HEAD(&rctx->chunks); + rctx->num_chunks =3D 0; + rctx->total_len =3D 0; + + return 0; +} + +static int cmh_kmac_update(struct ahash_request *req) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_kmac_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cmh_kmac_reqctx *rctx =3D ahash_request_ctx(req); + struct cmh_kmac_chunk *chunk; + int nents; + + if (rctx->error) + return rctx->error; + + if (!req->nbytes) + return 0; + + /* + * Cap cumulative input at KMAC_MAX_DATA (64 KB) before the alloc + * below, so sizeof(*chunk) + req->nbytes cannot overflow size_t. + */ + if (req->nbytes > KMAC_MAX_DATA - rctx->total_len) { + rctx->error =3D -EINVAL; + goto err_free_chunks; + } + + chunk =3D kmalloc(sizeof(*chunk) + req->nbytes, + req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC); + if (!chunk) { + rctx->error =3D -ENOMEM; + goto err_free_chunks; + } + + chunk->len =3D req->nbytes; + if (req->base.flags & CRYPTO_AHASH_REQ_VIRT) { + memcpy(chunk->data, req->svirt, req->nbytes); + } else { + nents =3D sg_nents_for_len(req->src, req->nbytes); + if (nents < 0 || + sg_copy_to_buffer(req->src, nents, + chunk->data, req->nbytes) !=3D req->nbytes) { + kfree(chunk); + rctx->error =3D -EINVAL; + goto err_free_chunks; + } + } + + spin_lock_bh(&tctx->chunk_lock); + if (tctx->tfm_buffered + chunk->len > CMH_KMAC_TFM_MAX_BUFFERED) { + spin_unlock_bh(&tctx->chunk_lock); + kfree(chunk); + rctx->error =3D -ENOMEM; + goto err_free_chunks; + } + list_add_tail(&chunk->list, &rctx->chunks); + list_add_tail(&chunk->tfm_node, &tctx->all_chunks); + tctx->tfm_buffered +=3D chunk->len; + spin_unlock_bh(&tctx->chunk_lock); + rctx->num_chunks++; + rctx->total_len +=3D req->nbytes; + + return 0; + +err_free_chunks: + /* + * Terminal error -- free all previously accumulated chunks. + * The crypto API hash path does not call .final() on error, + * so chunks would be orphaned otherwise. + */ + cmh_kmac_free_chunks(rctx, tctx); + return rctx->error; +} + +static void cmh_kmac_complete(void *data, int error) +{ + struct ahash_request *req =3D data; + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_kmac_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cmh_kmac_reqctx *rctx =3D ahash_request_ctx(req); + + if (error =3D=3D -EINPROGRESS) { + cmh_complete(&req->base, error); + return; + } + + cmh_dma_unmap_single(rctx->digest_dma, rctx->info->digest_size, + DMA_FROM_DEVICE); + + if (!error) + memcpy(req->result, rctx->digest_buf, + rctx->info->digest_size); + + kfree(rctx->digest_buf); + rctx->digest_buf =3D NULL; + cmh_dma_free_sg(rctx->sgm); + rctx->sgm =3D NULL; + cmh_kmac_free_chunks(rctx, tctx); + cmh_complete(&req->base, error); +} + +static int cmh_kmac_final(struct ahash_request *req) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_kmac_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cmh_kmac_reqctx *rctx =3D ahash_request_ctx(req); + const struct cmh_kmac_alg_info *info =3D rctx->info; + struct vcq_cmd cmds[CMH_KMAC_MAX_PAYLOAD]; + struct cmh_sg_map *sgm =3D NULL; + dma_addr_t digest_dma =3D DMA_MAPPING_ERROR, key_dma =3D DMA_MAPPING_ERRO= R; + u8 *digest_buf; + u64 key_ref; + u32 key_len; + struct core_dispatch d; + s32 target_mbx; + u32 core_id; + u32 idx; + int ret; + gfp_t gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + + if (rctx->error) { + ret =3D rctx->error; + goto out_free; + } + + if (tctx->key.mode =3D=3D CMH_KEY_NONE) { + ret =3D -ENOKEY; + goto out_free; + } + + if (rctx->num_chunks > 0) { + sgm =3D cmh_kmac_build_sg(rctx, gfp); + if (!sgm) { + ret =3D -ENOMEM; + goto out_free; + } + } + + digest_buf =3D kzalloc(info->digest_size, gfp); + if (!digest_buf) { + ret =3D -ENOMEM; + goto out_free_sg; + } + digest_dma =3D cmh_dma_map_single(digest_buf, info->digest_size, + DMA_FROM_DEVICE); + if (cmh_dma_map_error(digest_dma)) { + ret =3D -ENOMEM; + goto out_free_digest; + } + + /* Resolve key reference */ + idx =3D 0; + + key_dma =3D tctx->key.raw.dma; + vcq_add_sys_write(&cmds[idx++], SYS_REF_TEMP, (u64)key_dma, + SYS_REF_NONE, tctx->key.raw.len, + tctx->key.raw.sys_type); + key_ref =3D SYS_REF_TEMP; + key_len =3D tctx->key.raw.len; + d =3D cmh_core_select_instance(CMH_CORE_HC); + + target_mbx =3D d.mbx_idx; + + core_id =3D d.core_id; + + { + u32 span; + + vcq_add_hc_kmac(&cmds[idx], core_id, key_ref, key_len, + tctx->custom_len, info->hc_algo, + info->digest_size); + + /* Add inline S data after the KMAC slot */ + span =3D vcq_add_inline_data(&cmds[idx], tctx->custom, + tctx->custom_len); + idx +=3D span; + } + + if (sgm) + vcq_add_hc_gather(&cmds[idx++], core_id, (u64)sgm->items_dma, + HC_CMD_UPDATE); + + vcq_add_hc_final(&cmds[idx++], core_id, (u64)digest_dma, info->digest_siz= e); + vcq_add_flush(&cmds[idx++], core_id); + + rctx->digest_buf =3D digest_buf; + rctx->digest_dma =3D digest_dma; + rctx->sgm =3D sgm; + + ret =3D cmh_vcq_pack_and_submit_async(cmds, idx, rctx->packed, + CMH_KMAC_MAX_PACKED, + target_mbx, + cmh_kmac_complete, req, + !!(req->base.flags & + CRYPTO_TFM_REQ_MAY_BACKLOG), + cmh_tm_async_timeout_jiffies()); + if (ret =3D=3D -EBUSY) + return -EBUSY; + if (ret) { + /* + * Synchronous rejection (e.g. -EAGAIN: CMQ full, no backlog). + * Free only the per-submit transients and keep the accumulated + * chunks intact so the caller can retry the identical final(). + * If no retry comes, cra_exit reclaims the orphaned chunks; the + * per-tfm buffered-byte cap bounds how much stays pinned. + */ + cmh_dma_unmap_single(digest_dma, info->digest_size, + DMA_FROM_DEVICE); + kfree(digest_buf); + cmh_dma_free_sg(sgm); + rctx->digest_buf =3D NULL; + rctx->digest_dma =3D DMA_MAPPING_ERROR; + rctx->sgm =3D NULL; + return ret; + } + + return -EINPROGRESS; + +out_free_digest: + kfree(digest_buf); + +out_free_sg: + cmh_dma_free_sg(sgm); + +out_free: + cmh_kmac_free_chunks(rctx, tctx); + return ret; +} + +static int cmh_kmac_finup(struct ahash_request *req) +{ + int ret; + + ret =3D cmh_kmac_update(req); + if (ret) + return ret; + + return cmh_kmac_final(req); +} + +static int cmh_kmac_digest(struct ahash_request *req) +{ + int ret; + + ret =3D cmh_kmac_init(req); + if (ret) + return ret; + + return cmh_kmac_finup(req); +} + +/* + * export/import are intentionally unsupported for KMAC. + * + * Unlike the HMAC / AES-CMAC / SM4-CMAC / Poly1305 accumulators, KMAC has + * no software fallback to hand a large or mid-stream state off to, and its + * keyed-XOF input is buffered as a chunk list up to KMAC_MAX_DATA (64 KB). + * There is no bounded flat state we could serialise into a sane statesize + * (the siblings cap a small RAW window and switch to their fallback beyond + * it -- KMAC has neither). Returning -EOPNOTSUPP is non-breaking: the + * normal init/update/final and finup paths do not use export/import; only + * algif_hash transform cloning (accept() on an already-updated socket) + * does, and it degrades gracefully. statesize is kept non-zero so the + * crypto API still allocates a valid (unused) state buffer. + */ +static int cmh_kmac_export(struct ahash_request *req, void *out) +{ + return -EOPNOTSUPP; +} + +static int cmh_kmac_import(struct ahash_request *req, const void *in) +{ + return -EOPNOTSUPP; +} + +/* Transform init/exit */ + +static int cmh_kmac_cra_init(struct crypto_tfm *tfm) +{ + struct cmh_kmac_tfm_ctx *tctx =3D crypto_tfm_ctx(tfm); + + tctx->key.mode =3D CMH_KEY_NONE; + tctx->custom =3D NULL; + tctx->custom_len =3D 0; + spin_lock_init(&tctx->chunk_lock); + INIT_LIST_HEAD(&tctx->all_chunks); + crypto_ahash_set_reqsize(__crypto_ahash_cast(tfm), + sizeof(struct cmh_kmac_reqctx)); + return 0; +} + +static void cmh_kmac_cra_exit(struct crypto_tfm *tfm) +{ + struct cmh_kmac_tfm_ctx *tctx =3D crypto_tfm_ctx(tfm); + struct cmh_kmac_chunk *chunk, *tmp; + + /* Free any orphaned chunks (e.g. testmgr export/reimport poison) */ + spin_lock_bh(&tctx->chunk_lock); + list_for_each_entry_safe(chunk, tmp, &tctx->all_chunks, tfm_node) { + list_del(&chunk->tfm_node); + tctx->tfm_buffered -=3D chunk->len; + kfree(chunk); + } + spin_unlock_bh(&tctx->chunk_lock); + + cmh_key_destroy(&tctx->key); + kfree(tctx->custom); + tctx->custom =3D NULL; +} + +/* Registration */ + +static struct cmh_kmac_alg_drv cmh_kmac_drvs[CMH_KMAC_ALG_COUNT]; + +/** + * cmh_kmac_register() - Register KMAC-128/256 hash algorithms with the cr= ypto framework + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_kmac_register(void) +{ + unsigned int i; + int ret; + + if (!cmh_core_present(CMH_CORE_HC)) + return 0; + + for (i =3D 0; i < CMH_KMAC_ALG_COUNT; i++) { + const struct cmh_kmac_alg_info *info =3D + &cmh_kmac_algs_info[i]; + struct cmh_kmac_alg_drv *drv =3D &cmh_kmac_drvs[i]; + struct ahash_alg *alg =3D &drv->alg; + + drv->info =3D info; + + alg->init =3D cmh_kmac_init; + alg->update =3D cmh_kmac_update; + alg->final =3D cmh_kmac_final; + alg->finup =3D cmh_kmac_finup; + alg->digest =3D cmh_kmac_digest; + alg->export =3D cmh_kmac_export; + alg->import =3D cmh_kmac_import; + alg->setkey =3D cmh_kmac_setkey; + + alg->halg.digestsize =3D info->digest_size; + alg->halg.statesize =3D sizeof(struct cmh_kmac_reqctx); + + strscpy(alg->halg.base.cra_name, info->alg_name, + CRYPTO_MAX_ALG_NAME); + strscpy(alg->halg.base.cra_driver_name, info->drv_name, + CRYPTO_MAX_ALG_NAME); + alg->halg.base.cra_priority =3D 300; + alg->halg.base.cra_flags =3D CRYPTO_ALG_KERN_DRIVER_ONLY | + CRYPTO_ALG_NO_FALLBACK | + CRYPTO_ALG_ASYNC | + CRYPTO_ALG_REQ_VIRT; + alg->halg.base.cra_blocksize =3D 1; /* XOF/keyed XOF */ + alg->halg.base.cra_ctxsize =3D sizeof(struct cmh_kmac_tfm_ctx); + alg->halg.base.cra_init =3D cmh_kmac_cra_init; + alg->halg.base.cra_exit =3D cmh_kmac_cra_exit; + alg->halg.base.cra_module =3D THIS_MODULE; + + ret =3D crypto_register_ahash(alg); + if (ret) { + dev_err(cmh_dev(), "kmac: failed to register %s (rc=3D%d)\n", + info->drv_name, ret); + while (i--) + crypto_unregister_ahash(&cmh_kmac_drvs[i].alg); + return ret; + } + + dev_dbg(cmh_dev(), "kmac: registered %s (priority 300)\n", + info->drv_name); + } + + return 0; +} + +/** + * cmh_kmac_unregister() - Unregister KMAC hash algorithms from the crypto= framework + */ +void cmh_kmac_unregister(void) +{ + unsigned int i; + + if (!cmh_core_present(CMH_CORE_HC)) + return; + + for (i =3D 0; i < CMH_KMAC_ALG_COUNT; i++) { + crypto_unregister_ahash(&cmh_kmac_drvs[i].alg); + dev_dbg(cmh_dev(), "kmac: unregistered %s\n", + cmh_kmac_algs_info[i].drv_name); + } +} diff --git a/drivers/crypto/cmh/cmh_main.c b/drivers/crypto/cmh/cmh_main.c index 589363cc4b28..474b438fc18a 100644 --- a/drivers/crypto/cmh/cmh_main.c +++ b/drivers/crypto/cmh/cmh_main.c @@ -33,6 +33,8 @@ #include "cmh_rh.h" #include "cmh_hash.h" #include "cmh_hmac.h" +#include "cmh_cshake.h" +#include "cmh_kmac.h" #include "cmh_mgmt.h" #include "cmh_registers.h" #include "cmh_debugfs.h" @@ -214,6 +216,16 @@ static int cmh_probe(struct platform_device *pdev) if (ret) goto err_hmac_register; =20 + /* Register CSHAKE hash algorithms */ + ret =3D cmh_cshake_register(); + if (ret) + goto err_cshake_register; + + /* Register KMAC hash algorithms */ + ret =3D cmh_kmac_register(); + if (ret) + goto err_kmac_register; + /* Register key management device (/dev/cmh_mgmt) */ ret =3D cmh_mgmt_register(); if (ret) @@ -224,6 +236,10 @@ static int cmh_probe(struct platform_device *pdev) return 0; =20 err_mgmt_register: + cmh_kmac_unregister(); +err_kmac_register: + cmh_cshake_unregister(); +err_cshake_register: cmh_hmac_unregister(); err_hmac_register: cmh_hash_unregister(); @@ -254,6 +270,8 @@ static void cmh_remove(struct platform_device *pdev) cfg =3D &dev->config; =20 cmh_mgmt_unregister(); + cmh_kmac_unregister(); + cmh_cshake_unregister(); cmh_hmac_unregister(); cmh_hash_unregister(); cmh_rh_cleanup(cfg); diff --git a/drivers/crypto/cmh/include/cmh_cshake.h b/drivers/crypto/cmh/i= nclude/cmh_cshake.h new file mode 100644 index 000000000000..9bafe0baf52f --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_cshake.h @@ -0,0 +1,16 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Kernel Crypto API CSHAKE Driver + * + * Registers cSHAKE-128 and cSHAKE-256 ahash algorithms using + * HC_CMD_CSHAKE with inline customization string S. + */ + +#ifndef CMH_CSHAKE_H +#define CMH_CSHAKE_H + +int cmh_cshake_register(void); +void cmh_cshake_unregister(void); + +#endif /* CMH_CSHAKE_H */ diff --git a/drivers/crypto/cmh/include/cmh_kmac.h b/drivers/crypto/cmh/inc= lude/cmh_kmac.h new file mode 100644 index 000000000000..b3c92d71a0b6 --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_kmac.h @@ -0,0 +1,16 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Kernel Crypto API KMAC Driver + * + * Registers KMAC-128 and KMAC-256 ahash algorithms using + * HC_CMD_KMAC with inline customization string S. + */ + +#ifndef CMH_KMAC_H +#define CMH_KMAC_H + +int cmh_kmac_register(void); +void cmh_kmac_unregister(void); + +#endif /* CMH_KMAC_H */ --=20 2.43.7 From nobody Fri Sep 25 01:20:34 2026 Received: from CO1PR03CU002.outbound.protection.outlook.com (mail-westus2azon11020091.outbound.protection.outlook.com [52.101.46.91]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0E26B4B0481; Thu, 17 Sep 2026 22:59:40 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.46.91 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789685984; cv=fail; b=qGcYz9JLloqvBSHNQDaOwB+o43v9+phbRMh0bTUlw+xkIMGYkPI4sWdMSOKAJM902/Dn01w6v75+pTKXeV4kfdbmHdx9io1Znss3T6SzAgPu4tDVz5rCVrKOwuIilvbzk2M+72G/q5zb0WqnDHGMjGqLSMCw9BCcmW9lfMYPp7c= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789685984; c=relaxed/simple; bh=evSt4A/FvkoEBAv+dzUe+LOCUW3gjCyHmNEsaYrdtsk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=kuT87tLSFpC2O00qCVO1Gbd7BGJPidHTjo/r606enrb3YX0XnkDtWNHReMhpP0gXffmH5x3D90ZP6IMzL2+x7iFpCTw4Y+JG2kKCSqen1W8sVAv3kgRVrR6E5f2Inr2hso9ZxEiwp4dK73O6AhXMFYQPgdMroZjnyVPrScDBJlQ= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com; spf=fail smtp.mailfrom=rambus.com; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b=gObE+Ts+; arc=fail smtp.client-ip=52.101.46.91 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=rambus.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b="gObE+Ts+" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=d9SuSS6cfPxZBn+evcz6kodLSBwWJ8+queg7Zy6sQ/NaYYRGjqrJANpNMzmG/mc7N3XyGLT93+iBCQrhcIH1vfD3uUf/MJrT51LQX2RstRPTkkSq4asifFWaJVLhjg9aP06UnZgSq6uhiJ7KsHpXk61x6C1179+tLm4WJTIK9YlbwthmVO1W9COFE3df/FUwwugznTYXF0rutYW4d+EjpZ+F3Rtl2+LWUiSmO5y6WfkWwySElJuygMTlQpyprkJpfAeXLmzn2hdEGcYSOD/WdeHlOjs0ra9BP/U4pGy1pKYkxB1WxCJNuoJ1NS5wN4kxgXFoiUNl7H+GZ5TkBlFwRg== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=bDt7wAK5h+ciJIa07ayF73SKX4Bw7G7UKRr2NijQ5OM=; b=FLXD+QE5EDzeJ0dSBCv6M31u4kCaFrH3kFyHNy/5N6QFxz87z6ecFfmjSdiHrrpZ/IJuL0GJV1O9oYmCFHQ7ZaeNd7D9QqV6/TzKx1OPRCVVKgqWIF0gbOZ+GLTkcLQa9+pzpJwxiKcxxTdElCzDMTIsO1z2SDHEUZpevvdt877StHmHZcJLMredmmPfkH+nCLVyzx/B9ht7HzwC6IDuEdfQspI+3Pun0bCdOXLIXtBSZvu9W0TlK/JVJqAy/oGXUnTD/c0AvQ84b1I98PBGHSVmAkAefqMeJMQjE31pXHpQIn1/mTnnMyk8Gllf3REExLxxq4tibIxaiKCNu+y2Tw== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 192.86.86.210) smtp.rcpttodomain=cryptography.com smtp.mailfrom=rambus.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=rambus.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=rambus.com; s=selector1; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=bDt7wAK5h+ciJIa07ayF73SKX4Bw7G7UKRr2NijQ5OM=; b=gObE+Ts+W5bYqDoIyUP5jJ11fL/oJb2l3ll4gqEoTQfreXQaerlpGRS0bLUytKhyX4oRCuDzbAG/2+O1BTa6Ih+DDQ9hlWWdCIX4rt+PzI2Zo1XjUDrJOkoL0pkKARDkgZWX2T/PcW4sG8B2xwZ1/c03TNP9is/SVmzraj9JgLFNsL+pI6BiGWbRux+fAHG66SbzSmVorAxv2ay6O7dDxQ9HYsT5hIspV9/GjdJVAsiGBGT2qafxy59/yG0jazP6Z/W6QqHmgaxP0R4uVpuXhPXS3sWe1oYTrEbabrk/V2sMq1nZXy5BgYxD9PEzPRc1xbvhkVeKzxk6Iyi5sSjSaA== Received: from BLAPR03CA0022.namprd03.prod.outlook.com (2603:10b6:208:32b::27) by CY5PR04MB994499.namprd04.prod.outlook.com (2603:10b6:930:12c::12) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.428.13; Thu, 17 Sep 2026 22:59:34 +0000 Received: from BL6PEPF00020E66.namprd04.prod.outlook.com (2603:10b6:208:32b:cafe::55) by BLAPR03CA0022.outlook.office365.com (2603:10b6:208:32b::27) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.428.12 via Frontend Transport; Thu, 17 Sep 2026 22:59:34 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 192.86.86.210) smtp.mailfrom=rambus.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=rambus.com; Received-SPF: Pass (protection.outlook.com: domain of rambus.com designates 192.86.86.210 as permitted sender) receiver=protection.outlook.com; client-ip=192.86.86.210; helo=hqxsv-psmtppxy02.rambus.com; pr=C Received: from hqxsv-psmtppxy02.rambus.com (192.86.86.210) by BL6PEPF00020E66.mail.protection.outlook.com (10.167.249.27) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.451.8 via Frontend Transport; Thu, 17 Sep 2026 22:59:34 +0000 Received: from hqxsv-cmdev3-aousherovitch.rambus.com (hqn-lb-int-float.rambus.com [10.12.20.20]) by hqxsv-psmtppxy02.rambus.com (Postfix) with ESMTP id 16D401801763; Thu, 17 Sep 2026 22:59:32 +0000 (UTC) From: Alex Ousherovitch To: Albert Ou , Alex Ousherovitch , Conor Dooley , "David S. Miller" , Herbert Xu , Jonathan Corbet , Krzysztof Kozlowski , Palmer Dabbelt , Paul Walmsley , Rob Herring , Saravanakrishnan Krishnamoorthy , Shuah Khan Cc: Alexandre Ghiti , devicetree@vger.kernel.org, Joel Wittenauer , linux-api@vger.kernel.org, linux-crypto@vger.kernel.org, linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-riscv@lists.infradead.org, Shuah Khan , Thi Nguyen Subject: [PATCH v5 07/19] crypto: cmh - add SM3 ahash Date: Thu, 17 Sep 2026 15:59:16 -0700 Message-ID: <20260917225929.2494111-8-aousherovitch@rambus.com> X-Mailer: git-send-email 2.43.7 In-Reply-To: <20260917225929.2494111-1-aousherovitch@rambus.com> References: <20260917225929.2494111-1-aousherovitch@rambus.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: BL6PEPF00020E66:EE_|CY5PR04MB994499:EE_ X-MS-Office365-Filtering-Correlation-Id: ec77f20e-171f-476a-f68b-08df150f5762 X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|36860700016|82310400026|23010399003|1800799024|376014|7416014|3023799007|56012099006|10067099003|11063799006|921020|22082099003|18002099003; X-Microsoft-Antispam-Message-Info: Lz4ktcXv+GHHGVvNf4LwXJ1wktbeM0c2I0Sat8TlMMJU0UA50dx7E4ueQr0gcgfxkFh7e5KManHMklEcUAH1+AEkXQZNp/ckHekYBY+IeNWWmjWzYIwRovX3/+zKdsFUA/h+aox4vyHCU5+04yzkBPo+9brA5KUT8yrT+4v909aC1wAX9a5dN8crrCGHSwGeY5UlDQuW4k4a3nY3XkziesLTHIgH+j+BumIO5Y39c3bo1ZyAi7CZZMEVpSVaT3x2kmvpq+1O1H/ye6CvACHEWzAgDwKYybbqvMCOSqHe5zaiN2JOxvlQ2k9ud3bvNmqhxRcTuYl+g8rM6fmeo7JC66yrf6EWxJAmppuZzJw6+uHRCE61c3EFOYtzz2vCJxdIdvtInvpyyUIexlU7eWnnhAxvl1y74TJAqiUZgKN1N54rnZEEvzO/6u0aiwoiNN95rbjQhkn1gbXpaE10X0c1FEmyC0TIF2s+4KjySHpp2KzxIJ1+nzwoyS6qdQ9XvmaxI/eY4oe0U8St+o1j7gPEV0yAM0ewsmktNPrU29zmlW8NyJ6huH9GKPeuwTULg1ItTaSEFFo1R5R2mNmC/IuPD4vSee/G0KhCWFJ1e6kmT4OFVGT6q0OgwvJTmQDdIZukEbYWsKtdfl2LVWyUSrwiuu6CRenoP3IhHqowHvtcGMFWA/WpUxFPUPjXAiHSRLCFYQJiebm+IDvSYYTttftzSk016vfObTftZMkljhYulDvyTUuId7PHkXCtFKKhMjTx X-Forefront-Antispam-Report: CIP:192.86.86.210;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:hqxsv-psmtppxy02.rambus.com;PTR:InfoDomainNonexistent;CAT:NONE;SFS:(13230040)(36860700016)(82310400026)(23010399003)(1800799024)(376014)(7416014)(3023799007)(56012099006)(10067099003)(11063799006)(921020)(22082099003)(18002099003);DIR:OUT;SFP:1102; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: s9IrzjqNYuwPi3uMCSQsz9dJCswuH5Bq0JmiZoI68UGbmP/er/M/IOxABVPe3vKGVfVy0fqUl7ytMKP3otUmbrjqhB9h8MkG2FRfuXZTxL2YIJVntFJ7CR5RjAHVyO1sHm7f8wHu56tahEeEvV0IyRqE4zu2dOy+PTmsE2fjRoIcIpDbsqfFi7zpyJHNzBf06i438+ovqgkTVC4g0VW7aKTeUJO+9vigkGj7Nbcryecg3NoZ+RqqrskCERNE0X2wWBybJ2FboyMnmBbAWUNxH4RRn9b6tY1RF3FqXg3F0rUF8DPwhn48WAaGjYqwufl7u3txqYj+xQE5AEr/6EGmkuLJ9FDye6I5BVPSoxQFBkAMJrnV99G97vmr3xx5fEsUD+mQWjpoGqzoqSpHeX73XJ50Ot5HS4NY8G2EtV+eERHtSNIaaQOMGkaFV/jJRGvU X-OriginatorOrg: rambus.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 17 Sep 2026 22:59:34.2354 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: ec77f20e-171f-476a-f68b-08df150f5762 X-MS-Exchange-CrossTenant-Id: bd0ba799-c2b9-413c-9c56-5d1731c4827c X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=bd0ba799-c2b9-413c-9c56-5d1731c4827c;Ip=[192.86.86.210];Helo=[hqxsv-psmtppxy02.rambus.com] X-MS-Exchange-CrossTenant-AuthSource: BL6PEPF00020E66.namprd04.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: CY5PR04MB994499 Content-Type: text/plain; charset="utf-8" Register the SM3 ahash algorithm using the CMH SM3 core (core ID 0x05). Supports incremental update/finup/final and export/import. Signed-off-by: Alex Ousherovitch Co-developed-by: Saravanakrishnan Krishnamoorthy Signed-off-by: Saravanakrishnan Krishnamoorthy --- drivers/crypto/cmh/Makefile | 3 +- drivers/crypto/cmh/cmh_main.c | 9 + drivers/crypto/cmh/cmh_sm3.c | 604 +++++++++++++++++++++++++++ drivers/crypto/cmh/include/cmh_sm3.h | 28 ++ 4 files changed, 643 insertions(+), 1 deletion(-) create mode 100644 drivers/crypto/cmh/cmh_sm3.c create mode 100644 drivers/crypto/cmh/include/cmh_sm3.h diff --git a/drivers/crypto/cmh/Makefile b/drivers/crypto/cmh/Makefile index 3ab154744658..664b2a20bc19 100644 --- a/drivers/crypto/cmh/Makefile +++ b/drivers/crypto/cmh/Makefile @@ -18,7 +18,8 @@ cmh-y :=3D \ cmh_hash.o \ cmh_hmac.o \ cmh_cshake.o \ - cmh_kmac.o + cmh_kmac.o \ + cmh_sm3.o =20 # Management ioctl device (/dev/cmh_mgmt): key lifecycle, PKE, PQC ioctls. cmh-$(CONFIG_CRYPTO_DEV_CMH_MGMT) +=3D \ diff --git a/drivers/crypto/cmh/cmh_main.c b/drivers/crypto/cmh/cmh_main.c index 474b438fc18a..4ad1500d8e49 100644 --- a/drivers/crypto/cmh/cmh_main.c +++ b/drivers/crypto/cmh/cmh_main.c @@ -35,6 +35,7 @@ #include "cmh_hmac.h" #include "cmh_cshake.h" #include "cmh_kmac.h" +#include "cmh_sm3.h" #include "cmh_mgmt.h" #include "cmh_registers.h" #include "cmh_debugfs.h" @@ -226,6 +227,11 @@ static int cmh_probe(struct platform_device *pdev) if (ret) goto err_kmac_register; =20 + /* Register SM3 hash algorithm */ + ret =3D cmh_sm3_register(); + if (ret) + goto err_sm3_register; + /* Register key management device (/dev/cmh_mgmt) */ ret =3D cmh_mgmt_register(); if (ret) @@ -236,6 +242,8 @@ static int cmh_probe(struct platform_device *pdev) return 0; =20 err_mgmt_register: + cmh_sm3_unregister(); +err_sm3_register: cmh_kmac_unregister(); err_kmac_register: cmh_cshake_unregister(); @@ -270,6 +278,7 @@ static void cmh_remove(struct platform_device *pdev) cfg =3D &dev->config; =20 cmh_mgmt_unregister(); + cmh_sm3_unregister(); cmh_kmac_unregister(); cmh_cshake_unregister(); cmh_hmac_unregister(); diff --git a/drivers/crypto/cmh/cmh_sm3.c b/drivers/crypto/cmh/cmh_sm3.c new file mode 100644 index 000000000000..a08341512e5d --- /dev/null +++ b/drivers/crypto/cmh/cmh_sm3.c @@ -0,0 +1,604 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- SM3 Hash Driver (CORE_ID_SM3) + * + * Registers an asynchronous hash (ahash) algorithm for SM3 + * (GB/T 32905-2016) using the CMH SM3 core. This is a standalone + * driver separate from cmh_hash.c (which handles HC-based SHA-2/3/SHAKE) + * because SM3 runs on a different hardware core with its own command + * IDs and context layout. + * + * Incremental HW update model (same pattern as cmh_hash.c). The driver + * sets CRYPTO_AHASH_ALG_BLOCK_ONLY, so the Crypto API buffers partial + * blocks and .update() only ever sees whole-block-aligned data: + * + * .init() -> software-only: zero per-request context + * .update() -> SM3_CMD_INIT [+ RESTORE] + UPDATE + SAVE + FLUSH; + * return -EINPROGRESS, completing with the leftover byte + * count the API must buffer (0 when block-aligned) + * .finup() -> SM3_CMD_INIT [+ RESTORE] [+ UPDATE] + FINAL + FLUSH + * (also serves .final via the API: nbytes =3D=3D 0) + * .digest() -> INIT + UPDATE + FINAL + FLUSH (single-shot via finup) + * .export()/.import() -> software-only: copy the SM3 context + * checkpoint; the API appends its own partial-block buffer + * + * This is an sg-only driver (no CRYPTO_ALG_REQ_VIRT): BLOCK_ONLY buffer + * prepending assumes scatterlists. + */ + +#include +#include +#include +#include +#include +#include +#include +#include + +#include "cmh_sm3.h" +#include "cmh_vcq.h" +#include "cmh_txn.h" +#include "cmh_dma.h" + +/* Per-Request State */ + +/* + * Exported SM3 state -- serialised by .export(), deserialised by + * .import(). This is what statesize advertises; the Crypto API + * appends its own partial-block buffer on top. + */ +struct cmh_sm3_export_state { + u8 checkpoint[SM3_CONTEXT_SIZE]; /* SM3 context from last SAVE */ + u32 hw_started; /* non-zero if checkpoint valid */ +}; + +#define CMH_SM3_MAX_PAYLOAD 5 /* INIT + RESTORE + UPDATE + FINAL/SAVE= + FLUSH */ +#define CMH_SM3_MAX_PACKED (CMH_SM3_MAX_PAYLOAD * 2) + +/* + * Checkpoint embedded inline: the kernel ahash API has no per-request + * destructor, so a heap-allocated checkpoint leaks if a request is + * abandoned without .final(). + * + * Mapping the inline checkpoint for DMA is safe even on non-coherent + * platforms: it is only ever mapped DMA_TO_DEVICE and its bytes are + * frozen from dma_map_single() until the matching unmap (it is written + * solely in the completion, after the unmap). CPU writes to adjacent + * fields sharing a cacheline never alter the checkpoint bytes, and a + * TO_DEVICE unmap performs no cache invalidate; the shared-cacheline + * hazard applies only to FROM_DEVICE buffers, which are kmalloc'd. + */ +struct cmh_sm3_reqctx { + int error; + u32 hw_started; + u32 has_checkpoint; + u32 update_remainder; /* sub-block bytes the API must re-buffer */ + u8 checkpoint[SM3_CONTEXT_SIZE]; /* SM3 context from last SAVE */ + /* DMA state for current async operation */ + dma_addr_t ckpt_dma; + dma_addr_t save_dma; + dma_addr_t data_dma; + dma_addr_t digest_dma; + u8 *save_buf; + u8 *data_buf; + u32 data_len; + u8 *digest_buf; + struct vcq_cmd packed[CMH_SM3_MAX_PACKED]; +}; + +/* VCQ Builders -- SM3 core (CORE_ID_SM3); generic flush from cmh_vcq.h */ + +static void vcq_add_sm3_init(struct vcq_cmd *slot, u32 core_id) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, SM3_CMD_INIT); + /* SM3 has a single algorithm -- no algo selector field */ +} + +static void vcq_add_sm3_update(struct vcq_cmd *slot, u32 core_id, u64 inpu= t_phys, u32 len) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, SM3_CMD_UPDATE); + slot->hwc.sm3.cmd_update.input =3D input_phys; + slot->hwc.sm3.cmd_update.inlen =3D len; +} + +static void vcq_add_sm3_final(struct vcq_cmd *slot, u32 core_id, u64 diges= t_phys, u32 outlen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, SM3_CMD_FINAL); + slot->hwc.sm3.cmd_final.digest =3D digest_phys; + slot->hwc.sm3.cmd_final.outlen =3D outlen; +} + +static void vcq_add_sm3_save(struct vcq_cmd *slot, u32 core_id, u64 output= _phys, u32 outlen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, SM3_CMD_SAVE); + slot->hwc.sm3.cmd_save.output =3D output_phys; + slot->hwc.sm3.cmd_save.outlen =3D outlen; +} + +static void vcq_add_sm3_restore(struct vcq_cmd *slot, u32 core_id, u64 inp= ut_phys, u32 inlen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, SM3_CMD_RESTORE); + slot->hwc.sm3.cmd_restore.input =3D input_phys; + slot->hwc.sm3.cmd_restore.inlen =3D inlen; +} + +/* Request Context Cleanup */ + +static void cmh_sm3_free_reqctx(struct cmh_sm3_reqctx *rctx) +{ + rctx->has_checkpoint =3D 0; +} + +/* VCQ Packing + Submit */ + +/* ahash Operations */ + +static int cmh_sm3_init(struct ahash_request *req) +{ + struct cmh_sm3_reqctx *rctx =3D ahash_request_ctx(req); + + memset(rctx, 0, sizeof(*rctx)); + return 0; +} + +/* + * Update completion -- takes ownership of save_buf as new checkpoint. + */ +static void cmh_sm3_update_complete(void *data, int error) +{ + struct ahash_request *req =3D data; + struct cmh_sm3_reqctx *rctx =3D ahash_request_ctx(req); + + if (error =3D=3D -EINPROGRESS) { + cmh_complete(&req->base, error); + return; + } + + if (rctx->has_checkpoint) + cmh_dma_unmap_single(rctx->ckpt_dma, SM3_CONTEXT_SIZE, + DMA_TO_DEVICE); + cmh_dma_unmap_single(rctx->save_dma, SM3_CONTEXT_SIZE, + DMA_FROM_DEVICE); + cmh_dma_unmap_single(rctx->data_dma, rctx->data_len, + DMA_TO_DEVICE); + + if (!error) { + memcpy(rctx->checkpoint, rctx->save_buf, SM3_CONTEXT_SIZE); + rctx->has_checkpoint =3D 1; + kfree(rctx->save_buf); + rctx->save_buf =3D NULL; + rctx->hw_started =3D 1; + /* Hand the API the sub-block remainder it must re-buffer. */ + error =3D rctx->update_remainder; + } else { + kfree(rctx->save_buf); + rctx->save_buf =3D NULL; + rctx->error =3D error; + } + + kfree(rctx->data_buf); + rctx->data_buf =3D NULL; + rctx->data_len =3D 0; + + cmh_complete(&req->base, error); +} + +static int cmh_sm3_update(struct ahash_request *req) +{ + struct cmh_sm3_reqctx *rctx =3D ahash_request_ctx(req); + struct vcq_cmd cmds[CMH_SM3_MAX_PAYLOAD]; + struct core_dispatch d; + u32 full_len; + u32 idx; + int ret; + gfp_t gfp; + + if (rctx->error) + return rctx->error; + + if (!req->nbytes) + return 0; + + /* block size is a power of two, but modulo keeps the split exact. */ + rctx->update_remainder =3D req->nbytes % CMH_SM3_BLOCK_SIZE; + full_len =3D req->nbytes - rctx->update_remainder; + + gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + + /* + * Reject a single update whose linearisation would exceed the largest + * kmalloc: return a permanent -EMSGSIZE ("message too long") rather + * than a transient -ENOMEM the client would keep retrying. __GFP_NOWARN + * keeps a borderline-large (but sub-cap) request quiet if it still + * cannot be satisfied. + */ + if (full_len > KMALLOC_MAX_SIZE) + return -EMSGSIZE; + + rctx->data_buf =3D kmalloc(full_len, gfp | __GFP_NOWARN); + if (!rctx->data_buf) + return -ENOMEM; + + scatterwalk_map_and_copy(rctx->data_buf, req->src, 0, full_len, 0); + + rctx->data_len =3D full_len; + + rctx->save_buf =3D kzalloc(SM3_CONTEXT_SIZE, gfp); + if (!rctx->save_buf) { + ret =3D -ENOMEM; + goto err_free; + } + + rctx->data_dma =3D cmh_dma_map_single(rctx->data_buf, full_len, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->data_dma)) { + ret =3D -ENOMEM; + goto err_free; + } + + rctx->save_dma =3D cmh_dma_map_single(rctx->save_buf, SM3_CONTEXT_SIZE, + DMA_FROM_DEVICE); + if (cmh_dma_map_error(rctx->save_dma)) { + ret =3D -ENOMEM; + goto err_unmap_data; + } + + rctx->ckpt_dma =3D DMA_MAPPING_ERROR; + if (rctx->has_checkpoint) { + rctx->ckpt_dma =3D cmh_dma_map_single(rctx->checkpoint, + SM3_CONTEXT_SIZE, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->ckpt_dma)) { + ret =3D -ENOMEM; + goto err_unmap_save; + } + } + + d =3D cmh_core_select_instance(CMH_CORE_SM3); + idx =3D 0; + + vcq_add_sm3_init(&cmds[idx++], d.core_id); + + if (rctx->has_checkpoint) + vcq_add_sm3_restore(&cmds[idx++], d.core_id, + (u64)rctx->ckpt_dma, SM3_CONTEXT_SIZE); + + vcq_add_sm3_update(&cmds[idx++], d.core_id, + (u64)rctx->data_dma, full_len); + + vcq_add_sm3_save(&cmds[idx++], d.core_id, + (u64)rctx->save_dma, SM3_CONTEXT_SIZE); + + vcq_add_flush(&cmds[idx++], d.core_id); + + ret =3D cmh_vcq_pack_and_submit_async(cmds, idx, rctx->packed, + CMH_SM3_MAX_PACKED, + d.mbx_idx, + cmh_sm3_update_complete, req, + !!(req->base.flags & + CRYPTO_TFM_REQ_MAY_BACKLOG), + cmh_tm_async_timeout_jiffies()); + if (ret && ret !=3D -EBUSY) + goto err_unmap_ckpt; + + if (ret =3D=3D -EBUSY) + return -EBUSY; + return -EINPROGRESS; + +err_unmap_ckpt: + if (rctx->has_checkpoint) + cmh_dma_unmap_single(rctx->ckpt_dma, SM3_CONTEXT_SIZE, + DMA_TO_DEVICE); +err_unmap_save: + cmh_dma_unmap_single(rctx->save_dma, SM3_CONTEXT_SIZE, + DMA_FROM_DEVICE); +err_unmap_data: + cmh_dma_unmap_single(rctx->data_dma, full_len, DMA_TO_DEVICE); +err_free: + kfree(rctx->save_buf); + rctx->save_buf =3D NULL; + kfree(rctx->data_buf); + rctx->data_buf =3D NULL; + rctx->data_len =3D 0; + return ret; +} + +static void cmh_sm3_final_complete(void *data, int error) +{ + struct ahash_request *req =3D data; + struct cmh_sm3_reqctx *rctx =3D ahash_request_ctx(req); + + if (error =3D=3D -EINPROGRESS) { + cmh_complete(&req->base, error); + return; + } + + if (rctx->has_checkpoint) + cmh_dma_unmap_single(rctx->ckpt_dma, SM3_CONTEXT_SIZE, + DMA_TO_DEVICE); + if (rctx->data_buf) + cmh_dma_unmap_single(rctx->data_dma, rctx->data_len, + DMA_TO_DEVICE); + cmh_dma_unmap_single(rctx->digest_dma, CMH_SM3_DIGEST_SIZE, + DMA_FROM_DEVICE); + + if (!error) + memcpy(req->result, rctx->digest_buf, CMH_SM3_DIGEST_SIZE); + + kfree(rctx->digest_buf); + rctx->digest_buf =3D NULL; + kfree(rctx->data_buf); + rctx->data_buf =3D NULL; + cmh_sm3_free_reqctx(rctx); + cmh_complete(&req->base, error); +} + +static int cmh_sm3_submit_final(struct ahash_request *req, + u8 *data_buf, u32 data_len) +{ + struct cmh_sm3_reqctx *rctx =3D ahash_request_ctx(req); + struct vcq_cmd cmds[CMH_SM3_MAX_PAYLOAD]; + struct core_dispatch d; + u32 idx; + int ret; + gfp_t gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + + rctx->data_buf =3D data_buf; + rctx->data_len =3D data_len; + + rctx->digest_buf =3D kzalloc(CMH_SM3_DIGEST_SIZE, gfp); + if (!rctx->digest_buf) { + ret =3D -ENOMEM; + goto err_free_data; + } + + rctx->digest_dma =3D cmh_dma_map_single(rctx->digest_buf, + CMH_SM3_DIGEST_SIZE, + DMA_FROM_DEVICE); + if (cmh_dma_map_error(rctx->digest_dma)) { + ret =3D -ENOMEM; + goto err_free_digest; + } + + rctx->data_dma =3D DMA_MAPPING_ERROR; + if (data_buf && data_len > 0) { + rctx->data_dma =3D cmh_dma_map_single(data_buf, data_len, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->data_dma)) { + ret =3D -ENOMEM; + goto err_unmap_digest; + } + } + + rctx->ckpt_dma =3D DMA_MAPPING_ERROR; + if (rctx->has_checkpoint) { + rctx->ckpt_dma =3D cmh_dma_map_single(rctx->checkpoint, + SM3_CONTEXT_SIZE, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->ckpt_dma)) { + ret =3D -ENOMEM; + goto err_unmap_data; + } + } + + d =3D cmh_core_select_instance(CMH_CORE_SM3); + idx =3D 0; + + vcq_add_sm3_init(&cmds[idx++], d.core_id); + + if (rctx->has_checkpoint) + vcq_add_sm3_restore(&cmds[idx++], d.core_id, + (u64)rctx->ckpt_dma, SM3_CONTEXT_SIZE); + + if (data_buf && data_len > 0) + vcq_add_sm3_update(&cmds[idx++], d.core_id, + (u64)rctx->data_dma, data_len); + + vcq_add_sm3_final(&cmds[idx++], d.core_id, + (u64)rctx->digest_dma, CMH_SM3_DIGEST_SIZE); + + vcq_add_flush(&cmds[idx++], d.core_id); + + ret =3D cmh_vcq_pack_and_submit_async(cmds, idx, rctx->packed, + CMH_SM3_MAX_PACKED, + d.mbx_idx, + cmh_sm3_final_complete, req, + !!(req->base.flags & + CRYPTO_TFM_REQ_MAY_BACKLOG), + cmh_tm_async_timeout_jiffies()); + if (ret =3D=3D -EBUSY) + return -EBUSY; + if (ret) + goto err_unmap_ckpt; + + return -EINPROGRESS; + +err_unmap_ckpt: + if (rctx->has_checkpoint) + cmh_dma_unmap_single(rctx->ckpt_dma, SM3_CONTEXT_SIZE, + DMA_TO_DEVICE); +err_unmap_data: + if (data_buf && data_len > 0) + cmh_dma_unmap_single(rctx->data_dma, data_len, + DMA_TO_DEVICE); +err_unmap_digest: + cmh_dma_unmap_single(rctx->digest_dma, CMH_SM3_DIGEST_SIZE, + DMA_FROM_DEVICE); +err_free_digest: + kfree(rctx->digest_buf); + rctx->digest_buf =3D NULL; +err_free_data: + kfree(data_buf); + rctx->data_buf =3D NULL; + /* + * Preserve the SM3 checkpoint on failure: a synchronous rejection is + * retryable, and for a terminal error the inline checkpoint is freed + * with the request context, so it never leaks. It is cleared only in + * the completion after a successful final(). + */ + return ret; +} + +static int cmh_sm3_finup(struct ahash_request *req); + +/* + * One-shot digest -- delegates to init + finup so that all data is + * linearised and mapped through cmh_dma_map_single(), which is the + * only DMA mapping path aware of all supported DMA backends. + */ +static int cmh_sm3_digest(struct ahash_request *req) +{ + int ret; + + ret =3D cmh_sm3_init(req); + if (ret) + return ret; + return cmh_sm3_finup(req); +} + +/* + * .finup -- hash any remaining data and finalise in one transaction. + * With BLOCK_ONLY the Crypto API prepends the bytes it held back, so + * req->src already carries the full tail. Also serves .final (nbytes + * =3D=3D 0). Avoids ahash_def_finup(), which would clone via export/impo= rt. + */ +static int cmh_sm3_finup(struct ahash_request *req) +{ + struct cmh_sm3_reqctx *rctx =3D ahash_request_ctx(req); + u32 data_len =3D req->nbytes; + u8 *data_buf =3D NULL; + gfp_t gfp; + + if (rctx->error) + return rctx->error; + + if (data_len =3D=3D 0) + return cmh_sm3_submit_final(req, NULL, 0); + + /* Reject an oversized linearisation with a permanent -EMSGSIZE. */ + if (data_len > KMALLOC_MAX_SIZE) + return -EMSGSIZE; + + gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + + data_buf =3D kmalloc(data_len, gfp | __GFP_NOWARN); + if (!data_buf) + return -ENOMEM; + + scatterwalk_map_and_copy(data_buf, req->src, 0, data_len, 0); + + return cmh_sm3_submit_final(req, data_buf, data_len); +} + +static int cmh_sm3_export(struct ahash_request *req, void *out) +{ + struct cmh_sm3_reqctx *rctx =3D ahash_request_ctx(req); + struct cmh_sm3_export_state *state =3D out; + + /* + * Zero the whole exported state first: the struct may carry padding, + * so without this the padding would leak kernel memory to user space + * through the ahash export. + */ + memset(state, 0, sizeof(*state)); + + if (rctx->hw_started && rctx->has_checkpoint) + memcpy(state->checkpoint, rctx->checkpoint, SM3_CONTEXT_SIZE); + + state->hw_started =3D rctx->hw_started; + + return 0; +} + +static int cmh_sm3_import(struct ahash_request *req, const void *in) +{ + struct cmh_sm3_reqctx *rctx =3D ahash_request_ctx(req); + const struct cmh_sm3_export_state *state =3D in; + + memset(rctx, 0, sizeof(*rctx)); + + rctx->hw_started =3D state->hw_started; + + if (state->hw_started) { + memcpy(rctx->checkpoint, state->checkpoint, SM3_CONTEXT_SIZE); + rctx->has_checkpoint =3D 1; + } + + return 0; +} + +/* Registration */ + +static struct ahash_alg cmh_sm3_ahash_alg =3D { + .init =3D cmh_sm3_init, + .update =3D cmh_sm3_update, + .finup =3D cmh_sm3_finup, + .digest =3D cmh_sm3_digest, + .export =3D cmh_sm3_export, + .import =3D cmh_sm3_import, + + .halg =3D { + .digestsize =3D CMH_SM3_DIGEST_SIZE, + .statesize =3D sizeof(struct cmh_sm3_export_state), + .base =3D { + .cra_name =3D "sm3", + .cra_driver_name =3D "rambus-cmh-sm3", + .cra_priority =3D 300, + .cra_flags =3D CRYPTO_ALG_KERN_DRIVER_ONLY | + CRYPTO_ALG_NO_FALLBACK | + CRYPTO_ALG_ASYNC | + CRYPTO_AHASH_ALG_BLOCK_ONLY, + .cra_blocksize =3D CMH_SM3_BLOCK_SIZE, + .cra_ctxsize =3D 0, + .cra_reqsize =3D sizeof(struct cmh_sm3_reqctx), + .cra_module =3D THIS_MODULE, + }, + }, +}; + +/** + * cmh_sm3_register() - Register SM3 hash algorithm with the crypto framew= ork + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_sm3_register(void) +{ + int ret; + + if (!cmh_core_present(CMH_CORE_SM3)) + return 0; + + ret =3D crypto_register_ahash(&cmh_sm3_ahash_alg); + if (ret) { + dev_err(cmh_dev(), "sm3: failed to register cmh-sm3 (rc=3D%d)\n", + ret); + return ret; + } + + return 0; +} + +/** + * cmh_sm3_unregister() - Unregister SM3 hash algorithm from the crypto fr= amework + */ +void cmh_sm3_unregister(void) +{ + if (!cmh_core_present(CMH_CORE_SM3)) + return; + + crypto_unregister_ahash(&cmh_sm3_ahash_alg); +} diff --git a/drivers/crypto/cmh/include/cmh_sm3.h b/drivers/crypto/cmh/incl= ude/cmh_sm3.h new file mode 100644 index 000000000000..962078b051b4 --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_sm3.h @@ -0,0 +1,28 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- SM3 Hash Driver + * + * Registers an ahash algorithm for SM3 (GB/T 32905-2016) with the + * Linux crypto subsystem using the CMH SM3 core (CORE_ID_SM3). + * This is a CRYPTO_AHASH_ALG_BLOCK_ONLY driver (same model as + * cmh_hash.c), so the Crypto API buffers partial blocks and hands the + * driver only whole-block-aligned data: + * + * .init() -> software-only: zero per-request context + * .update() -> SM3_CMD_INIT [+ RESTORE] + UPDATE(full blocks) + SAVE + = FLUSH + * (also serves final: the API calls finup with nbytes =3D= =3D 0) + * .digest() -> INIT + UPDATE + FINAL + FLUSH (single-shot) + * .export() -> software-only: copy the SM3 checkpoint + * .import() -> software-only: restore the SM3 checkpoint + */ + +#ifndef CMH_SM3_H +#define CMH_SM3_H + +#include "cmh_config.h" + +int cmh_sm3_register(void); +void cmh_sm3_unregister(void); + +#endif /* CMH_SM3_H */ --=20 2.43.7 From nobody Fri Sep 25 01:20:34 2026 Received: from DM5PR21CU001.outbound.protection.outlook.com (mail-centralusazon11021084.outbound.protection.outlook.com [52.101.62.84]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5997F4B4051; Thu, 17 Sep 2026 22:59:43 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.62.84 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789685997; cv=fail; b=WJ/M/gxRt+jE4AP/sc8NBDWnktmcYLQFeVlRL5OREquoaCHAGmkifqE8D0x5rtLZIZQfGupxmPOY15+WVy291Youzsg/3nYfVEvckdPwP02cdtVoMedfDniRTVeYPIVjeJiBVsTRSdadbCUfzl69yUPVWCPoFOPQMz4ZGQAXVHE= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789685997; c=relaxed/simple; bh=Aep5WfgTvo3/aSyGj9gVJu7GKIShUbk8fFUiiM8xncg=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=sn2XFZqq3kDueTMqE+FtXlXq2e5GCrretjqZD6VJo6eiVSwQ4I6Pnc2xRnf5ZraG3nozG6huqpufIkrcSYlpfnxz5si5MW7Wluh7d6R+jcSzWmZQ7aTc04HinqQ7ZmTko0aSpJujxkoyfpk/GCbsPfRXwohKVQLyq0ugJ/4mS+o= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com; spf=fail smtp.mailfrom=rambus.com; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b=IhxUYk/g; arc=fail smtp.client-ip=52.101.62.84 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=rambus.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b="IhxUYk/g" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=iCI/bvaQiiasi5Bexl5hVZ3VYOSuDLk0z5/994oZFBCn5ngkdENK5UoSJPfnj/TsseSRRM4UrIIHFVC/3aNY1JS0Cu68sBsYJAFMEGTTvLDgRHwblwVHvbfcdI1Q9xopnfAfPySkVCXF+rYLHcAx//3AW/oZmsw4vS27LR12AgnCJ53tqkDYqJbiRxOzkufBEhqpbMN64inNH4C43wLNywySuotC9g3J2SB4IbcnHNJjxV3ERsb0RkrWuynded5lOpLU/Izuowm4lAagS/WLFrEgRy+Ia3w6d2aYVD37i8s3hY4Dg9fmLF3aQt2rw0PIAsCO59uGtJm5vBIEOn+8OQ== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=y3U8pIuVv0W/HGn3YEnW7xBVSl4Fos44ZFPeM3fULd8=; b=A6Y6G2StVsIE9eO0kUNN0KjTBoCGKy128yf983Whk4B9TJ4ZeGVJs/FrBB+PAFZYt2/MnwC3EJyIna8zZdM1KsJREoHC6Vvf1a8cLQqXxD7weIFrTVbvtSbEZ/4hXIQhqPBR1tW3HxfcFa6oN1rUFtEk0LwpGqxN/dABsaDIfvym4X12AHWGkXjAhoAOa0hs4/cs5TrD9OjihhUPunUe0wtBGVbSE8A/IWFwfEaZ4c37SjOsDWyq+8tX0domLIuU09UIIXt4ZfT8KV1fz3W/rz+5lSuinzFJHnFm2i+lzXl/q+3zBGr1XlkBnI7po/D9h3grCXs8Rj05/zzXiL7P1w== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 192.86.86.210) smtp.rcpttodomain=cryptography.com smtp.mailfrom=rambus.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=rambus.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=rambus.com; s=selector1; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=y3U8pIuVv0W/HGn3YEnW7xBVSl4Fos44ZFPeM3fULd8=; b=IhxUYk/gwDwxJJqmenw1OlIiGBS0D/h812Xolyl1rlKGX/64oAZliEKnc5fYndBSWIAl9njJtbE3sL07bRXpqDPFr4SmP1rhFyf12lJTAcTfipXDduXetWficZDarg5L0yzSSy8QQALJ+Gr1BHR8YuMMDa5H8Walc3PbrsABxOD+TqseinBFs2buAAhGJhbcqCnb5ASAoxqyJXfW9hWCWBUqqTMXirT3prD0UaNgW4VJ5qMoEeCOTlT31/6DEubgAMEeFktHiaGC77CkDaZJ9K2wPx5DiO9+XYX3wfOofzCwSZakChfJswKcuw2kuvsqoIHsvmf5PTsf2CD/4phNkg== Received: from SJ0PR13CA0056.namprd13.prod.outlook.com (2603:10b6:a03:2c2::31) by CO6PR04MB7523.namprd04.prod.outlook.com (2603:10b6:303:a8::17) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.428.13; Thu, 17 Sep 2026 22:59:35 +0000 Received: from CO1PEPF000075EF.namprd03.prod.outlook.com (2603:10b6:a03:2c2:cafe::44) by SJ0PR13CA0056.outlook.office365.com (2603:10b6:a03:2c2::31) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.451.7 via Frontend Transport; Thu, 17 Sep 2026 22:59:34 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 192.86.86.210) smtp.mailfrom=rambus.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=rambus.com; Received-SPF: Pass (protection.outlook.com: domain of rambus.com designates 192.86.86.210 as permitted sender) receiver=protection.outlook.com; client-ip=192.86.86.210; helo=hqxsv-psmtppxy02.rambus.com; pr=C Received: from hqxsv-psmtppxy02.rambus.com (192.86.86.210) by CO1PEPF000075EF.mail.protection.outlook.com (10.167.249.38) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.451.8 via Frontend Transport; Thu, 17 Sep 2026 22:59:33 +0000 Received: from hqxsv-cmdev3-aousherovitch.rambus.com (hqn-lb-int-float.rambus.com [10.12.20.20]) by hqxsv-psmtppxy02.rambus.com (Postfix) with ESMTP id 21F111801764; Thu, 17 Sep 2026 22:59:32 +0000 (UTC) From: Alex Ousherovitch To: Albert Ou , Alex Ousherovitch , Conor Dooley , "David S. Miller" , Herbert Xu , Jonathan Corbet , Krzysztof Kozlowski , Palmer Dabbelt , Paul Walmsley , Rob Herring , Saravanakrishnan Krishnamoorthy , Shuah Khan Cc: Alexandre Ghiti , devicetree@vger.kernel.org, Joel Wittenauer , linux-api@vger.kernel.org, linux-crypto@vger.kernel.org, linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-riscv@lists.infradead.org, Shuah Khan , Thi Nguyen Subject: [PATCH v5 08/19] crypto: cmh - add AES skcipher/aead/cmac Date: Thu, 17 Sep 2026 15:59:17 -0700 Message-ID: <20260917225929.2494111-9-aousherovitch@rambus.com> X-Mailer: git-send-email 2.43.7 In-Reply-To: <20260917225929.2494111-1-aousherovitch@rambus.com> References: <20260917225929.2494111-1-aousherovitch@rambus.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: CO1PEPF000075EF:EE_|CO6PR04MB7523:EE_ X-MS-Office365-Filtering-Correlation-Id: ece18a2e-70e7-4f62-68cc-08df150f5720 X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|23010399003|1800799024|376014|7416014|82310400026|36860700016|921020|22082099003|18002099003|56012099006|3023799007|11063799006|10067099003|6133799003; X-Microsoft-Antispam-Message-Info: IPSb4mRHSPh7FKz5iuypDBiDTEbD9hcSkoZpRiXZ4oGwRRX2eV+zwn4fymAzqyWhp7sDgs3Pvn1GO/TzQZfvS9FadCHQZ5fYoeRWv1SqSWpPqKH6E9mzRzl+eEFok7G7uGdBz9jWQxhhkhlqnxkrzFp91Mub5T7EN9/2zBawMGzz6OcLO1OsnuSKcRIqkh5aQvV2A3xmfemCb2A7+OpCnIcLWVig1q2xT8kdNAOdRs8eiG38iXniQLAbPoStfKGpN67l6H/NdKCGUImLxphL++FOpTLXlzBHzMdDKb/OHH6sAFnFhG4fIS8/C4RAoTZMzolO/iYrDZP4sXYNR+9h4zSo4skpoyNs4CYaCUU8YVmuBYRdGunjsqTb4BWTrTpz7kExAVXXQdkn7gMKK1Wujkv4Irqt27t97QG7EWZkgUrRZGKzea+QAgY7xvoxLJh5KLU8H89QYD4UfCASxL7jthjGS8RITWd/P1aysfucFr/qjaMmzT1O9VXEA1J77Tvi2vpePTh3uzuT9qWdtCKIdjpJzh0VTxFTp8BZSugnyqvqBqFFL6Jcp6hnikMIv8D/ivu0iNWry6w66bBhoUZo4lbH/olaJnDJ7FyiJvLct1pHCQskadskHwlZO+gpdQKWjDXJD6TargcSjY8/ZE7MzFDOdy7G3tyXrMjiQjPe8ZWOvNGnusKOUcbSXZiFr9a54bYNC13tk77/2rFT9ClzGKE2tGUGWY94pR/0eUFd9oN/wGSiwSbm+f3q5MNWjobc X-Forefront-Antispam-Report: CIP:192.86.86.210;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:hqxsv-psmtppxy02.rambus.com;PTR:InfoDomainNonexistent;CAT:NONE;SFS:(13230040)(23010399003)(1800799024)(376014)(7416014)(82310400026)(36860700016)(921020)(22082099003)(18002099003)(56012099006)(3023799007)(11063799006)(10067099003)(6133799003);DIR:OUT;SFP:1102; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: xYF7OOrUNRXrjrgMYOVkjPuKYw+ftvcV65Q8ccKfSWk+f4E+YpEO7JYPRz+pNUfcednq3KeMQAODTCNnvsJS8Hqv6KjK4cTyNwRxsZQi65JlW74E1YXSn6kurC7BJJ3WihmL4j4mlRb5ZJEhAResSHLe+gHRWpORuG7idmDBuX9lZvIUA1MfVihdhHS8DTKpjFjwBmzGEh8EDQPblhacXTAIseXwDBqzxIZtA3bf+HDl6zLvMMb89y3/ifVlaNS0xLfqDcJ3suCI+JH+PyMaRERmC2q3rxUmEiCq6LzAdKaITF4owGchg6DSpOXn2bMgOZgfz9zxDxXYQoNodQR77iOWW0+Ibu2v1m6gaOt5uppI9usfF16I72KhQhqMuRH0drYCe4YxoHMewaSonY/k0jzegQJ3kt6Wu3uLiz+MQIAvjkhssFzUA2MtoRiBias7 X-OriginatorOrg: rambus.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 17 Sep 2026 22:59:33.9212 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: ece18a2e-70e7-4f62-68cc-08df150f5720 X-MS-Exchange-CrossTenant-Id: bd0ba799-c2b9-413c-9c56-5d1731c4827c X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=bd0ba799-c2b9-413c-9c56-5d1731c4827c;Ip=[192.86.86.210];Helo=[hqxsv-psmtppxy02.rambus.com] X-MS-Exchange-CrossTenant-AuthSource: CO1PEPF000075EF.namprd03.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: CO6PR04MB7523 Content-Type: text/plain; charset="utf-8" Register AES algorithms using the CMH AES core (core ID 0x03): - skcipher: AES-ECB, AES-CBC, AES-CTR, AES-XTS, AES-CFB - aead: AES-GCM, AES-CCM - ahash: AES-CMAC Supports 128, 192, and 256-bit keys. AEAD algorithms handle associated data, payload, and authentication tag with correct encrypt/decrypt separation. Signed-off-by: Alex Ousherovitch Co-developed-by: Saravanakrishnan Krishnamoorthy Signed-off-by: Saravanakrishnan Krishnamoorthy --- drivers/crypto/cmh/Makefile | 5 +- drivers/crypto/cmh/cmh_aes.c | 755 +++++++++++++++++++ drivers/crypto/cmh/cmh_aes_aead.c | 1013 ++++++++++++++++++++++++++ drivers/crypto/cmh/cmh_aes_cmac.c | 741 +++++++++++++++++++ drivers/crypto/cmh/cmh_main.c | 25 + drivers/crypto/cmh/include/cmh_aes.h | 24 + 6 files changed, 2562 insertions(+), 1 deletion(-) create mode 100644 drivers/crypto/cmh/cmh_aes.c create mode 100644 drivers/crypto/cmh/cmh_aes_aead.c create mode 100644 drivers/crypto/cmh/cmh_aes_cmac.c create mode 100644 drivers/crypto/cmh/include/cmh_aes.h diff --git a/drivers/crypto/cmh/Makefile b/drivers/crypto/cmh/Makefile index 664b2a20bc19..c5dcfa4f7794 100644 --- a/drivers/crypto/cmh/Makefile +++ b/drivers/crypto/cmh/Makefile @@ -19,7 +19,10 @@ cmh-y :=3D \ cmh_hmac.o \ cmh_cshake.o \ cmh_kmac.o \ - cmh_sm3.o + cmh_sm3.o \ + cmh_aes.o \ + cmh_aes_aead.o \ + cmh_aes_cmac.o =20 # Management ioctl device (/dev/cmh_mgmt): key lifecycle, PKE, PQC ioctls. cmh-$(CONFIG_CRYPTO_DEV_CMH_MGMT) +=3D \ diff --git a/drivers/crypto/cmh/cmh_aes.c b/drivers/crypto/cmh/cmh_aes.c new file mode 100644 index 000000000000..7d15b58d7194 --- /dev/null +++ b/drivers/crypto/cmh/cmh_aes.c @@ -0,0 +1,755 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Kernel Crypto API AES (skcipher) Driver + * + * Registers skcipher algorithms with the Linux crypto subsystem: + * ecb(aes), cbc(aes), ctr(aes), cfb(aes), xts(aes) + * + * Uses the CMH AES Core via VCQ commands: + * [SYS_CMD_WRITE] + AES_CMD_INIT + [AES_CMD_UPDATE] + AES_CMD_FINAL + * + VCQ_CMD_FLUSH + * + * The AES core requires bidirectional DMA -- both input and output + * buffers are mapped and passed in a single AES_CMD_FINAL command. + * + * Raw-key atomicity: SYS_CMD_WRITE to SYS_REF_TEMP is packed into + * the same VCQ as AES commands (see cmh_key.h for details). + * + */ + +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "cmh_aes.h" +#include "cmh_vcq.h" +#include "cmh_aes_abi.h" +#include "cmh_sys_abi.h" +#include "cmh_sys.h" +#include "cmh_txn.h" +#include "cmh_dma.h" +#include "cmh_key.h" + +/* Algorithm Table */ + +struct cmh_aes_alg_info { + u32 aes_mode; /* AES_MODE_* */ + u32 ivsize; /* bytes (0 for ECB) */ + u32 min_keysize; /* minimum key bytes */ + u32 max_keysize; /* maximum key bytes */ + const char *alg_name; /* Linux crypto name: "ecb(aes)" */ + const char *drv_name; /* driver name: "rambus-cmh-ecb-aes" */ +}; + +static const struct cmh_aes_alg_info aes_algs[] =3D { + { AES_MODE_ECB, 0, AES_KEYSIZE_128, AES_KEYSIZE_256, + "ecb(aes)", "rambus-cmh-ecb-aes" }, + { AES_MODE_CBC, CMH_AES_IV_SIZE, AES_KEYSIZE_128, AES_KEYSIZE_256, + "cbc(aes)", "rambus-cmh-cbc-aes" }, + { AES_MODE_CTR, CMH_AES_IV_SIZE, AES_KEYSIZE_128, AES_KEYSIZE_256, + "ctr(aes)", "rambus-cmh-ctr-aes" }, + { AES_MODE_CFB, CMH_AES_IV_SIZE, AES_KEYSIZE_128, AES_KEYSIZE_256, + "cfb(aes)", "rambus-cmh-cfb-aes" }, + { AES_MODE_XTS, CMH_AES_IV_SIZE, 2 * AES_KEYSIZE_128, 2 * AES_KEYSIZE_25= 6, + "xts(aes)", "rambus-cmh-xts-aes" }, +}; + +/* Per-transform context (allocated by crypto framework) */ + +struct cmh_aes_tfm_ctx { + struct cmh_key_ctx key; +}; + +/* Per-request context (lives in skcipher_request::__ctx) */ + +/* + * Maximum payload commands: + * [SYS_CMD_WRITE] + AES_CMD_INIT + [AES_CMD_UPDATE] + AES_CMD_FINAL + * + VCQ_CMD_FLUSH =3D 5 + * UPDATE is used for XTS data > 2 blocks (see cmh_aes_crypt). + */ +#define CMH_AES_MAX_PAYLOAD 5 +#define CMH_AES_MAX_PACKED (CMH_AES_MAX_PAYLOAD * 2) + +struct cmh_aes_reqctx { + dma_addr_t in_dma; + dma_addr_t out_dma; + dma_addr_t iv_dma; + dma_addr_t iv2_dma; + dma_addr_t key_dma; + u8 *in_buf; + u8 *out_buf; + u8 *iv_buf; + u8 *iv2_buf; + u32 cryptlen; + u32 ivsize; + u32 keylen; + u32 aes_mode; + u32 aes_op; + /* CTR counter-wrap split state */ + u32 ctr_chunk1_len; + u32 core_id; + s32 target_mbx; + u64 key_ref; + struct vcq_cmd packed[CMH_AES_MAX_PACKED]; +}; + +/* VCQ Builders -- AES-specific */ + +static void vcq_add_aes_init(struct vcq_cmd *slot, u32 core_id, u64 key_re= f, u64 iv_dma, + u32 keylen, u32 ivlen, u32 mode, u32 op, + u32 iolen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, AES_CMD_INIT); + slot->hwc.aes.cmd_init.key =3D key_ref; + slot->hwc.aes.cmd_init.iv =3D iv_dma; + slot->hwc.aes.cmd_init.keylen =3D keylen; + slot->hwc.aes.cmd_init.ivlen =3D ivlen; + slot->hwc.aes.cmd_init.mode =3D mode; + slot->hwc.aes.cmd_init.op =3D op; + slot->hwc.aes.cmd_init.aadlen =3D 0; + slot->hwc.aes.cmd_init.iolen =3D iolen; + slot->hwc.aes.cmd_init.taglen =3D 0; +} + +static void vcq_add_aes_update(struct vcq_cmd *slot, u32 core_id, u64 inpu= t_dma, + u64 output_dma, u32 iolen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, AES_CMD_UPDATE); + slot->hwc.aes.cmd_update.input =3D input_dma; + slot->hwc.aes.cmd_update.output =3D output_dma; + slot->hwc.aes.cmd_update.iolen =3D iolen; +} + +static void vcq_add_aes_final(struct vcq_cmd *slot, u32 core_id, u64 input= _dma, + u64 output_dma, u32 iolen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, AES_CMD_FINAL); + slot->hwc.aes.cmd_final.input =3D input_dma; + slot->hwc.aes.cmd_final.output =3D output_dma; + slot->hwc.aes.cmd_final.iolen =3D iolen; + slot->hwc.aes.cmd_final.tag =3D 0; + slot->hwc.aes.cmd_final.taglen =3D 0; +} + +/* + * We wrap each skcipher_alg with its info pointer in a compound struct, + * then use container_of() in cmh_aes_get_info() to recover it. + * This is the same pattern used by hash, hmac, cshake, kmac. + */ +struct cmh_aes_alg_drv { + struct skcipher_alg alg; + const struct cmh_aes_alg_info *info; +}; + +static bool aes_is_stream_mode(u32 mode) +{ + return mode =3D=3D AES_MODE_CTR || mode =3D=3D AES_MODE_CFB; +} + +/* + * Update req->iv after a successful encrypt/decrypt. + * + * The Linux skcipher API contract requires that req->iv is updated to + * reflect the state needed to continue processing in a chained call: + * CBC encrypt: IV <- last ciphertext block + * CBC decrypt: IV <- last ciphertext block of the *input* + * CTR: IV <- counter incremented by ceil(cryptlen / blocksize) + * CFB: IV <- last ciphertext block + */ +static void cmh_aes_update_iv(struct skcipher_request *req, u32 mode, + u32 op, const u8 *in_buf, const u8 *out_buf) +{ + u32 bs =3D CMH_AES_BLOCK_SIZE; + u32 nblocks; + + switch (mode) { + case AES_MODE_CBC: + if (op =3D=3D AES_OP_ENCRYPT) + memcpy(req->iv, out_buf + req->cryptlen - bs, bs); + else + memcpy(req->iv, in_buf + req->cryptlen - bs, bs); + break; + case AES_MODE_CTR: + /* + * Arithmetic big-endian 128-bit counter increment. + * Process from the least-significant byte (index 15) + * upward, carrying as needed. + */ + nblocks =3D DIV_ROUND_UP(req->cryptlen, bs); + { + u8 *iv =3D req->iv; + int i; + + for (i =3D bs - 1; i >=3D 0 && nblocks; i--) { + u32 sum =3D (u32)iv[i] + (nblocks & 0xff); + + iv[i] =3D (u8)sum; + nblocks =3D (nblocks >> 8) + (sum >> 8); + } + } + break; + case AES_MODE_CFB: + /* + * CFB-128 chains on the last ciphertext block. On encrypt, + * that is out_buf; on decrypt, it is in_buf. + * + * For sub-block requests (cryptlen < 16), there is no + * complete ciphertext block to chain, so the IV is left + * unchanged -- CFB-128 has no defined chaining semantic + * for partial blocks (shift-register CFB-n is a different + * mode). Without this guard the pointer arithmetic + * underflows and reads before the buffer. + */ + if (req->cryptlen >=3D bs) { + if (op =3D=3D AES_OP_ENCRYPT) + memcpy(req->iv, out_buf + req->cryptlen - bs, + bs); + else + memcpy(req->iv, in_buf + req->cryptlen - bs, + bs); + } + break; + default: + break; + } +} + +/* skcipher Operations */ + +static const struct cmh_aes_alg_info * +cmh_aes_get_info(struct crypto_skcipher *tfm) +{ + struct skcipher_alg *alg =3D crypto_skcipher_alg(tfm); + + return container_of(alg, struct cmh_aes_alg_drv, alg)->info; +} + +static int cmh_aes_setkey(struct crypto_skcipher *tfm, const u8 *key, + unsigned int keylen) +{ + struct cmh_aes_tfm_ctx *tctx =3D crypto_skcipher_ctx(tfm); + const struct cmh_aes_alg_info *info =3D cmh_aes_get_info(tfm); + + if (info->aes_mode =3D=3D AES_MODE_XTS) { + int err; + + /* XTS: double key (32, 48, or 64 bytes) */ + if (keylen !=3D 2 * AES_KEYSIZE_128 && + keylen !=3D 2 * AES_KEYSIZE_192 && + keylen !=3D 2 * AES_KEYSIZE_256) + return -EINVAL; + err =3D xts_verify_key(tfm, key, keylen); + if (err) + return err; + } else { + /* Standard: 16, 24, or 32 bytes */ + if (keylen !=3D AES_KEYSIZE_128 && + keylen !=3D AES_KEYSIZE_192 && + keylen !=3D AES_KEYSIZE_256) + return -EINVAL; + } + + return cmh_key_setkey_raw(&tctx->key, key, keylen, CORE_ID_AES); +} + +static int cmh_aes_init_tfm(struct crypto_skcipher *tfm) +{ + struct cmh_aes_tfm_ctx *tctx =3D crypto_skcipher_ctx(tfm); + + memset(tctx, 0, sizeof(*tctx)); + crypto_skcipher_set_reqsize(tfm, sizeof(struct cmh_aes_reqctx)); + return 0; +} + +static void cmh_aes_exit_tfm(struct crypto_skcipher *tfm) +{ + struct cmh_aes_tfm_ctx *tctx =3D crypto_skcipher_ctx(tfm); + + cmh_key_destroy(&tctx->key); +} + +#define CMH_AES_MAX_CRYPTLEN SZ_32M + +/* DMA unmap helper */ +static void cmh_aes_unmap_dma(struct cmh_aes_reqctx *rctx) +{ + if (rctx->iv2_buf) + cmh_dma_unmap_single(rctx->iv2_dma, rctx->ivsize, + DMA_TO_DEVICE); + if (rctx->ivsize > 0) + cmh_dma_unmap_single(rctx->iv_dma, rctx->ivsize, + DMA_TO_DEVICE); + cmh_dma_unmap_single(rctx->out_dma, rctx->cryptlen, DMA_FROM_DEVICE); + cmh_dma_unmap_single(rctx->in_dma, rctx->cryptlen, DMA_TO_DEVICE); +} + +static void cmh_aes_free_bufs(struct cmh_aes_reqctx *rctx) +{ + kfree(rctx->iv2_buf); + rctx->iv2_buf =3D NULL; + kfree(rctx->iv_buf); + rctx->iv_buf =3D NULL; + kfree_sensitive(rctx->out_buf); + rctx->out_buf =3D NULL; + kfree_sensitive(rctx->in_buf); + rctx->in_buf =3D NULL; +} + +/* + * Submit the second CTR chunk after the first completes. + * Called from cmh_aes_complete when ctr_chunk1_len > 0. + */ +static void cmh_aes_complete(void *data, int error); + +static int cmh_aes_ctr_submit_chunk2(struct skcipher_request *req) +{ + struct crypto_skcipher *tfm =3D crypto_skcipher_reqtfm(req); + struct cmh_aes_tfm_ctx *tctx =3D crypto_skcipher_ctx(tfm); + struct cmh_aes_reqctx *rctx =3D skcipher_request_ctx(req); + struct vcq_cmd cmds[CMH_AES_MAX_PAYLOAD]; + u32 chunk1 =3D rctx->ctr_chunk1_len; + u32 chunk2 =3D rctx->cryptlen - chunk1; + u64 key_ref; + u32 keylen; + u32 idx =3D 0; + + /* Clear split flag so next completion is final */ + rctx->ctr_chunk1_len =3D 0; + + vcq_add_sys_write(&cmds[idx++], SYS_REF_TEMP, + (u64)rctx->key_dma, SYS_REF_NONE, + tctx->key.raw.len, + tctx->key.raw.sys_type); + key_ref =3D SYS_REF_TEMP; + keylen =3D tctx->key.raw.len; + + vcq_add_aes_init(&cmds[idx++], rctx->core_id, key_ref, + (u64)rctx->iv2_dma, keylen, rctx->ivsize, + rctx->aes_mode, rctx->aes_op, 0); + vcq_add_aes_final(&cmds[idx++], rctx->core_id, + (u64)(rctx->in_dma + chunk1), + (u64)(rctx->out_dma + chunk1), chunk2); + vcq_add_flush(&cmds[idx++], rctx->core_id); + + return cmh_vcq_pack_and_submit_async(cmds, idx, rctx->packed, + CMH_AES_MAX_PACKED, + rctx->target_mbx, + cmh_aes_complete, req, + !!(req->base.flags & + CRYPTO_TFM_REQ_MAY_BACKLOG), + cmh_tm_async_timeout_jiffies()); +} + +/* + * Async completion callback -- fires from RH threaded IRQ context. + * + * Unmaps DMA buffers, copies output to req->dst scatterlist, + * updates the IV state, frees temporaries, and completes the request. + * + * For CTR counter-wrap splits, the first chunk completion chains + * into a second VCQ submission rather than finalizing immediately. + */ +static void cmh_aes_complete(void *data, int error) +{ + struct skcipher_request *req =3D data; + struct cmh_aes_reqctx *rctx =3D skcipher_request_ctx(req); + + if (error =3D=3D -EINPROGRESS) { + cmh_complete(&req->base, error); + return; + } + + /* + * CTR counter-wrap: first chunk completed, submit second. + * DMA mappings remain valid (they cover the full buffer). + * + * Recursion depth bounded: chunk2 clears ctr_chunk1_len before + * submission, so the second cmh_aes_complete invocation sees 0 + * and finalizes (max depth =3D 2). + */ + if (rctx->ctr_chunk1_len && !error) { + int ret; + + ret =3D cmh_aes_ctr_submit_chunk2(req); + + if (!ret || ret =3D=3D -EBUSY) + return; + /* Submission failed; clean up below */ + error =3D ret; + } + + cmh_aes_unmap_dma(rctx); + + if (!error) { + scatterwalk_map_and_copy(rctx->out_buf, req->dst, + 0, rctx->cryptlen, 1); + cmh_aes_update_iv(req, rctx->aes_mode, rctx->aes_op, + rctx->in_buf, rctx->out_buf); + } + + cmh_aes_free_bufs(rctx); + cmh_complete(&req->base, error); +} + +/* + * Core encrypt/decrypt -- builds a VCQ transaction and submits async. + * + * Returns -EINPROGRESS on successful submission (completion callback + * will fire later). Returns 0 for trivial cases (zero-length). + * Returns negative errno on pre-submission errors. + */ +static int cmh_aes_crypt(struct skcipher_request *req, u32 aes_op) +{ + struct crypto_skcipher *tfm =3D crypto_skcipher_reqtfm(req); + struct cmh_aes_tfm_ctx *tctx =3D crypto_skcipher_ctx(tfm); + const struct cmh_aes_alg_info *info =3D cmh_aes_get_info(tfm); + struct cmh_aes_reqctx *rctx =3D skcipher_request_ctx(req); + struct vcq_cmd cmds[CMH_AES_MAX_PAYLOAD]; + u64 key_ref; + u32 keylen; + struct core_dispatch d; + s32 target_mbx; + u32 core_id; + u32 idx; + int ret; + gfp_t gfp; + + if (tctx->key.mode =3D=3D CMH_KEY_NONE) + return -ENOKEY; + + if (!req->cryptlen) + return 0; + + if (req->cryptlen > CMH_AES_MAX_CRYPTLEN) + return -EINVAL; + + switch (info->aes_mode) { + case AES_MODE_CTR: + case AES_MODE_CFB: + break; + case AES_MODE_XTS: + if (req->cryptlen < CMH_AES_BLOCK_SIZE) + return -EINVAL; + break; + default: + if (req->cryptlen & (CMH_AES_BLOCK_SIZE - 1)) + return -EINVAL; + break; + } + + gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + + /* Initialise reqctx */ + memset(rctx, 0, sizeof(*rctx)); + rctx->cryptlen =3D req->cryptlen; + rctx->ivsize =3D info->ivsize; + rctx->aes_mode =3D info->aes_mode; + rctx->aes_op =3D aes_op; + rctx->iv2_buf =3D NULL; + + /* + * Linearise input from scatterlist. cryptlen is user-controlled up + * to CMH_AES_MAX_CRYPTLEN (well above KMALLOC_MAX_SIZE), so use + * __GFP_NOWARN: an oversized request fails cleanly with -ENOMEM + * instead of splatting the page allocator. + */ + rctx->in_buf =3D kmalloc(req->cryptlen, gfp | __GFP_NOWARN); + if (!rctx->in_buf) + return -ENOMEM; + + scatterwalk_map_and_copy(rctx->in_buf, req->src, 0, req->cryptlen, 0); + + rctx->in_dma =3D cmh_dma_map_single(rctx->in_buf, req->cryptlen, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->in_dma)) { + ret =3D -ENOMEM; + goto out_free_in; + } + + /* Allocate and map output buffer */ + rctx->out_buf =3D kmalloc(req->cryptlen, gfp | __GFP_NOWARN); + if (!rctx->out_buf) { + ret =3D -ENOMEM; + goto out_unmap_in; + } + + rctx->out_dma =3D cmh_dma_map_single(rctx->out_buf, req->cryptlen, + DMA_FROM_DEVICE); + if (cmh_dma_map_error(rctx->out_dma)) { + ret =3D -ENOMEM; + goto out_free_out; + } + + /* Map IV if required */ + if (info->ivsize > 0) { + rctx->iv_buf =3D kmemdup(req->iv, info->ivsize, gfp); + if (!rctx->iv_buf) { + ret =3D -ENOMEM; + goto out_unmap_out; + } + rctx->iv_dma =3D cmh_dma_map_single(rctx->iv_buf, info->ivsize, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->iv_dma)) { + ret =3D -ENOMEM; + goto out_free_iv; + } + } + + /* Resolve key reference */ + idx =3D 0; + + rctx->key_dma =3D tctx->key.raw.dma; + rctx->keylen =3D tctx->key.raw.len; + vcq_add_sys_write(&cmds[idx++], SYS_REF_TEMP, + (u64)rctx->key_dma, SYS_REF_NONE, + tctx->key.raw.len, + tctx->key.raw.sys_type); + key_ref =3D SYS_REF_TEMP; + keylen =3D tctx->key.raw.len; + d =3D cmh_core_select_instance(CMH_CORE_AES); + target_mbx =3D d.mbx_idx; + core_id =3D d.core_id; + + /* + * iolen in INIT: XTS needs total length upfront for tweak + * computation; all other modes use 0 (streaming). + */ + vcq_add_aes_init(&cmds[idx++], core_id, key_ref, (u64)rctx->iv_dma, + keylen, info->ivsize, info->aes_mode, aes_op, + info->aes_mode =3D=3D AES_MODE_XTS ? + req->cryptlen : 0); + + if (info->aes_mode =3D=3D AES_MODE_XTS && + req->cryptlen > 2 * CMH_AES_BLOCK_SIZE) { + u32 final_len, update_len; + + if (req->cryptlen & (CMH_AES_BLOCK_SIZE - 1)) + final_len =3D CMH_AES_BLOCK_SIZE + + (req->cryptlen & (CMH_AES_BLOCK_SIZE - 1)); + else + final_len =3D 2 * CMH_AES_BLOCK_SIZE; + + update_len =3D req->cryptlen - final_len; + + vcq_add_aes_update(&cmds[idx++], core_id, + (u64)rctx->in_dma, + (u64)rctx->out_dma, update_len); + vcq_add_aes_final(&cmds[idx++], core_id, + (u64)(rctx->in_dma + update_len), + (u64)(rctx->out_dma + update_len), + final_len); + } else if (info->aes_mode =3D=3D AES_MODE_CTR) { + /* + * CTR counter-wrap workaround: + * The AES-SCA hardware uses a 64-bit block counter. + * If the lower 64 bits of the IV would wrap during + * this operation, split into two separate VCQ + * transactions -- the completion callback for the + * first chunk submits the second. + */ + u64 lower64 =3D get_unaligned_be64(rctx->iv_buf + 8); + u32 nblocks =3D DIV_ROUND_UP(req->cryptlen, + CMH_AES_BLOCK_SIZE); + u64 bwrap =3D lower64 ? (~lower64 + 1ULL) : U64_MAX; + + if (nblocks > bwrap) { + u32 chunk1 =3D (u32)bwrap * CMH_AES_BLOCK_SIZE; + u64 upper64; + + /* Prepare second IV for chained submission */ + rctx->iv2_buf =3D kmalloc(info->ivsize, gfp); + if (!rctx->iv2_buf) { + ret =3D -ENOMEM; + goto out_unmap_iv; + } + upper64 =3D get_unaligned_be64(rctx->iv_buf); + put_unaligned_be64(upper64 + 1, rctx->iv2_buf); + put_unaligned_be64(0, rctx->iv2_buf + 8); + + rctx->iv2_dma =3D + cmh_dma_map_single(rctx->iv2_buf, + info->ivsize, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->iv2_dma)) { + ret =3D -ENOMEM; + goto out_free_iv2; + } + + /* Store state for the chained second submission */ + rctx->ctr_chunk1_len =3D chunk1; + rctx->core_id =3D core_id; + rctx->target_mbx =3D target_mbx; + rctx->key_ref =3D key_ref; + + /* First transaction: only chunk1 */ + vcq_add_aes_final(&cmds[idx++], core_id, + (u64)rctx->in_dma, + (u64)rctx->out_dma, chunk1); + } else { + /* No wrap: single FINAL with all data */ + vcq_add_aes_final(&cmds[idx++], core_id, + (u64)rctx->in_dma, + (u64)rctx->out_dma, + req->cryptlen); + } + } else { + vcq_add_aes_final(&cmds[idx++], core_id, + (u64)rctx->in_dma, + (u64)rctx->out_dma, req->cryptlen); + } + + vcq_add_flush(&cmds[idx++], core_id); + + ret =3D cmh_vcq_pack_and_submit_async(cmds, idx, rctx->packed, + CMH_AES_MAX_PACKED, target_mbx, + cmh_aes_complete, req, + !!(req->base.flags & + CRYPTO_TFM_REQ_MAY_BACKLOG), + cmh_tm_async_timeout_jiffies()); + if (ret =3D=3D -EBUSY) + return -EBUSY; + if (ret) + goto out_cleanup_all; + + return -EINPROGRESS; + +out_cleanup_all: + if (rctx->iv2_buf) { + cmh_dma_unmap_single(rctx->iv2_dma, info->ivsize, + DMA_TO_DEVICE); + } +out_free_iv2: + kfree(rctx->iv2_buf); +out_unmap_iv: + if (info->ivsize > 0) + cmh_dma_unmap_single(rctx->iv_dma, info->ivsize, + DMA_TO_DEVICE); +out_free_iv: + kfree(rctx->iv_buf); +out_unmap_out: + cmh_dma_unmap_single(rctx->out_dma, req->cryptlen, DMA_FROM_DEVICE); +out_free_out: + kfree_sensitive(rctx->out_buf); +out_unmap_in: + cmh_dma_unmap_single(rctx->in_dma, req->cryptlen, DMA_TO_DEVICE); +out_free_in: + kfree_sensitive(rctx->in_buf); + return ret; +} + +static int cmh_aes_encrypt(struct skcipher_request *req) +{ + return cmh_aes_crypt(req, AES_OP_ENCRYPT); +} + +static int cmh_aes_decrypt(struct skcipher_request *req) +{ + return cmh_aes_crypt(req, AES_OP_DECRYPT); +} + +/* Registration */ + +static struct cmh_aes_alg_drv aes_drv_algs[ARRAY_SIZE(aes_algs)]; + +/** + * cmh_aes_register() - Register AES-CBC/CTR/ECB/XTS skcipher algorithms w= ith the crypto framework + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_aes_register(void) +{ + unsigned int i; + int ret; + + if (!cmh_core_present(CMH_CORE_AES)) + return 0; + + for (i =3D 0; i < ARRAY_SIZE(aes_algs); i++) { + const struct cmh_aes_alg_info *info =3D &aes_algs[i]; + struct cmh_aes_alg_drv *drv =3D &aes_drv_algs[i]; + struct skcipher_alg *alg =3D &drv->alg; + + drv->info =3D info; + + memset(alg, 0, sizeof(*alg)); + + alg->setkey =3D cmh_aes_setkey; + alg->encrypt =3D cmh_aes_encrypt; + alg->decrypt =3D cmh_aes_decrypt; + alg->init =3D cmh_aes_init_tfm; + alg->exit =3D cmh_aes_exit_tfm; + alg->min_keysize =3D info->min_keysize; + alg->max_keysize =3D info->max_keysize; + alg->ivsize =3D info->ivsize; + + strscpy(alg->base.cra_name, info->alg_name, + CRYPTO_MAX_ALG_NAME); + strscpy(alg->base.cra_driver_name, info->drv_name, + CRYPTO_MAX_ALG_NAME); + alg->base.cra_priority =3D 300; + alg->base.cra_flags =3D CRYPTO_ALG_KERN_DRIVER_ONLY | + CRYPTO_ALG_ASYNC; + alg->base.cra_blocksize =3D aes_is_stream_mode(info->aes_mode) + ? 1 : CMH_AES_BLOCK_SIZE; + /* + * Stream modes (CTR/CFB) use cra_blocksize=3D1 but consume a + * full cipher block per counter/feedback step; declare + * chunksize so the crypto layer buffers sub-block data instead + * of splitting the keystream across requests. + */ + if (aes_is_stream_mode(info->aes_mode)) + alg->chunksize =3D CMH_AES_BLOCK_SIZE; + alg->base.cra_ctxsize =3D sizeof(struct cmh_aes_tfm_ctx); + alg->base.cra_module =3D THIS_MODULE; + + ret =3D crypto_register_skcipher(alg); + if (ret) { + dev_err(cmh_dev(), "cmh_aes: failed to register %s (rc=3D%d)\n", + info->alg_name, ret); + goto err_unregister; + } + + dev_dbg(cmh_dev(), "cmh_aes: registered %s\n", info->alg_name); + } + + return 0; + +err_unregister: + while (i--) + crypto_unregister_skcipher(&aes_drv_algs[i].alg); + return ret; +} + +/** + * cmh_aes_unregister() - Unregister AES skcipher algorithms from the cryp= to framework + */ +void cmh_aes_unregister(void) +{ + unsigned int i; + + if (!cmh_core_present(CMH_CORE_AES)) + return; + + for (i =3D 0; i < ARRAY_SIZE(aes_algs); i++) { + crypto_unregister_skcipher(&aes_drv_algs[i].alg); + dev_dbg(cmh_dev(), "cmh_aes: unregistered %s\n", aes_algs[i].alg_name); + } +} diff --git a/drivers/crypto/cmh/cmh_aes_aead.c b/drivers/crypto/cmh/cmh_aes= _aead.c new file mode 100644 index 000000000000..ed2a0b9449bc --- /dev/null +++ b/drivers/crypto/cmh/cmh_aes_aead.c @@ -0,0 +1,1013 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Kernel Crypto API AES AEAD Driver (GCM/CCM) + * + * Registers AEAD algorithms with the Linux crypto subsystem: + * gcm(aes), ccm(aes) + * + * GCM: AES_CMD_INIT(mode=3DGCM) + [AAD_FINAL] + AES_CMD_FINAL + FLUSH + * - Standard 12-byte IV (nonce), 16-byte tag + * - AES_CMD_INIT carries aadlen/iolen/taglen + * - AES_CMD_FINAL carries tag DMA for encrypt (produce) / decrypt (veri= fy) + * + * CCM: AES_CMD_CCM_INIT + [AAD_FINAL] + AES_CMD_FINAL + FLUSH + * - Variable nonce (7--13 bytes), variable tag (4--16 bytes) + * - Uses AES_CMD_CCM_INIT (0x0A) with aes_cmd_init struct + * - Nonce passed via IV field, taglen in init + */ + +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "cmh_aes.h" +#include "cmh_vcq.h" +#include "cmh_aes_abi.h" +#include "cmh_sys_abi.h" +#include "cmh_sys.h" +#include "cmh_txn.h" +#include "cmh_dma.h" +#include "cmh_key.h" + +/* + * GCM IV contract: + * + * The AES core requires exactly 16 bytes loaded into its IV register. + * For standard 96-bit nonce GCM, the driver passes: + * + * IV[0..11] =3D user-supplied 12-byte nonce + * IV[12..15] =3D 0x00000000 + * + * The hardware internally sets the last 32 bits to the big-endian + * counter value 1 (forming J0 =3D nonce || 0x00000001) before + * processing AAD. The driver must NOT pre-set the counter. + * + * If the IV format is incorrect, GCM authentication will fail + * (encrypt produces wrong ciphertext/tag, decrypt rejects). + */ +#define AES_GCM_IV_SIZE 12U /* GCM nonce size (standard) */ +#define AES_GCM_HW_IV_SIZE 16U /* HW requires 16-byte IV buffer */ +#define AES_GCM_TAG_SIZE 16U + +/* CCM: callers pass a 16-byte IV in RFC 3610 format: + * iv[0] =3D L-1, iv[1..14-iv[0]] =3D nonce, rest =3D counter (zeroed). + * Nonce length =3D 14 - iv[0], range 7..13. + */ +#define AES_CCM_IV_SIZE 16U + +enum cmh_aes_aead_type { + CMH_AES_AEAD_GCM, + CMH_AES_AEAD_CCM, +}; + +struct cmh_aes_aead_info { + enum cmh_aes_aead_type type; + u32 aes_mode; /* AES_MODE_GCM or AES_MODE_CCM */ + u32 ivsize; + u32 maxauthsize; + const char *alg_name; + const char *drv_name; +}; + +static const struct cmh_aes_aead_info aes_aead_algs[] =3D { + { CMH_AES_AEAD_GCM, AES_MODE_GCM, AES_GCM_IV_SIZE, + AES_GCM_TAG_SIZE, "gcm(aes)", "rambus-cmh-gcm-aes" }, + { CMH_AES_AEAD_CCM, AES_MODE_CCM, AES_CCM_IV_SIZE, + AES_GCM_TAG_SIZE, "ccm(aes)", "rambus-cmh-ccm-aes" }, +}; + +struct cmh_aes_aead_tfm_ctx { + struct cmh_key_ctx key; + u32 authsize; /* tag length set by setauthsize */ + struct crypto_cipher *sw_cipher; /* CCM empty-input fallback */ + struct crypto_aead *fallback; /* CCM authsize=3D10 fallback */ +}; + +/* Per-request context (lives in aead_request::__ctx) */ + +/* + * Maximum payload commands: + * [SYS_CMD_WRITE] + AES_CMD_INIT + AAD_FINAL + AES_CMD_FINAL + FLUSH = =3D 5 + */ +#define CMH_AES_AEAD_MAX_PAYLOAD 5 +#define CMH_AES_AEAD_MAX_PACKED (CMH_AES_AEAD_MAX_PAYLOAD * 2) + +struct cmh_aes_aead_reqctx { + dma_addr_t in_dma; + dma_addr_t out_dma; + dma_addr_t iv_dma; + dma_addr_t key_dma; + dma_addr_t aad_dma; + dma_addr_t tag_dma; + u8 *in_buf; + u8 *out_buf; + u8 *iv_buf; + u8 *aad_buf; + u8 *tag_buf; + u32 cryptlen; + u32 assoclen; + u32 authsize; + u32 iv_map_len; + u32 keylen; + bool encrypting; + bool empty_gcm_fallback; + struct vcq_cmd packed[CMH_AES_AEAD_MAX_PACKED]; +}; + +struct cmh_aes_aead_drv { + struct aead_alg alg; + const struct cmh_aes_aead_info *info; +}; + +static const struct cmh_aes_aead_info * +cmh_aes_aead_get_info(struct crypto_aead *tfm) +{ + struct aead_alg *alg =3D crypto_aead_alg(tfm); + + return container_of(alg, struct cmh_aes_aead_drv, alg)->info; +} + +/* VCQ Builders -- AEAD-specific */ + +static void vcq_add_aes_aead_init(struct vcq_cmd *slot, u32 core_id, u64 k= ey_ref, + u64 iv_dma, u32 keylen, u32 ivlen, + u32 mode, u32 op, u32 aadlen, u32 iolen, + u32 taglen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, AES_CMD_INIT); + slot->hwc.aes.cmd_init.key =3D key_ref; + slot->hwc.aes.cmd_init.iv =3D iv_dma; + slot->hwc.aes.cmd_init.keylen =3D keylen; + slot->hwc.aes.cmd_init.ivlen =3D ivlen; + slot->hwc.aes.cmd_init.mode =3D mode; + slot->hwc.aes.cmd_init.op =3D op; + slot->hwc.aes.cmd_init.aadlen =3D aadlen; + slot->hwc.aes.cmd_init.iolen =3D iolen; + slot->hwc.aes.cmd_init.taglen =3D taglen; +} + +static void vcq_add_aes_ccm_init(struct vcq_cmd *slot, u32 core_id, u64 ke= y_ref, + u64 nonce_dma, u32 keylen, u32 noncelen, + u32 op, u32 aadlen, u32 iolen, u32 taglen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, AES_CMD_CCM_INIT); + slot->hwc.aes.cmd_init.key =3D key_ref; + slot->hwc.aes.cmd_init.iv =3D nonce_dma; + slot->hwc.aes.cmd_init.keylen =3D keylen; + slot->hwc.aes.cmd_init.ivlen =3D noncelen; + slot->hwc.aes.cmd_init.mode =3D AES_MODE_CCM; + slot->hwc.aes.cmd_init.op =3D op; + slot->hwc.aes.cmd_init.aadlen =3D aadlen; + slot->hwc.aes.cmd_init.iolen =3D iolen; + slot->hwc.aes.cmd_init.taglen =3D taglen; +} + +static void vcq_add_aes_aad_final(struct vcq_cmd *slot, u32 core_id, u64 a= ad_dma, + u32 aadlen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, AES_CMD_AAD_FINAL); + slot->hwc.aes.cmd_aad_final.data =3D aad_dma; + slot->hwc.aes.cmd_aad_final.datalen =3D aadlen; +} + +static void vcq_add_aes_aead_final(struct vcq_cmd *slot, u32 core_id, u64 = input_dma, + u64 output_dma, u64 tag_dma, + u32 iolen, u32 taglen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, AES_CMD_FINAL); + slot->hwc.aes.cmd_final.input =3D input_dma; + slot->hwc.aes.cmd_final.output =3D output_dma; + slot->hwc.aes.cmd_final.tag =3D tag_dma; + slot->hwc.aes.cmd_final.iolen =3D iolen; + slot->hwc.aes.cmd_final.taglen =3D taglen; +} + +/* setkey */ +static int cmh_aes_aead_setkey(struct crypto_aead *tfm, const u8 *key, + unsigned int keylen) +{ + struct cmh_aes_aead_tfm_ctx *tctx =3D crypto_aead_ctx(tfm); + int ret; + + if (keylen !=3D 16 && keylen !=3D 24 && keylen !=3D 32) + return -EINVAL; + + /* + * Program the HW key first; only mirror it to the SW fallback + * ciphers on success so a failed HW step cannot leave the SW + * ciphers (new) inconsistent with the HW key (old). + */ + ret =3D cmh_key_setkey_raw(&tctx->key, key, keylen, CORE_ID_AES); + if (ret) + return ret; + + if (tctx->sw_cipher) { + ret =3D crypto_cipher_setkey(tctx->sw_cipher, key, keylen); + if (ret) + return ret; + } + if (tctx->fallback) { + ret =3D crypto_aead_setkey(tctx->fallback, key, keylen); + if (ret) + return ret; + } + + return 0; +} + +static int cmh_aes_aead_setauthsize(struct crypto_aead *tfm, + unsigned int authsize) +{ + struct cmh_aes_aead_tfm_ctx *tctx =3D crypto_aead_ctx(tfm); + const struct cmh_aes_aead_info *info =3D cmh_aes_aead_get_info(tfm); + int ret; + + if (info->type =3D=3D CMH_AES_AEAD_GCM) { + /* GCM: accept 4, 8, 12, 13, 14, 15, 16 per NIST SP 800-38D */ + if (authsize < 4 || authsize > 16 || + (authsize > 4 && authsize < 8) || + (authsize > 8 && authsize < 12)) + return -EINVAL; + } else { + /* CCM: accept all RFC 3610 values {4,6,8,10,12,14,16} */ + if (authsize < 4 || authsize > 16 || (authsize & 1)) + return -EINVAL; + /* Forward to SW fallback for authsize=3D10 (HW unsupported) */ + if (tctx->fallback) { + ret =3D crypto_aead_setauthsize(tctx->fallback, + authsize); + if (ret) + return ret; + } + } + + tctx->authsize =3D authsize; + return 0; +} + +static int cmh_aes_aead_init_tfm(struct crypto_aead *tfm) +{ + struct cmh_aes_aead_tfm_ctx *tctx =3D crypto_aead_ctx(tfm); + const struct cmh_aes_aead_info *info =3D cmh_aes_aead_get_info(tfm); + + memset(tctx, 0, sizeof(*tctx)); + tctx->authsize =3D info->maxauthsize; + + if (info->type =3D=3D CMH_AES_AEAD_CCM) { + struct crypto_aead *fb; + struct crypto_cipher *ci; + + ci =3D crypto_alloc_cipher("aes", 0, 0); + if (IS_ERR(ci)) + return PTR_ERR(ci); + tctx->sw_cipher =3D ci; + + fb =3D crypto_alloc_aead("ccm(aes)", 0, + CRYPTO_ALG_NEED_FALLBACK); + if (IS_ERR(fb)) { + crypto_free_cipher(ci); + tctx->sw_cipher =3D NULL; + return PTR_ERR(fb); + } + tctx->fallback =3D fb; + + /* + * The fallback subrequest is placed after cmh_aes_aead_reqctx + * and PTR_ALIGN()'d to crypto_tfm_ctx_alignment() in + * cmh_aes_ccm_fallback(). Reserve that alignment as slack so + * the aligned subrequest still fits within the request context + * even when ARCH_KMALLOC_MINALIGN exceeds the struct's natural + * alignment. + */ + crypto_aead_set_reqsize(tfm, + sizeof(struct cmh_aes_aead_reqctx) + + crypto_tfm_ctx_alignment() + + sizeof(struct aead_request) + + crypto_aead_reqsize(fb)); + } else { + crypto_aead_set_reqsize(tfm, + sizeof(struct cmh_aes_aead_reqctx)); + } + + return 0; +} + +static void cmh_aes_aead_exit_tfm(struct crypto_aead *tfm) +{ + struct cmh_aes_aead_tfm_ctx *tctx =3D crypto_aead_ctx(tfm); + + if (tctx->fallback) + crypto_free_aead(tctx->fallback); + if (tctx->sw_cipher) + crypto_free_cipher(tctx->sw_cipher); + cmh_key_destroy(&tctx->key); +} + +/* DMA unmap helper */ +static void cmh_aes_aead_unmap_dma(struct cmh_aes_aead_reqctx *rctx) +{ + u32 tag_map_len; + + cmh_dma_unmap_single(rctx->iv_dma, rctx->iv_map_len, DMA_TO_DEVICE); + /* + * The empty-GCM fallback maps a full AES block (16 bytes) for the + * ECB output regardless of authsize, so unmap with the mapped size. + */ + tag_map_len =3D rctx->empty_gcm_fallback ? + AES_GCM_HW_IV_SIZE : rctx->authsize; + cmh_dma_unmap_single(rctx->tag_dma, tag_map_len, + (rctx->encrypting || rctx->empty_gcm_fallback) ? + DMA_FROM_DEVICE : DMA_TO_DEVICE); + if (rctx->cryptlen > 0) { + cmh_dma_unmap_single(rctx->out_dma, rctx->cryptlen, + DMA_FROM_DEVICE); + cmh_dma_unmap_single(rctx->in_dma, rctx->cryptlen, + DMA_TO_DEVICE); + } + if (rctx->assoclen > 0) + cmh_dma_unmap_single(rctx->aad_dma, rctx->assoclen, + DMA_TO_DEVICE); +} + +static void cmh_aes_aead_free_bufs(struct cmh_aes_aead_reqctx *rctx) +{ + kfree(rctx->iv_buf); + rctx->iv_buf =3D NULL; + kfree(rctx->tag_buf); + rctx->tag_buf =3D NULL; + kfree_sensitive(rctx->out_buf); + rctx->out_buf =3D NULL; + kfree_sensitive(rctx->in_buf); + rctx->in_buf =3D NULL; + kfree(rctx->aad_buf); + rctx->aad_buf =3D NULL; +} + +static void cmh_aes_aead_complete(void *data, int error) +{ + struct aead_request *req =3D data; + struct cmh_aes_aead_reqctx *rctx =3D aead_request_ctx(req); + + if (error =3D=3D -EINPROGRESS) { + cmh_complete(&req->base, error); + return; + } + + cmh_aes_aead_unmap_dma(rctx); + + /* + * Map HW error on decrypt to -EBADMSG. The eSW AES core uses a + * single error code (-EIO) for both authentication failures and + * other core errors (e.g. DMA timeout), so we cannot distinguish + * them from the MBX_STATUS alone. In practice the only error + * during a well-formed AEAD decrypt is auth-tag mismatch; a DMA + * timeout would indicate a fatal HW problem where -EBADMSG vs + * -EIO is moot. The kernel crypto API requires -EBADMSG for + * AEAD authentication failures. + */ + if (error =3D=3D -EIO && !rctx->encrypting) + error =3D -EBADMSG; + + if (!error) { + /* GCM empty-input decrypt: compare computed tag with expected */ + if (rctx->empty_gcm_fallback && !rctx->encrypting) { + if (crypto_memneq(rctx->tag_buf, rctx->in_buf, + rctx->authsize)) + error =3D -EBADMSG; + } + if (!error && rctx->cryptlen > 0) + scatterwalk_map_and_copy(rctx->out_buf, req->dst, + req->assoclen, + rctx->cryptlen, 1); + /* + * Out-of-place (src !=3D dst): the AAD is authenticated but not + * transformed, so the crypto API expects it copied verbatim + * into the dst AAD region (in-place dst already holds it). + */ + if (!error && req->src !=3D req->dst && req->assoclen > 0) + scatterwalk_map_and_copy(rctx->aad_buf, req->dst, + 0, req->assoclen, 1); + if (!error && rctx->encrypting) + scatterwalk_map_and_copy(rctx->tag_buf, req->dst, + req->assoclen + + rctx->cryptlen, + rctx->authsize, 1); + } + + cmh_aes_aead_free_bufs(rctx); + cmh_complete(&req->base, error); +} + +/* + * GCM empty-input fallback. + * + * When both AAD and plaintext are empty, GCM reduces to: + * tag =3D E(K, J0) where J0 =3D nonce || 0x00000001 + * + * The eSW GCM engine rejects this degenerate case, so we compute it + * via a single ECB block encryption of J0. + * + * VCQ: [SYS_CMD_WRITE] + AES_CMD_INIT(ECB) + AES_CMD_FINAL + FLUSH + */ +static int cmh_aes_gcm_empty(struct aead_request *req, u32 aes_op) +{ + struct crypto_aead *tfm =3D crypto_aead_reqtfm(req); + struct cmh_aes_aead_tfm_ctx *tctx =3D crypto_aead_ctx(tfm); + struct cmh_aes_aead_reqctx *rctx =3D aead_request_ctx(req); + struct vcq_cmd cmds[CMH_AES_AEAD_MAX_PAYLOAD]; + u64 key_ref; + u32 keylen, authsize; + struct core_dispatch d; + s32 target_mbx; + u32 core_id; + u32 idx; + int ret; + gfp_t gfp; + + authsize =3D tctx->authsize; + + gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + + memset(rctx, 0, sizeof(*rctx)); + rctx->cryptlen =3D 0; + rctx->assoclen =3D 0; + rctx->authsize =3D authsize; + rctx->encrypting =3D (aes_op =3D=3D AES_OP_ENCRYPT); + rctx->empty_gcm_fallback =3D true; + + /* Build J0 =3D nonce || 0x00000001 in iv_buf */ + rctx->iv_buf =3D kzalloc(AES_GCM_HW_IV_SIZE, gfp); + if (!rctx->iv_buf) + return -ENOMEM; + memcpy(rctx->iv_buf, req->iv, AES_GCM_IV_SIZE); + rctx->iv_buf[15] =3D 0x01; /* big-endian counter =3D 1 */ + rctx->iv_map_len =3D AES_GCM_HW_IV_SIZE; + + rctx->iv_dma =3D cmh_dma_map_single(rctx->iv_buf, AES_GCM_HW_IV_SIZE, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->iv_dma)) { + ret =3D -ENOMEM; + goto out_free_iv; + } + + /* Tag buffer -- receives E(K, J0) output */ + rctx->tag_buf =3D kzalloc(AES_GCM_HW_IV_SIZE, gfp); + if (!rctx->tag_buf) { + ret =3D -ENOMEM; + goto out_unmap_iv; + } + rctx->tag_dma =3D cmh_dma_map_single(rctx->tag_buf, AES_GCM_HW_IV_SIZE, + DMA_FROM_DEVICE); + if (cmh_dma_map_error(rctx->tag_dma)) { + ret =3D -ENOMEM; + goto out_free_tag; + } + + /* For decrypt: read expected tag from request for later comparison */ + if (!rctx->encrypting) { + rctx->in_buf =3D kmalloc(authsize, gfp); + if (!rctx->in_buf) { + ret =3D -ENOMEM; + goto out_unmap_tag; + } + scatterwalk_map_and_copy(rctx->in_buf, req->src, 0, + authsize, 0); + } + + /* Resolve key */ + idx =3D 0; + rctx->key_dma =3D tctx->key.raw.dma; + vcq_add_sys_write(&cmds[idx++], SYS_REF_TEMP, + (u64)rctx->key_dma, SYS_REF_NONE, + tctx->key.raw.len, + tctx->key.raw.sys_type); + key_ref =3D SYS_REF_TEMP; + keylen =3D tctx->key.raw.len; + d =3D cmh_core_select_instance(CMH_CORE_AES); + target_mbx =3D d.mbx_idx; + core_id =3D d.core_id; + + /* ECB INIT: single block encryption of J0 */ + vcq_add_aes_aead_init(&cmds[idx++], core_id, key_ref, + 0, keylen, 0, AES_MODE_ECB, AES_OP_ENCRYPT, + 0, AES_GCM_HW_IV_SIZE, 0); + + /* FINAL: J0 in, E(K,J0) out */ + vcq_add_aes_aead_final(&cmds[idx++], core_id, + (u64)rctx->iv_dma, (u64)rctx->tag_dma, + 0, AES_GCM_HW_IV_SIZE, 0); + + vcq_add_flush(&cmds[idx++], core_id); + + ret =3D cmh_vcq_pack_and_submit_async(cmds, idx, rctx->packed, + CMH_AES_AEAD_MAX_PACKED, + target_mbx, + cmh_aes_aead_complete, req, + !!(req->base.flags & + CRYPTO_TFM_REQ_MAY_BACKLOG), + cmh_tm_async_timeout_jiffies()); + if (ret =3D=3D -EBUSY) + return -EBUSY; + if (ret) + goto out_free_in; + + return -EINPROGRESS; + +out_free_in: + kfree_sensitive(rctx->in_buf); +out_unmap_tag: + cmh_dma_unmap_single(rctx->tag_dma, AES_GCM_HW_IV_SIZE, + DMA_FROM_DEVICE); +out_free_tag: + kfree(rctx->tag_buf); +out_unmap_iv: + cmh_dma_unmap_single(rctx->iv_dma, AES_GCM_HW_IV_SIZE, DMA_TO_DEVICE); +out_free_iv: + kfree(rctx->iv_buf); + return ret; +} + +/* + * CCM empty-input fallback. + * + * When both AAD and plaintext are empty, CCM reduces to: + * T =3D E(K, B0) -- CBC-MAC of the single formatting block + * S0 =3D E(K, A0) -- CTR block zero + * tag =3D (T XOR S0)[0..authsize-1] + * + * The eSW rejects this degenerate case, so the driver computes it + * synchronously via two crypto_cipher single-block encryptions. + */ +static int cmh_aes_ccm_empty(struct aead_request *req, u32 aes_op) +{ + struct crypto_aead *tfm =3D crypto_aead_reqtfm(req); + struct cmh_aes_aead_tfm_ctx *tctx =3D crypto_aead_ctx(tfm); + u32 authsize =3D tctx->authsize; + u8 b0[CMH_AES_BLOCK_SIZE], a0[CMH_AES_BLOCK_SIZE]; + u8 t[CMH_AES_BLOCK_SIZE], s0[CMH_AES_BLOCK_SIZE]; + u8 tag[CMH_AES_BLOCK_SIZE]; + u8 L; + u32 i; + + /* Defense-in-depth: iv[0] =3D L-1, valid L is 2..8 per RFC 3610 S2.1 */ + if (WARN_ON_ONCE(req->iv[0] < 1 || req->iv[0] > 7)) + return -EINVAL; + + L =3D req->iv[0] + 1; + + if (tctx->key.mode !=3D CMH_KEY_RAW) + return -EOPNOTSUPP; + + /* B0: flags || nonce || Q(=3D0). Adata=3D0, t=3Dauthsize, q=3DL. */ + memset(b0, 0, CMH_AES_BLOCK_SIZE); + b0[0] =3D (u8)(8 * ((authsize - 2) / 2) + (L - 1)); + memcpy(&b0[1], &req->iv[1], 15 - L); + + /* A0: (L-1) || nonce || counter(=3D0) */ + memset(a0, 0, CMH_AES_BLOCK_SIZE); + a0[0] =3D (u8)(L - 1); + memcpy(&a0[1], &req->iv[1], 15 - L); + + crypto_cipher_encrypt_one(tctx->sw_cipher, t, b0); + crypto_cipher_encrypt_one(tctx->sw_cipher, s0, a0); + + for (i =3D 0; i < authsize; i++) + tag[i] =3D t[i] ^ s0[i]; + + if (aes_op =3D=3D AES_OP_ENCRYPT) { + scatterwalk_map_and_copy(tag, req->dst, + req->assoclen, authsize, 1); + } else { + u8 expected[CMH_AES_BLOCK_SIZE]; + + scatterwalk_map_and_copy(expected, req->src, + req->assoclen, authsize, 0); + if (crypto_memneq(tag, expected, authsize)) + return -EBADMSG; + } + + return 0; +} + +/* + * CCM authsize=3D10 fallback. + * + * The eSW AES CCM core does not support authsize=3D10 (valid per RFC 3610= ). + * Forward the entire request to the generic CCM implementation. + */ +static void cmh_aes_ccm_fb_done(void *data, int err) +{ + struct aead_request *req =3D data; + + cmh_complete(&req->base, err); +} + +static int cmh_aes_ccm_fallback(struct aead_request *req, u32 aes_op) +{ + struct crypto_aead *tfm =3D crypto_aead_reqtfm(req); + struct cmh_aes_aead_tfm_ctx *tctx =3D crypto_aead_ctx(tfm); + struct cmh_aes_aead_reqctx *rctx =3D aead_request_ctx(req); + struct aead_request *subreq =3D + PTR_ALIGN((void *)(rctx + 1), crypto_tfm_ctx_alignment()); + + aead_request_set_tfm(subreq, tctx->fallback); + aead_request_set_callback(subreq, req->base.flags, + cmh_aes_ccm_fb_done, req); + aead_request_set_crypt(subreq, req->src, req->dst, + req->cryptlen, req->iv); + aead_request_set_ad(subreq, req->assoclen); + + return (aes_op =3D=3D AES_OP_ENCRYPT) ? + crypto_aead_encrypt(subreq) : crypto_aead_decrypt(subreq); +} + +/* + * Core AEAD encrypt/decrypt -- async path. + * + * Encrypt: plaintext -> ciphertext + tag appended + * Decrypt: ciphertext + tag -> plaintext (tag verified by eSW) + * + * VCQ: [SYS_CMD_WRITE] + INIT/CCM_INIT + [AAD_FINAL] + FINAL + FLUSH + */ +static int cmh_aes_aead_crypt(struct aead_request *req, u32 aes_op) +{ + struct crypto_aead *tfm =3D crypto_aead_reqtfm(req); + struct cmh_aes_aead_tfm_ctx *tctx =3D crypto_aead_ctx(tfm); + const struct cmh_aes_aead_info *info =3D cmh_aes_aead_get_info(tfm); + struct cmh_aes_aead_reqctx *rctx =3D aead_request_ctx(req); + struct vcq_cmd cmds[CMH_AES_AEAD_MAX_PAYLOAD]; + u64 key_ref; + u32 keylen, authsize, cryptlen; + struct core_dispatch d; + s32 target_mbx; + u32 core_id; + u32 idx; + int ret; + gfp_t gfp; + + if (tctx->key.mode =3D=3D CMH_KEY_NONE) + return -ENOKEY; + + authsize =3D tctx->authsize; + + if (aes_op =3D=3D AES_OP_ENCRYPT) { + cryptlen =3D req->cryptlen; + } else { + if (req->cryptlen < authsize) + return -EINVAL; + cryptlen =3D req->cryptlen - authsize; + } + + /* + * Validate CCM IV format early -- the empty-input fallback and + * nonce extraction both depend on iv[0] being in range [1,7]. + */ + if (info->type =3D=3D CMH_AES_AEAD_CCM) { + if (req->iv[0] < 1 || req->iv[0] > 7) + return -EINVAL; + } + + /* + * The CMH eSW rejects GCM/CCM when both aadlen and iolen are zero. + * For GCM, the tag is simply E(K, J0) -- handle with ECB fallback. + * For CCM, compute tag =3D E(K,B0) XOR E(K,A0) in software. + */ + if (cryptlen =3D=3D 0 && req->assoclen =3D=3D 0) { + if (info->type =3D=3D CMH_AES_AEAD_GCM) + return cmh_aes_gcm_empty(req, aes_op); + return cmh_aes_ccm_empty(req, aes_op); + } + + /* + * HW does not support authsize=3D10 for CCM. Forward the entire + * request to the generic CCM implementation. + */ + if (info->type =3D=3D CMH_AES_AEAD_CCM && authsize =3D=3D 10) + return cmh_aes_ccm_fallback(req, aes_op); + + /* + * HW uses a proprietary LLI scatter-gather format that is + * incompatible with struct scatterlist, so the payload is + * linearised into contiguous buffers for DMA. Cap total + * size to prevent excessive memory consumption. + */ + if ((u64)cryptlen + req->assoclen > SZ_1M) + return -EINVAL; + + gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + + memset(rctx, 0, sizeof(*rctx)); + rctx->cryptlen =3D cryptlen; + rctx->assoclen =3D req->assoclen; + rctx->authsize =3D authsize; + rctx->encrypting =3D (aes_op =3D=3D AES_OP_ENCRYPT); + + /* Linearise AAD */ + if (req->assoclen > 0) { + rctx->aad_buf =3D kmalloc(req->assoclen, gfp | __GFP_NOWARN); + if (!rctx->aad_buf) + return -ENOMEM; + scatterwalk_map_and_copy(rctx->aad_buf, req->src, + 0, req->assoclen, 0); + rctx->aad_dma =3D cmh_dma_map_single(rctx->aad_buf, + req->assoclen, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->aad_dma)) { + ret =3D -ENOMEM; + goto out_free_aad; + } + } + + /* Linearise input */ + if (cryptlen > 0) { + rctx->in_buf =3D kmalloc(cryptlen, gfp | __GFP_NOWARN); + if (!rctx->in_buf) { + ret =3D -ENOMEM; + goto out_unmap_aad; + } + scatterwalk_map_and_copy(rctx->in_buf, req->src, + req->assoclen, cryptlen, 0); + rctx->in_dma =3D cmh_dma_map_single(rctx->in_buf, cryptlen, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->in_dma)) { + ret =3D -ENOMEM; + goto out_free_in; + } + } + + /* Allocate output buffer */ + if (cryptlen > 0) { + rctx->out_buf =3D kmalloc(cryptlen, gfp | __GFP_NOWARN); + if (!rctx->out_buf) { + ret =3D -ENOMEM; + goto out_unmap_in; + } + rctx->out_dma =3D cmh_dma_map_single(rctx->out_buf, cryptlen, + DMA_FROM_DEVICE); + if (cmh_dma_map_error(rctx->out_dma)) { + ret =3D -ENOMEM; + goto out_free_out; + } + } + + /* Tag buffer */ + rctx->tag_buf =3D kmalloc(authsize, gfp); + if (!rctx->tag_buf) { + ret =3D -ENOMEM; + goto out_unmap_out; + } + + if (!rctx->encrypting) { + scatterwalk_map_and_copy(rctx->tag_buf, req->src, + req->assoclen + cryptlen, + authsize, 0); + } else { + memset(rctx->tag_buf, 0, authsize); + } + + rctx->tag_dma =3D cmh_dma_map_single(rctx->tag_buf, authsize, + rctx->encrypting ? + DMA_FROM_DEVICE : DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->tag_dma)) { + ret =3D -ENOMEM; + goto out_free_tag; + } + + /* Map IV/nonce */ + if (info->type =3D=3D CMH_AES_AEAD_GCM) { + rctx->iv_buf =3D kzalloc(AES_GCM_HW_IV_SIZE, gfp); + if (!rctx->iv_buf) { + ret =3D -ENOMEM; + goto out_unmap_tag; + } + memcpy(rctx->iv_buf, req->iv, AES_GCM_IV_SIZE); + rctx->iv_map_len =3D AES_GCM_HW_IV_SIZE; + rctx->iv_dma =3D cmh_dma_map_single(rctx->iv_buf, + rctx->iv_map_len, + DMA_TO_DEVICE); + } else { + u32 noncelen; + + if (req->iv[0] < 1 || req->iv[0] > 7) { + ret =3D -EINVAL; + goto out_unmap_tag; + } + noncelen =3D 14 - req->iv[0]; + + rctx->iv_buf =3D kmemdup(req->iv + 1, noncelen, gfp); + if (!rctx->iv_buf) { + ret =3D -ENOMEM; + goto out_unmap_tag; + } + rctx->iv_map_len =3D noncelen; + rctx->iv_dma =3D cmh_dma_map_single(rctx->iv_buf, + rctx->iv_map_len, + DMA_TO_DEVICE); + } + if (cmh_dma_map_error(rctx->iv_dma)) { + ret =3D -ENOMEM; + goto out_free_iv; + } + + /* Resolve key reference */ + idx =3D 0; + + rctx->key_dma =3D tctx->key.raw.dma; + rctx->keylen =3D tctx->key.raw.len; + vcq_add_sys_write(&cmds[idx++], SYS_REF_TEMP, + (u64)rctx->key_dma, SYS_REF_NONE, + tctx->key.raw.len, + tctx->key.raw.sys_type); + key_ref =3D SYS_REF_TEMP; + keylen =3D tctx->key.raw.len; + d =3D cmh_core_select_instance(CMH_CORE_AES); + target_mbx =3D d.mbx_idx; + core_id =3D d.core_id; + + /* Build INIT command */ + if (info->type =3D=3D CMH_AES_AEAD_CCM) { + vcq_add_aes_ccm_init(&cmds[idx++], core_id, key_ref, + (u64)rctx->iv_dma, keylen, + rctx->iv_map_len, aes_op, + req->assoclen, cryptlen, authsize); + } else { + vcq_add_aes_aead_init(&cmds[idx++], core_id, key_ref, + (u64)rctx->iv_dma, keylen, + AES_GCM_HW_IV_SIZE, info->aes_mode, + aes_op, req->assoclen, cryptlen, + authsize); + } + + if (req->assoclen > 0) + vcq_add_aes_aad_final(&cmds[idx++], core_id, + (u64)rctx->aad_dma, req->assoclen); + + vcq_add_aes_aead_final(&cmds[idx++], core_id, + cryptlen > 0 ? (u64)rctx->in_dma : 0, + cryptlen > 0 ? (u64)rctx->out_dma : 0, + (u64)rctx->tag_dma, cryptlen, authsize); + + vcq_add_flush(&cmds[idx++], core_id); + + ret =3D cmh_vcq_pack_and_submit_async(cmds, idx, rctx->packed, + CMH_AES_AEAD_MAX_PACKED, + target_mbx, + cmh_aes_aead_complete, req, + !!(req->base.flags & + CRYPTO_TFM_REQ_MAY_BACKLOG), + cmh_tm_async_timeout_jiffies()); + if (ret =3D=3D -EBUSY) + return -EBUSY; + if (ret) + goto out_cleanup_all; + + return -EINPROGRESS; + +out_cleanup_all: + cmh_dma_unmap_single(rctx->iv_dma, rctx->iv_map_len, DMA_TO_DEVICE); +out_free_iv: + kfree(rctx->iv_buf); +out_unmap_tag: + cmh_dma_unmap_single(rctx->tag_dma, authsize, + rctx->encrypting ? DMA_FROM_DEVICE : + DMA_TO_DEVICE); +out_free_tag: + kfree(rctx->tag_buf); +out_unmap_out: + if (cryptlen > 0) + cmh_dma_unmap_single(rctx->out_dma, cryptlen, DMA_FROM_DEVICE); +out_free_out: + kfree_sensitive(rctx->out_buf); +out_unmap_in: + if (cryptlen > 0) + cmh_dma_unmap_single(rctx->in_dma, cryptlen, DMA_TO_DEVICE); +out_free_in: + kfree_sensitive(rctx->in_buf); +out_unmap_aad: + if (req->assoclen > 0) + cmh_dma_unmap_single(rctx->aad_dma, req->assoclen, + DMA_TO_DEVICE); +out_free_aad: + kfree(rctx->aad_buf); + return ret; +} + +static int cmh_aes_aead_encrypt(struct aead_request *req) +{ + return cmh_aes_aead_crypt(req, AES_OP_ENCRYPT); +} + +static int cmh_aes_aead_decrypt(struct aead_request *req) +{ + return cmh_aes_aead_crypt(req, AES_OP_DECRYPT); +} + +/* Registration */ + +static struct cmh_aes_aead_drv aes_aead_drv_algs[ARRAY_SIZE(aes_aead_algs)= ]; + +/** + * cmh_aes_aead_register() - Register AES-GCM/CCM AEAD algorithms with the= crypto framework + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_aes_aead_register(void) +{ + unsigned int i; + int ret; + + if (!cmh_core_present(CMH_CORE_AES)) + return 0; + + for (i =3D 0; i < ARRAY_SIZE(aes_aead_algs); i++) { + const struct cmh_aes_aead_info *info =3D &aes_aead_algs[i]; + struct cmh_aes_aead_drv *drv =3D &aes_aead_drv_algs[i]; + struct aead_alg *alg =3D &drv->alg; + + drv->info =3D info; + + memset(alg, 0, sizeof(*alg)); + + alg->setkey =3D cmh_aes_aead_setkey; + alg->setauthsize =3D cmh_aes_aead_setauthsize; + alg->encrypt =3D cmh_aes_aead_encrypt; + alg->decrypt =3D cmh_aes_aead_decrypt; + alg->init =3D cmh_aes_aead_init_tfm; + alg->exit =3D cmh_aes_aead_exit_tfm; + alg->ivsize =3D info->ivsize; + alg->maxauthsize =3D info->maxauthsize; + + strscpy(alg->base.cra_name, info->alg_name, + CRYPTO_MAX_ALG_NAME); + strscpy(alg->base.cra_driver_name, info->drv_name, + CRYPTO_MAX_ALG_NAME); + alg->base.cra_priority =3D 300; + alg->base.cra_flags =3D CRYPTO_ALG_KERN_DRIVER_ONLY | + CRYPTO_ALG_ASYNC; + if (info->type =3D=3D CMH_AES_AEAD_CCM) { + alg->base.cra_flags |=3D CRYPTO_ALG_NEED_FALLBACK; + /* + * Bump priority above 300 so we beat the generic + * ccm_base template instance. That template inherits + * priority (ctr + cbcmac) / 2 =3D 300 when both + * constituents are at 300, and list ordering would + * otherwise let it shadow our driver. + */ + alg->base.cra_priority =3D 301; + } + alg->base.cra_blocksize =3D 1; + alg->base.cra_ctxsize =3D sizeof(struct cmh_aes_aead_tfm_ctx); + alg->base.cra_module =3D THIS_MODULE; + + ret =3D crypto_register_aead(alg); + if (ret) { + dev_err(cmh_dev(), "cmh_aes_aead: failed to register %s (rc=3D%d)\n", + info->alg_name, ret); + goto err_unregister; + } + + dev_dbg(cmh_dev(), "cmh_aes_aead: registered %s\n", info->alg_name); + } + + return 0; + +err_unregister: + while (i--) + crypto_unregister_aead(&aes_aead_drv_algs[i].alg); + return ret; +} + +/** + * cmh_aes_aead_unregister() - Unregister AES AEAD algorithms from the cry= pto framework + */ +void cmh_aes_aead_unregister(void) +{ + unsigned int i; + + if (!cmh_core_present(CMH_CORE_AES)) + return; + + for (i =3D 0; i < ARRAY_SIZE(aes_aead_algs); i++) { + crypto_unregister_aead(&aes_aead_drv_algs[i].alg); + dev_dbg(cmh_dev(), "cmh_aes_aead: unregistered %s\n", + aes_aead_algs[i].alg_name); + } +} diff --git a/drivers/crypto/cmh/cmh_aes_cmac.c b/drivers/crypto/cmh/cmh_aes= _cmac.c new file mode 100644 index 000000000000..c19af5400fc0 --- /dev/null +++ b/drivers/crypto/cmh/cmh_aes_cmac.c @@ -0,0 +1,741 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Kernel Crypto API AES-CMAC (ahash) Driver + * + * Registers cmac(aes) as an ahash algorithm. + * + * CMAC produces a 16-byte tag (MAC) from a key and message. + * VCQ sequence: [SYS_CMD_WRITE] + AES_CMD_INIT(CMAC) + + * AES_CMD_AAD_FINAL_AUTH + FLUSH + * + * The ahash interface accumulates data in a kernel buffer via .update(), + * then .final() builds and submits the VCQ asynchronously. + */ + +#include +#include +#include +#include +#include +#include +#include +#include + +#include "cmh_aes.h" +#include "cmh_vcq.h" +#include "cmh_aes_abi.h" +#include "cmh_sys_abi.h" +#include "cmh_sys.h" +#include "cmh_txn.h" +#include "cmh_dma.h" +#include "cmh_key.h" + +#define AES_CMAC_DIGEST_SIZE 16U +#define AES_CMAC_BLOCK_SIZE 16U + +/* + * Maximum accumulated data for CMAC -- driver-imposed, not HW. + * + * The AES core does not expose external save/restore VCQ commands, + * so the driver must accumulate all data in kernel memory via + * .update() and submit it atomically in .final(). This cap limits + * the per-request kernel allocation. + */ +#define AES_CMAC_MAX_DATA (64 * 1024) + +/* Per-transform context */ +struct cmh_aes_cmac_tfm_ctx { + struct cmh_key_ctx key; + struct crypto_ahash *fb; /* generic SW fallback (oversized ops) */ + spinlock_t chunk_lock; /* protects all_chunks + tfm_buffered */ + struct list_head all_chunks; /* orphan-safe chunk tracking */ + size_t tfm_buffered; /* bytes on all_chunks; DoS cap */ +}; + +/* + * Per-transform cap on total bytes buffered across all_chunks. Bounds + * memory an AF_ALG client can pin via repeated open/update/abandon of + * request sockets (the crypto API has no per-request destructor). + */ +#define CMH_AES_CMAC_TFM_MAX_BUFFERED (16 * 1024 * 1024) + +/* One chunk per .update() call -- data is embedded via flexible array */ +struct cmh_aes_cmac_chunk { + struct list_head list; + struct list_head tfm_node; /* per-tfm orphan tracking */ + u32 len; + u8 data[]; +}; + +/* Per-request context (lives in ahash_request::__ctx) */ + +/* + * Maximum payload commands: + * [SYS_CMD_WRITE] + AES_CMD_INIT + AES_CMD_AAD_FINAL_AUTH + FLUSH =3D 4 + */ +#define CMH_AES_CMAC_MAX_PAYLOAD 4 +#define CMH_AES_CMAC_MAX_PACKED (CMH_AES_CMAC_MAX_PAYLOAD * 2) + +struct cmh_aes_cmac_reqctx { + struct list_head chunks; + u32 total_len; + bool switched; /* handed off to SW fallback */ + u8 *buf; /* linearised in final() for DMA */ + /* DMA state for async final */ + dma_addr_t key_dma; + dma_addr_t in_dma; + dma_addr_t tag_dma; + u8 *tag_buf; + u32 keylen; + struct vcq_cmd packed[CMH_AES_CMAC_MAX_PACKED]; +}; + +/* + * Flat state for export/import, tagged by the leading @format byte: + * CMH_AES_CMAC_FMT_RAW is the accumulated input bytes (flat window); + * CMH_AES_CMAC_FMT_FB is the software fallback's own exported state, + * used once the request has switched to the fallback so export/import + * (transform clone) works at any input length. + */ +#define CMH_AES_CMAC_FMT_RAW 0 +#define CMH_AES_CMAC_FMT_FB 1 + +struct cmh_aes_cmac_export_state { + u8 format; + u8 __pad[3]; + u32 total_len; + u8 data[]; +}; + +/* + * The crypto subsystem pre-allocates statesize bytes per request. + * CMH_AES_CMAC_STATE_SIZE (4096) sizes both the CMH_AES_CMAC_FMT_RAW + * window (CMH_AES_CMAC_EXPORT_MAX accumulated bytes) and the + * CMH_AES_CMAC_FMT_FB software state. A RAW export past + * CMH_AES_CMAC_EXPORT_MAX transparently switches to the fallback and + * emits CMH_AES_CMAC_FMT_FB instead, so export/import is not capped. + */ +#define CMH_AES_CMAC_STATE_SIZE 4096 +#define CMH_AES_CMAC_EXPORT_MAX \ + (CMH_AES_CMAC_STATE_SIZE - sizeof(struct cmh_aes_cmac_export_state)) + +/* + * Export/import (transform clone): the AES core lacks external + * save/restore VCQ commands, so the driver accumulates input in kernel + * memory and serialises that buffer for the common (bounded) case. + * When the accumulated input exceeds the HW cap (AES_CMAC_MAX_DATA, + * 64 KB) or the flat export window, the request transparently switches + * to a generic software cmac(aes) fallback the driver allocates itself, + * so arbitrary-length MACs and transform clone both stay conformant + * with O(1) driver memory. + */ + +static int cmh_aes_cmac_setkey(struct crypto_ahash *tfm, const u8 *key, + unsigned int keylen) +{ + struct cmh_aes_cmac_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + int ret; + + if (keylen !=3D 16 && keylen !=3D 24 && keylen !=3D 32) + return -EINVAL; + + ret =3D cmh_key_setkey_raw(&tctx->key, key, keylen, CORE_ID_AES); + if (ret) + return ret; + + /* Keep the software fallback keyed in lock-step for oversized ops. */ + return crypto_ahash_setkey(tctx->fb, key, keylen); +} + +static void cmh_aes_cmac_free_chunks(struct cmh_aes_cmac_reqctx *rctx, + struct cmh_aes_cmac_tfm_ctx *tctx) +{ + struct cmh_aes_cmac_chunk *c, *tmp; + + spin_lock_bh(&tctx->chunk_lock); + list_for_each_entry_safe(c, tmp, &rctx->chunks, list) { + list_del(&c->list); + list_del(&c->tfm_node); + tctx->tfm_buffered -=3D c->len; + kfree_sensitive(c); + } + spin_unlock_bh(&tctx->chunk_lock); + rctx->total_len =3D 0; +} + +/* Software-fallback helpers (arbitrary-length + transform-clone support) = */ + +/* + * The fallback ahash_request lives immediately after the reqctx; + * cmh_aes_cmac_init_tfm() reserves crypto_ahash_reqsize(fb) bytes for it. + */ +static struct ahash_request * +cmh_aes_cmac_fb_req(struct cmh_aes_cmac_reqctx *rctx) +{ + return PTR_ALIGN((void *)(rctx + 1), crypto_tfm_ctx_alignment()); +} + +static int cmh_aes_cmac_fb_update_virt(struct ahash_request *fb_req, + const u8 *data, u32 len) +{ + ahash_request_set_virt(fb_req, data, NULL, len); + return crypto_ahash_update(fb_req); +} + +/* + * Switch a request from the HW-buffered path to the software fallback: + * initialise the fallback request, replay every accumulated chunk + * through it, then drop the chunks. The fallback transform was keyed by + * cmh_aes_cmac_setkey() when the caller installed the MAC key. + */ +static int cmh_aes_cmac_switch_to_fb(struct ahash_request *req) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_aes_cmac_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cmh_aes_cmac_reqctx *rctx =3D ahash_request_ctx(req); + struct ahash_request *fb_req =3D cmh_aes_cmac_fb_req(rctx); + struct cmh_aes_cmac_chunk *c; + int ret; + + ahash_request_set_tfm(fb_req, tctx->fb); + ahash_request_set_callback(fb_req, 0, NULL, NULL); + + ret =3D crypto_ahash_init(fb_req); + if (ret) + return ret; + + list_for_each_entry(c, &rctx->chunks, list) { + ret =3D cmh_aes_cmac_fb_update_virt(fb_req, c->data, c->len); + if (ret) + return ret; + } + + cmh_aes_cmac_free_chunks(rctx, tctx); + rctx->switched =3D true; + return 0; +} + +/* Forward the current update() payload to the fallback. */ +static int cmh_aes_cmac_fb_forward(struct ahash_request *req, + struct cmh_aes_cmac_reqctx *rctx) +{ + struct ahash_request *fb_req =3D cmh_aes_cmac_fb_req(rctx); + + if (req->base.flags & CRYPTO_AHASH_REQ_VIRT) + return cmh_aes_cmac_fb_update_virt(fb_req, req->svirt, + req->nbytes); + ahash_request_set_crypt(fb_req, req->src, NULL, req->nbytes); + return crypto_ahash_update(fb_req); +} + +static int cmh_aes_cmac_init(struct ahash_request *req) +{ + struct cmh_aes_cmac_reqctx *rctx =3D ahash_request_ctx(req); + + memset(rctx, 0, sizeof(*rctx)); + INIT_LIST_HEAD(&rctx->chunks); + return 0; +} + +static int cmh_aes_cmac_update(struct ahash_request *req) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_aes_cmac_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cmh_aes_cmac_reqctx *rctx =3D ahash_request_ctx(req); + struct cmh_aes_cmac_chunk *chunk; + gfp_t gfp; + int ret; + + if (!req->nbytes) + return 0; + + /* Already handed off to the fallback: forward directly (O(1) mem). */ + if (rctx->switched) + return cmh_aes_cmac_fb_forward(req, rctx); + + /* + * Exceeding the HW input cap: switch to the software fallback + * (replaying the buffered chunks) rather than failing, then + * forward this update. + */ + if (req->nbytes > AES_CMAC_MAX_DATA - rctx->total_len) { + ret =3D cmh_aes_cmac_switch_to_fb(req); + if (ret) + goto err_free_chunks; + return cmh_aes_cmac_fb_forward(req, rctx); + } + + gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + + chunk =3D kmalloc(sizeof(*chunk) + req->nbytes, gfp); + if (!chunk) { + ret =3D -ENOMEM; + goto err_free_chunks; + } + + chunk->len =3D req->nbytes; + if (req->base.flags & CRYPTO_AHASH_REQ_VIRT) + memcpy(chunk->data, req->svirt, req->nbytes); + else + scatterwalk_map_and_copy(chunk->data, req->src, + 0, req->nbytes, 0); + + spin_lock_bh(&tctx->chunk_lock); + if (tctx->tfm_buffered + chunk->len > CMH_AES_CMAC_TFM_MAX_BUFFERED) { + spin_unlock_bh(&tctx->chunk_lock); + kfree_sensitive(chunk); + ret =3D -ENOMEM; + goto err_free_chunks; + } + list_add_tail(&chunk->list, &rctx->chunks); + list_add_tail(&chunk->tfm_node, &tctx->all_chunks); + tctx->tfm_buffered +=3D chunk->len; + spin_unlock_bh(&tctx->chunk_lock); + rctx->total_len +=3D req->nbytes; + return 0; + +err_free_chunks: + /* + * Terminal error -- free all previously accumulated chunks. + * callers may not call .final() on error, so they would leak. + */ + cmh_aes_cmac_free_chunks(rctx, tctx); + return ret; +} + +static void cmh_aes_cmac_complete(void *data, int error) +{ + struct ahash_request *req =3D data; + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_aes_cmac_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cmh_aes_cmac_reqctx *rctx =3D ahash_request_ctx(req); + + if (error =3D=3D -EINPROGRESS) { + cmh_complete(&req->base, error); + return; + } + + /* Unmap DMA */ + if (rctx->total_len > 0) + cmh_dma_unmap_single(rctx->in_dma, rctx->total_len, + DMA_TO_DEVICE); + cmh_dma_unmap_single(rctx->tag_dma, AES_CMAC_DIGEST_SIZE, + DMA_FROM_DEVICE); + + if (!error) + memcpy(req->result, rctx->tag_buf, AES_CMAC_DIGEST_SIZE); + + kfree(rctx->tag_buf); + rctx->tag_buf =3D NULL; + kfree_sensitive(rctx->buf); + rctx->buf =3D NULL; + cmh_aes_cmac_free_chunks(rctx, tctx); + cmh_complete(&req->base, error); +} + +static int cmh_aes_cmac_final(struct ahash_request *req) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_aes_cmac_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cmh_aes_cmac_reqctx *rctx =3D ahash_request_ctx(req); + struct vcq_cmd cmds[CMH_AES_CMAC_MAX_PAYLOAD]; + u64 key_ref; + u32 keylen; + struct core_dispatch d; + s32 target_mbx; + u32 core_id; + u32 idx; + int ret; + gfp_t gfp; + + /* Switched to the software fallback: complete there (synchronous). */ + if (rctx->switched) { + struct ahash_request *fb_req =3D cmh_aes_cmac_fb_req(rctx); + + ahash_request_set_crypt(fb_req, NULL, req->result, 0); + return crypto_ahash_final(fb_req); + } + + if (tctx->key.mode =3D=3D CMH_KEY_NONE) { + ret =3D -ENOKEY; + goto out_free_buf; + } + + gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + + /* Linearise accumulated chunks into a contiguous buffer for DMA */ + if (rctx->total_len > 0) { + struct cmh_aes_cmac_chunk *c; + u32 off =3D 0; + + rctx->buf =3D kmalloc(rctx->total_len, gfp); + if (!rctx->buf) { + ret =3D -ENOMEM; + goto out_free_chunks; + } + list_for_each_entry(c, &rctx->chunks, list) { + memcpy(rctx->buf + off, c->data, c->len); + off +=3D c->len; + } + } + + /* Tag output buffer */ + rctx->tag_buf =3D kzalloc(AES_CMAC_DIGEST_SIZE, gfp); + if (!rctx->tag_buf) { + ret =3D -ENOMEM; + goto out_free_buf; + } + + rctx->tag_dma =3D cmh_dma_map_single(rctx->tag_buf, + AES_CMAC_DIGEST_SIZE, + DMA_FROM_DEVICE); + if (cmh_dma_map_error(rctx->tag_dma)) { + ret =3D -ENOMEM; + goto out_free_tag; + } + + /* Map input data (may be zero-length for empty CMAC) */ + if (rctx->total_len > 0) { + rctx->in_dma =3D cmh_dma_map_single(rctx->buf, rctx->total_len, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->in_dma)) { + ret =3D -ENOMEM; + goto out_unmap_tag; + } + } + + /* Resolve key */ + idx =3D 0; + + rctx->key_dma =3D tctx->key.raw.dma; + rctx->keylen =3D tctx->key.raw.len; + vcq_add_sys_write(&cmds[idx++], SYS_REF_TEMP, + (u64)rctx->key_dma, SYS_REF_NONE, + tctx->key.raw.len, + tctx->key.raw.sys_type); + key_ref =3D SYS_REF_TEMP; + keylen =3D tctx->key.raw.len; + d =3D cmh_core_select_instance(CMH_CORE_AES); + target_mbx =3D d.mbx_idx; + core_id =3D d.core_id; + + /* + * INIT: mode=3DCMAC, op=3DENCRYPT (CMAC always "encrypts") + * CMAC data goes through the AAD path: + * aadlen =3D total data length, iolen =3D 0 + */ + { + struct vcq_cmd *slot =3D &cmds[idx++]; + + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, AES_CMD_INIT); + slot->hwc.aes.cmd_init.key =3D key_ref; + slot->hwc.aes.cmd_init.iv =3D 0; + slot->hwc.aes.cmd_init.keylen =3D keylen; + slot->hwc.aes.cmd_init.ivlen =3D 0; + slot->hwc.aes.cmd_init.mode =3D AES_MODE_CMAC; + slot->hwc.aes.cmd_init.op =3D AES_OP_ENCRYPT; + slot->hwc.aes.cmd_init.aadlen =3D rctx->total_len; + slot->hwc.aes.cmd_init.iolen =3D 0; + slot->hwc.aes.cmd_init.taglen =3D AES_CMAC_DIGEST_SIZE; + } + + /* AAD_FINAL_AUTH: final AAD + tag extraction in one atomic step */ + { + struct vcq_cmd *slot =3D &cmds[idx++]; + + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, AES_CMD_AAD_FINAL_AUTH); + slot->hwc.aes.cmd_aad_final_auth.data =3D + rctx->total_len > 0 ? (u64)rctx->in_dma : 0; + slot->hwc.aes.cmd_aad_final_auth.datalen =3D rctx->total_len; + slot->hwc.aes.cmd_aad_final_auth.tag =3D (u64)rctx->tag_dma; + slot->hwc.aes.cmd_aad_final_auth.taglen =3D AES_CMAC_DIGEST_SIZE; + } + + vcq_add_flush(&cmds[idx++], core_id); + + ret =3D cmh_vcq_pack_and_submit_async(cmds, idx, rctx->packed, + CMH_AES_CMAC_MAX_PACKED, + target_mbx, + cmh_aes_cmac_complete, req, + !!(req->base.flags & + CRYPTO_TFM_REQ_MAY_BACKLOG), + cmh_tm_async_timeout_jiffies()); + /* -EBUSY =3D backlogged; ownership transferred to callback. */ + if (ret =3D=3D -EBUSY) + return -EBUSY; + if (ret) { + /* + * Synchronous rejection (e.g. -EAGAIN: CMQ full, no backlog). + * Free only the per-submit transients and keep the accumulated + * chunks intact so the caller can retry the identical final(). + * If no retry comes, cra_exit reclaims the orphaned chunks; the + * per-tfm buffered-byte cap bounds how much stays pinned. + */ + if (rctx->total_len > 0 && !cmh_dma_map_error(rctx->in_dma)) + cmh_dma_unmap_single(rctx->in_dma, rctx->total_len, + DMA_TO_DEVICE); + cmh_dma_unmap_single(rctx->tag_dma, AES_CMAC_DIGEST_SIZE, + DMA_FROM_DEVICE); + kfree(rctx->tag_buf); + rctx->tag_buf =3D NULL; + kfree_sensitive(rctx->buf); + rctx->buf =3D NULL; + /* Keep chunks + total_len so the retry rebuilds the buffer. */ + return ret; + } + + return -EINPROGRESS; + +out_unmap_tag: + cmh_dma_unmap_single(rctx->tag_dma, AES_CMAC_DIGEST_SIZE, + DMA_FROM_DEVICE); +out_free_tag: + kfree(rctx->tag_buf); +out_free_buf: +out_free_chunks: + cmh_aes_cmac_free_chunks(rctx, tctx); + kfree_sensitive(rctx->buf); + rctx->buf =3D NULL; + rctx->total_len =3D 0; + return ret; +} + +/* + * ahash .export()/.import(): serialize/deserialize the software + * accumulation buffer. No HW state is involved -- the AES core + * does not support save/restore, but we only export the input queue. + */ + +static int cmh_aes_cmac_export(struct ahash_request *req, void *out) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_aes_cmac_reqctx *rctx =3D ahash_request_ctx(req); + struct cmh_aes_cmac_export_state *state =3D out; + struct cmh_aes_cmac_chunk *chunk; + u32 offset =3D 0; + int ret; + + /* + * If more data is buffered than the flat window holds, switch to + * the software fallback so a bounded, fixed-size state can be + * exported -- making export/import (clone) work at any length. + */ + if (!rctx->switched && rctx->total_len > CMH_AES_CMAC_EXPORT_MAX) { + ret =3D cmh_aes_cmac_switch_to_fb(req); + if (ret) + return ret; + } + + /* Zero the whole state buffer so no kernel memory leaks out. */ + memset(state, 0, crypto_ahash_statesize(tfm)); + + if (rctx->switched) { + state->format =3D CMH_AES_CMAC_FMT_FB; + return crypto_ahash_export(cmh_aes_cmac_fb_req(rctx), + state->data); + } + + state->format =3D CMH_AES_CMAC_FMT_RAW; + state->total_len =3D rctx->total_len; + list_for_each_entry(chunk, &rctx->chunks, list) { + memcpy(state->data + offset, chunk->data, chunk->len); + offset +=3D chunk->len; + } + return 0; +} + +static int cmh_aes_cmac_import(struct ahash_request *req, const void *in) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_aes_cmac_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cmh_aes_cmac_reqctx *rctx =3D ahash_request_ctx(req); + const struct cmh_aes_cmac_export_state *state =3D in; + struct cmh_aes_cmac_chunk *chunk; + + /* + * Do NOT call free_chunks() here: the crypto API does not + * guarantee the request context is in a valid state before + * import(), so the list pointers may be stale or invalid. + * Re-initialize from scratch instead. Any pre-existing chunks + * are tracked on tctx->all_chunks and freed in exit_tfm. + */ + memset(rctx, 0, sizeof(*rctx)); + INIT_LIST_HEAD(&rctx->chunks); + + /* Fallback-format state: replay it into a fallback request. */ + if (state->format =3D=3D CMH_AES_CMAC_FMT_FB) { + struct ahash_request *fb_req =3D cmh_aes_cmac_fb_req(rctx); + int ret; + + ahash_request_set_tfm(fb_req, tctx->fb); + ahash_request_set_callback(fb_req, 0, NULL, NULL); + ret =3D crypto_ahash_import(fb_req, state->data); + if (ret) + return ret; + rctx->switched =3D true; + return 0; + } + + if (state->format !=3D CMH_AES_CMAC_FMT_RAW) + return -EINVAL; + + if (state->total_len > CMH_AES_CMAC_EXPORT_MAX) + return -EINVAL; + + if (state->total_len) { + chunk =3D kmalloc(sizeof(*chunk) + state->total_len, + req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC); + if (!chunk) + return -ENOMEM; + chunk->len =3D state->total_len; + memcpy(chunk->data, state->data, state->total_len); + spin_lock_bh(&tctx->chunk_lock); + list_add_tail(&chunk->list, &rctx->chunks); + list_add_tail(&chunk->tfm_node, &tctx->all_chunks); + tctx->tfm_buffered +=3D chunk->len; + spin_unlock_bh(&tctx->chunk_lock); + rctx->total_len =3D state->total_len; + } + return 0; +} + +static int cmh_aes_cmac_finup(struct ahash_request *req) +{ + int err; + + err =3D cmh_aes_cmac_update(req); + if (err) + return err; + return cmh_aes_cmac_final(req); +} + +static int cmh_aes_cmac_digest(struct ahash_request *req) +{ + int err; + + err =3D cmh_aes_cmac_init(req); + if (err) + return err; + return cmh_aes_cmac_finup(req); +} + +static int cmh_aes_cmac_init_tfm(struct crypto_ahash *tfm) +{ + struct cmh_aes_cmac_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct crypto_ahash *fb; + + memset(tctx, 0, sizeof(*tctx)); + spin_lock_init(&tctx->chunk_lock); + INIT_LIST_HEAD(&tctx->all_chunks); + + /* + * Generic software fallback for oversized input / clone. Masking + * out CRYPTO_ALG_ASYNC excludes this (async) driver so the allocator + * picks the generic cmac(aes); its request is embedded after the + * reqctx. The generic cmac has no core export/import state, so it + * cannot serve as a CRYPTO_ALG_NEED_FALLBACK fallback -- allocate it + * explicitly here instead. + */ + fb =3D crypto_alloc_ahash(crypto_ahash_alg_name(tfm), 0, + CRYPTO_ALG_ASYNC); + if (IS_ERR(fb)) + return PTR_ERR(fb); + tctx->fb =3D fb; + + crypto_ahash_set_reqsize(tfm, + sizeof(struct cmh_aes_cmac_reqctx) + + crypto_tfm_ctx_alignment() + + sizeof(struct ahash_request) + + crypto_ahash_reqsize(fb)); + return 0; +} + +static void cmh_aes_cmac_exit_tfm(struct crypto_ahash *tfm) +{ + struct cmh_aes_cmac_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cmh_aes_cmac_chunk *c, *tmp; + + /* Free any orphaned chunks (e.g. testmgr export/reimport poison) */ + spin_lock_bh(&tctx->chunk_lock); + list_for_each_entry_safe(c, tmp, &tctx->all_chunks, tfm_node) { + list_del(&c->tfm_node); + tctx->tfm_buffered -=3D c->len; + kfree_sensitive(c); + } + spin_unlock_bh(&tctx->chunk_lock); + + if (tctx->fb) + crypto_free_ahash(tctx->fb); + cmh_key_destroy(&tctx->key); +} + +static struct ahash_alg cmh_aes_cmac_alg =3D { + .init =3D cmh_aes_cmac_init, + .update =3D cmh_aes_cmac_update, + .final =3D cmh_aes_cmac_final, + .finup =3D cmh_aes_cmac_finup, + .digest =3D cmh_aes_cmac_digest, + .export =3D cmh_aes_cmac_export, + .import =3D cmh_aes_cmac_import, + .setkey =3D cmh_aes_cmac_setkey, + .init_tfm =3D cmh_aes_cmac_init_tfm, + .exit_tfm =3D cmh_aes_cmac_exit_tfm, + .halg =3D { + .digestsize =3D AES_CMAC_DIGEST_SIZE, + .statesize =3D CMH_AES_CMAC_STATE_SIZE, + .base =3D { + .cra_name =3D "cmac(aes)", + .cra_driver_name =3D "rambus-cmh-cmac-aes", + .cra_priority =3D 300, + .cra_flags =3D CRYPTO_ALG_KERN_DRIVER_ONLY | + CRYPTO_ALG_NO_FALLBACK | + CRYPTO_ALG_ASYNC | + CRYPTO_ALG_REQ_VIRT, + .cra_blocksize =3D AES_CMAC_BLOCK_SIZE, + .cra_ctxsize =3D sizeof(struct cmh_aes_cmac_tfm_ctx), + .cra_module =3D THIS_MODULE, + }, + }, +}; + +/** + * cmh_aes_cmac_register() - Register AES-CMAC hash algorithm with the cry= pto framework + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_aes_cmac_register(void) +{ + int ret; + + if (!cmh_core_present(CMH_CORE_AES)) + return 0; + + ret =3D crypto_register_ahash(&cmh_aes_cmac_alg); + if (ret) + dev_err(cmh_dev(), "cmh_aes_cmac: failed to register cmac(aes) (rc=3D%d)= \n", + ret); + else + dev_dbg(cmh_dev(), "cmh_aes_cmac: registered cmac(aes)\n"); + + return ret; +} + +/** + * cmh_aes_cmac_unregister() - Unregister AES-CMAC hash algorithm from the= crypto framework + */ +void cmh_aes_cmac_unregister(void) +{ + if (!cmh_core_present(CMH_CORE_AES)) + return; + + crypto_unregister_ahash(&cmh_aes_cmac_alg); + dev_dbg(cmh_dev(), "cmh_aes_cmac: unregistered cmac(aes)\n"); +} diff --git a/drivers/crypto/cmh/cmh_main.c b/drivers/crypto/cmh/cmh_main.c index 4ad1500d8e49..a9985ad95c7d 100644 --- a/drivers/crypto/cmh/cmh_main.c +++ b/drivers/crypto/cmh/cmh_main.c @@ -36,6 +36,7 @@ #include "cmh_cshake.h" #include "cmh_kmac.h" #include "cmh_sm3.h" +#include "cmh_aes.h" #include "cmh_mgmt.h" #include "cmh_registers.h" #include "cmh_debugfs.h" @@ -232,6 +233,21 @@ static int cmh_probe(struct platform_device *pdev) if (ret) goto err_sm3_register; =20 + /* Register AES skcipher algorithms */ + ret =3D cmh_aes_register(); + if (ret) + goto err_aes_register; + + /* Register AES AEAD algorithms (GCM, CCM) */ + ret =3D cmh_aes_aead_register(); + if (ret) + goto err_aes_aead_register; + + /* Register AES CMAC algorithm */ + ret =3D cmh_aes_cmac_register(); + if (ret) + goto err_aes_cmac_register; + /* Register key management device (/dev/cmh_mgmt) */ ret =3D cmh_mgmt_register(); if (ret) @@ -242,6 +258,12 @@ static int cmh_probe(struct platform_device *pdev) return 0; =20 err_mgmt_register: + cmh_aes_cmac_unregister(); +err_aes_cmac_register: + cmh_aes_aead_unregister(); +err_aes_aead_register: + cmh_aes_unregister(); +err_aes_register: cmh_sm3_unregister(); err_sm3_register: cmh_kmac_unregister(); @@ -278,6 +300,9 @@ static void cmh_remove(struct platform_device *pdev) cfg =3D &dev->config; =20 cmh_mgmt_unregister(); + cmh_aes_cmac_unregister(); + cmh_aes_aead_unregister(); + cmh_aes_unregister(); cmh_sm3_unregister(); cmh_kmac_unregister(); cmh_cshake_unregister(); diff --git a/drivers/crypto/cmh/include/cmh_aes.h b/drivers/crypto/cmh/incl= ude/cmh_aes.h new file mode 100644 index 000000000000..591afaa36f85 --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_aes.h @@ -0,0 +1,24 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- AES Crypto API Drivers + * + * Registers AES algorithms with the Linux crypto subsystem: + * skcipher: ecb/cbc/ctr/cfb/xts(aes) + * aead: gcm/ccm(aes) + * shash: cmac(aes) + */ + +#ifndef CMH_AES_H +#define CMH_AES_H + +int cmh_aes_register(void); +void cmh_aes_unregister(void); + +int cmh_aes_aead_register(void); +void cmh_aes_aead_unregister(void); + +int cmh_aes_cmac_register(void); +void cmh_aes_cmac_unregister(void); + +#endif /* CMH_AES_H */ --=20 2.43.7 From nobody Fri Sep 25 01:20:34 2026 Received: from CH4PR04CU002.outbound.protection.outlook.com (mail-northcentralusazon11023093.outbound.protection.outlook.com [40.107.201.93]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 804773ACA54; Thu, 17 Sep 2026 22:59:41 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=40.107.201.93 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789685990; cv=fail; b=JTg6fiP+RUt0Yr8Htu7mWxch3+gzRA/7S5AMKGPVE8id0UD/Rah2lsHJwPI4TaN85esZdZdMb4UQligywIQtwgPDTqF3YGSMu8fuszNuEeexF1t+KDuvnf6NsCpCxYt01hW3N2jokDKV3zvDPPk/HM4tF7aR7lj0UXWHI257qCE= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789685990; c=relaxed/simple; bh=DqLQocR4sclAG26zn4LkBrSQ0qQb1Unknly1JWxcocs=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=RZ82GWel281Ws/5A88oxkCSEitj43QbMzo5b1lQWW+g2IEjkyOZ3IcE1wkKgBB7Og4W2weyXISVc5H1869pUuKvW9lfz3VQGfN+hmGDF4zc9SAfteODuQqNYGpkGDZXOmxT2Hmb7ULbEkyD3vJK4m1ackQ8RLTIqC7uPWJfl/tM= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com; spf=fail smtp.mailfrom=rambus.com; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b=VnVWfTd+; arc=fail smtp.client-ip=40.107.201.93 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=rambus.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b="VnVWfTd+" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=XDEDips/yX/kw03zJZ58NHWWu2LF9SRQ66XPVM+NYqb+vYhpAZZgIw7e4ff/ukZHsnW/d1+WTe9IQz83yzYGfpe1HCnlNHY+RzjsMiEzHsgQdNIoRq0CHDXRm7PB96NtFhHYvqauy7/B9MQTeOLi7VQ8hPI/AZNUd/P/STYM+r0RazBha/jRwcYcdnGT8SXQLzuem0JxfxwqkWiovNV7NNeUAJiaCdQDg1qPD4iYoPgAs5oHUpbdDvJRG0Cadh5O9VHcFCTaH518F7fuK7hVNUMXm/mTV4lQaLW3GrdgVQZkJm0l/fE1SjKKuAQVlARGsQC6ib0p5j91oGLY9KAXEA== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=sVlEh+p45ITtFuAmBtxvpfv7+u0k5R9uOaHrWbpw/gE=; b=DTYoTcaQvoZd0hir8/DwKyXnsqtssP9tFobBTRVg9vSwEoBCoU8xj9/cMQjwQnATJU8cGRf9KogdaKn/lhSg7Kg/V80lPNaHFKVifzZ3CoaLjbGMrkQhOt4PJwh8E9Ttnv5SSpPCLeIRM5Ot3YVPSIgRwOE1CbYWSw/LD+oUtYxauYo9yF2S59DSEaGYYG1v796CuEFHXPAmxnW6YGxLcNhmLkc8Od+foW55DlrrVPwoG/8UfWLfrCT7xCCD2TPTihPyG+AHtDU0Sm38vH3Q9M3GEHfMNsbg7GItnM2JXQKtDrJmaimVQfUxyNRaYaXZHfDKJD/FAIAvNkc8IN+E/A== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 192.86.86.210) smtp.rcpttodomain=cryptography.com smtp.mailfrom=rambus.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=rambus.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=rambus.com; s=selector1; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=sVlEh+p45ITtFuAmBtxvpfv7+u0k5R9uOaHrWbpw/gE=; b=VnVWfTd+mQKbvtIa11Z169YGifV2LeVMkvBNIKV9Ve1XBrYDckxbG+UZsJ7IxE1NOv1L1AzCnGNmwe6T2hfHBXQiAUK2uReRvI8Y7zT/BjdXktyboCELjI7KRZ+uhcX7mFcVxarBCkYdaYgEt/OY/V3kP+lnRX0Upp4MH5HQtwA1tCI+APlSCUigrwJK+CBiWh1p7A/0pGDg19xW2IXGQSTqUs2Wtl1Ko13fm+6xDR6+Qn3cjzoEgtVuS570SWO7rI4IO+Rb0DskflkZKyFBZ2BAqhyhMYfAppuOO8gv3xN2kTIXcpgo9e4GRxBsHEjBy6CuwUNji5Ku3sHLkITGpg== Received: from YQBPR0101CA0354.CANPRD01.PROD.OUTLOOK.COM (2603:10b6:c01:6b::11) by FT1PR04MB269544.namprd04.prod.outlook.com (2603:10b6:170:b::11) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.428.13; Thu, 17 Sep 2026 22:59:35 +0000 Received: from BL6PEPF00020E64.namprd04.prod.outlook.com (2603:10b6:c01:6b:cafe::f) by YQBPR0101CA0354.outlook.office365.com (2603:10b6:c01:6b::11) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.428.12 via Frontend Transport; Thu, 17 Sep 2026 22:59:34 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 192.86.86.210) smtp.mailfrom=rambus.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=rambus.com; Received-SPF: Pass (protection.outlook.com: domain of rambus.com designates 192.86.86.210 as permitted sender) receiver=protection.outlook.com; client-ip=192.86.86.210; helo=hqxsv-psmtppxy02.rambus.com; pr=C Received: from hqxsv-psmtppxy02.rambus.com (192.86.86.210) by BL6PEPF00020E64.mail.protection.outlook.com (10.167.249.25) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.451.8 via Frontend Transport; Thu, 17 Sep 2026 22:59:34 +0000 Received: from hqxsv-cmdev3-aousherovitch.rambus.com (hqn-lb-int-float.rambus.com [10.12.20.20]) by hqxsv-psmtppxy02.rambus.com (Postfix) with ESMTP id 328441801765; Thu, 17 Sep 2026 22:59:32 +0000 (UTC) From: Alex Ousherovitch To: Albert Ou , Alex Ousherovitch , Conor Dooley , "David S. Miller" , Herbert Xu , Jonathan Corbet , Krzysztof Kozlowski , Palmer Dabbelt , Paul Walmsley , Rob Herring , Saravanakrishnan Krishnamoorthy , Shuah Khan Cc: Alexandre Ghiti , devicetree@vger.kernel.org, Joel Wittenauer , linux-api@vger.kernel.org, linux-crypto@vger.kernel.org, linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-riscv@lists.infradead.org, Shuah Khan , Thi Nguyen Subject: [PATCH v5 09/19] crypto: cmh - add SM4 skcipher/aead/cmac/xcbc Date: Thu, 17 Sep 2026 15:59:18 -0700 Message-ID: <20260917225929.2494111-10-aousherovitch@rambus.com> X-Mailer: git-send-email 2.43.7 In-Reply-To: <20260917225929.2494111-1-aousherovitch@rambus.com> References: <20260917225929.2494111-1-aousherovitch@rambus.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: BL6PEPF00020E64:EE_|FT1PR04MB269544:EE_ X-MS-Office365-Filtering-Correlation-Id: f982e427-b517-4489-80f9-08df150f5762 X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|376014|7416014|36860700016|1800799024|82310400026|23010399003|6133799003|3023799007|10067099003|11063799006|56012099006|921020|22082099003|18002099003; X-Microsoft-Antispam-Message-Info: ARqcCucJb2BvHJBHUdT4dU14cCINkwif6MRdY66Zig8Mbrp9djVk/BVTMgAz59ja287az2HO3CsBaxkymFQ/YOWGSqNH2DJchTetV1IM2OC3Idmf0eu/qFC1kLAeGiWicNZrmjdq7nfm7tt0NNt/Qs+7cyWc7jwlcXhYKjHNIsGz8S8O5fIW9MdF4NDYyyvrrC368zIHsb9Tm02NAUFN/nayUKxBMMlhQxYCqf+sOpQC1cbPKS+xo/8yRf8AxE9nIoIVdrwn40K6PlR6vSY2HgtPAYfk26+X3D4AoVyWOuTtXdW8FnkvUEGtVQoAJsrIJKfzxeL74P5I+bS+wqL94c/KHFcLjJGQI9QnOq/YKYN4recNSN5QsOtTi4IXowP6/i/YHdWoEEo7JNby/KsSlqObHdFTNz/N173SrK35I3s5vcbRGCXm4mqekNx5DUW22pe3im3upJPFhDPY6GXsZVHnZ/Blxv+r7H+c3t8cbbNiG8nBR1J/+9UsowtMe7zF9a5gLVL1nj3zRIBZug9dClotfR/JYKJkJhzjlUqdEewH6IhHMoIQqd5Gm4nlr7V+rUbbKDAwi5gkPPGQ4yOJlZpo4ei2dhhebajn24C7aO3AXO6w/CBGK6qKAVt6FZlBRYsNmX/1bkMO/sO15rbGyouYj847POna5BoUuizsnDJb4mGO2xROfFHxJPYJkOoovsZDm7UwAPnT6dhcz14bFoA3Uabq8P1YJd6jqxkXgeU1EkgE2Y3p0L8DtskhAUh0 X-Forefront-Antispam-Report: CIP:192.86.86.210;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:hqxsv-psmtppxy02.rambus.com;PTR:InfoDomainNonexistent;CAT:NONE;SFS:(13230040)(376014)(7416014)(36860700016)(1800799024)(82310400026)(23010399003)(6133799003)(3023799007)(10067099003)(11063799006)(56012099006)(921020)(22082099003)(18002099003);DIR:OUT;SFP:1102; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: 947jDc8wPnOF0beNMMSeCW4oJqw5/HLoDwFhwUS7vOe/bi5uuKzMC4sx6AbReLNxGLoWx+AgdZc39XRhwIEdTcTIGPFSQ1jrqRWmfPdL8qpVvcE4cgkeIUxjCgBo8OHzyV2jkYCuAImK+lramnsVoOZujPSaCwEsl1znFm0xwOKgvWSC7WZH6QE8wENquPNAzI1kuf3a3RRVj89Yb6vmTj2BZ8eA2xQzRllKvVcvLqukXssmUWkBsFQSGWSwdOtTUi1xPYo5VUVXcUZ/VwRid+Pb+TmyAHKcYgJF5r2rD1N9anD1TAq+97Y8XcmsEjBM9w+IiJ6eL4IhB1kjy9wfb2fggSRCQvmT42Ydq8R4D1TLZIQmVlFQsGfo82AkaLvx4FYHfk/lFHyP5FSVBPoaAoiQXXc3FEWF3iPMcJvTuIh8wtFPui2XXXgpzP5jkmO8 X-OriginatorOrg: rambus.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 17 Sep 2026 22:59:34.2309 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: f982e427-b517-4489-80f9-08df150f5762 X-MS-Exchange-CrossTenant-Id: bd0ba799-c2b9-413c-9c56-5d1731c4827c X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=bd0ba799-c2b9-413c-9c56-5d1731c4827c;Ip=[192.86.86.210];Helo=[hqxsv-psmtppxy02.rambus.com] X-MS-Exchange-CrossTenant-AuthSource: BL6PEPF00020E64.namprd04.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: FT1PR04MB269544 Content-Type: text/plain; charset="utf-8" Register SM4 algorithms using the CMH SM4 core (core ID 0x04): - skcipher: SM4-ECB, SM4-CBC, SM4-CTR, SM4-XTS, SM4-CFB - aead: SM4-GCM, SM4-CCM - ahash: SM4-CMAC, SM4-XCBC Signed-off-by: Alex Ousherovitch Co-developed-by: Saravanakrishnan Krishnamoorthy Signed-off-by: Saravanakrishnan Krishnamoorthy --- drivers/crypto/cmh/Makefile | 5 +- drivers/crypto/cmh/cmh_main.c | 25 + drivers/crypto/cmh/cmh_sm4_aead.c | 884 +++++++++++++++++++++++ drivers/crypto/cmh/cmh_sm4_cmac.c | 968 ++++++++++++++++++++++++++ drivers/crypto/cmh/cmh_sm4_skcipher.c | 709 +++++++++++++++++++ drivers/crypto/cmh/include/cmh_sm4.h | 24 + 6 files changed, 2614 insertions(+), 1 deletion(-) create mode 100644 drivers/crypto/cmh/cmh_sm4_aead.c create mode 100644 drivers/crypto/cmh/cmh_sm4_cmac.c create mode 100644 drivers/crypto/cmh/cmh_sm4_skcipher.c create mode 100644 drivers/crypto/cmh/include/cmh_sm4.h diff --git a/drivers/crypto/cmh/Makefile b/drivers/crypto/cmh/Makefile index c5dcfa4f7794..032c863d856e 100644 --- a/drivers/crypto/cmh/Makefile +++ b/drivers/crypto/cmh/Makefile @@ -22,7 +22,10 @@ cmh-y :=3D \ cmh_sm3.o \ cmh_aes.o \ cmh_aes_aead.o \ - cmh_aes_cmac.o + cmh_aes_cmac.o \ + cmh_sm4_skcipher.o \ + cmh_sm4_aead.o \ + cmh_sm4_cmac.o =20 # Management ioctl device (/dev/cmh_mgmt): key lifecycle, PKE, PQC ioctls. cmh-$(CONFIG_CRYPTO_DEV_CMH_MGMT) +=3D \ diff --git a/drivers/crypto/cmh/cmh_main.c b/drivers/crypto/cmh/cmh_main.c index a9985ad95c7d..410941cc04e4 100644 --- a/drivers/crypto/cmh/cmh_main.c +++ b/drivers/crypto/cmh/cmh_main.c @@ -37,6 +37,7 @@ #include "cmh_kmac.h" #include "cmh_sm3.h" #include "cmh_aes.h" +#include "cmh_sm4.h" #include "cmh_mgmt.h" #include "cmh_registers.h" #include "cmh_debugfs.h" @@ -248,6 +249,21 @@ static int cmh_probe(struct platform_device *pdev) if (ret) goto err_aes_cmac_register; =20 + /* Register SM4 skcipher algorithms */ + ret =3D cmh_sm4_register(); + if (ret) + goto err_sm4_register; + + /* Register SM4 AEAD algorithms (GCM, CCM) */ + ret =3D cmh_sm4_aead_register(); + if (ret) + goto err_sm4_aead_register; + + /* Register SM4 CMAC/XCBC algorithms */ + ret =3D cmh_sm4_cmac_register(); + if (ret) + goto err_sm4_cmac_register; + /* Register key management device (/dev/cmh_mgmt) */ ret =3D cmh_mgmt_register(); if (ret) @@ -258,6 +274,12 @@ static int cmh_probe(struct platform_device *pdev) return 0; =20 err_mgmt_register: + cmh_sm4_cmac_unregister(); +err_sm4_cmac_register: + cmh_sm4_aead_unregister(); +err_sm4_aead_register: + cmh_sm4_unregister(); +err_sm4_register: cmh_aes_cmac_unregister(); err_aes_cmac_register: cmh_aes_aead_unregister(); @@ -300,6 +322,9 @@ static void cmh_remove(struct platform_device *pdev) cfg =3D &dev->config; =20 cmh_mgmt_unregister(); + cmh_sm4_cmac_unregister(); + cmh_sm4_aead_unregister(); + cmh_sm4_unregister(); cmh_aes_cmac_unregister(); cmh_aes_aead_unregister(); cmh_aes_unregister(); diff --git a/drivers/crypto/cmh/cmh_sm4_aead.c b/drivers/crypto/cmh/cmh_sm4= _aead.c new file mode 100644 index 000000000000..e89b4a9f97e5 --- /dev/null +++ b/drivers/crypto/cmh/cmh_sm4_aead.c @@ -0,0 +1,884 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Kernel Crypto API SM4 AEAD Driver (GCM/CCM) + * + * Registers AEAD algorithms with the Linux crypto subsystem: + * gcm(sm4), ccm(sm4) + * + * GCM: SM4_CMD_INIT(mode=3DGCM) + [AAD_FINAL] + SM4_CMD_FINAL + FLUSH + * CCM: SM4_CMD_CCM_INIT + [AAD_FINAL] + SM4_CMD_FINAL + FLUSH + * - SM4 CCM uses a distinct sm4_cmd_ccm_init struct + */ + +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "cmh_sm4.h" +#include "cmh_vcq.h" +#include "cmh_sm4_abi.h" +#include "cmh_sys_abi.h" +#include "cmh_sys.h" +#include "cmh_txn.h" +#include "cmh_dma.h" +#include "cmh_key.h" + +/* + * GCM IV contract: + * + * The SM4 core requires exactly 16 bytes loaded into its IV register. + * For standard 96-bit nonce GCM, the driver passes: + * + * IV[0..11] =3D user-supplied 12-byte nonce + * IV[12..15] =3D 0x00000000 + * + * The hardware internally sets the last 32 bits to the big-endian + * counter value 1 (forming J0 =3D nonce || 0x00000001) before + * processing AAD. The driver must NOT pre-set the counter. + * + * If the IV format is incorrect, GCM authentication will fail + * (encrypt produces wrong ciphertext/tag, decrypt rejects). + */ +#define SM4_GCM_IV_SIZE 12U /* GCM nonce size (standard) */ +#define SM4_GCM_HW_IV_SIZE 16U /* HW requires 16-byte IV buffer */ +#define SM4_GCM_TAG_SIZE 16U + +/* CCM: callers pass a 16-byte IV in RFC 3610 format: + * iv[0] =3D L-1, iv[1..14-iv[0]] =3D nonce, rest =3D counter (zeroed). + * Nonce length =3D 14 - iv[0], range 7..13. + */ +#define SM4_CCM_IV_SIZE 16U + +enum cmh_sm4_aead_type { + CMH_SM4_AEAD_GCM, + CMH_SM4_AEAD_CCM, +}; + +struct cmh_sm4_aead_info { + enum cmh_sm4_aead_type type; + u32 sm4_mode; + u32 ivsize; + u32 maxauthsize; + const char *alg_name; + const char *drv_name; +}; + +static const struct cmh_sm4_aead_info sm4_aead_algs[] =3D { + { CMH_SM4_AEAD_GCM, SM4_MODE_GCM, SM4_GCM_IV_SIZE, + SM4_GCM_TAG_SIZE, "gcm(sm4)", "rambus-cmh-gcm-sm4" }, + { CMH_SM4_AEAD_CCM, SM4_MODE_CCM, SM4_CCM_IV_SIZE, + SM4_GCM_TAG_SIZE, "ccm(sm4)", "rambus-cmh-ccm-sm4" }, +}; + +struct cmh_sm4_aead_tfm_ctx { + struct cmh_key_ctx key; + u32 authsize; + struct crypto_cipher *sw_cipher; /* CCM empty-input fallback */ +}; + +/* Per-request context (lives in aead_request::__ctx) */ + +#define CMH_SM4_AEAD_MAX_PAYLOAD 5 +#define CMH_SM4_AEAD_MAX_PACKED (CMH_SM4_AEAD_MAX_PAYLOAD * 2) + +struct cmh_sm4_aead_reqctx { + dma_addr_t in_dma; + dma_addr_t out_dma; + dma_addr_t iv_dma; + dma_addr_t key_dma; + dma_addr_t aad_dma; + dma_addr_t tag_dma; + u8 *in_buf; + u8 *out_buf; + u8 *iv_buf; + u8 *aad_buf; + u8 *tag_buf; + u32 cryptlen; + u32 assoclen; + u32 authsize; + u32 iv_map_len; + u32 keylen; + bool encrypting; + bool empty_gcm_fallback; + struct vcq_cmd packed[CMH_SM4_AEAD_MAX_PACKED]; +}; + +struct cmh_sm4_aead_drv { + struct aead_alg alg; + const struct cmh_sm4_aead_info *info; +}; + +static const struct cmh_sm4_aead_info * +cmh_sm4_aead_get_info(struct crypto_aead *tfm) +{ + struct aead_alg *alg =3D crypto_aead_alg(tfm); + + return container_of(alg, struct cmh_sm4_aead_drv, alg)->info; +} + +/* VCQ Builders -- SM4 AEAD-specific */ + +static void vcq_add_sm4_aead_init(struct vcq_cmd *slot, u32 core_id, u64 k= ey_ref, + u64 iv_dma, u32 keylen, u32 ivlen, + u32 mode, u32 op, u32 aadlen, u32 iolen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, SM4_CMD_INIT); + slot->hwc.sm4.cmd_init.key =3D key_ref; + slot->hwc.sm4.cmd_init.iv =3D iv_dma; + slot->hwc.sm4.cmd_init.keylen =3D keylen; + slot->hwc.sm4.cmd_init.ivlen =3D ivlen; + slot->hwc.sm4.cmd_init.mode =3D mode; + slot->hwc.sm4.cmd_init.op =3D op; + slot->hwc.sm4.cmd_init.aadlen =3D aadlen; + slot->hwc.sm4.cmd_init.iolen =3D iolen; +} + +static void vcq_add_sm4_ccm_init(struct vcq_cmd *slot, u32 core_id, u64 ke= y_ref, + u64 nonce_dma, u32 keylen, u32 noncelen, + u32 op, u32 aadlen, u32 iolen, u32 taglen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, SM4_CMD_CCM_INIT); + slot->hwc.sm4.cmd_ccm_init.key =3D key_ref; + slot->hwc.sm4.cmd_ccm_init.nonce =3D nonce_dma; + slot->hwc.sm4.cmd_ccm_init.keylen =3D keylen; + slot->hwc.sm4.cmd_ccm_init.noncelen =3D noncelen; + slot->hwc.sm4.cmd_ccm_init.op =3D op; + slot->hwc.sm4.cmd_ccm_init.aadlen =3D aadlen; + slot->hwc.sm4.cmd_ccm_init.iolen =3D iolen; + slot->hwc.sm4.cmd_ccm_init.taglen =3D taglen; +} + +static void vcq_add_sm4_aad_final(struct vcq_cmd *slot, u32 core_id, u64 a= ad_dma, + u32 aadlen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, SM4_CMD_AAD_FINAL); + slot->hwc.sm4.cmd_aad_final.data =3D aad_dma; + slot->hwc.sm4.cmd_aad_final.datalen =3D aadlen; +} + +static void vcq_add_sm4_aead_final(struct vcq_cmd *slot, u32 core_id, u64 = input_dma, + u64 output_dma, u64 tag_dma, + u32 iolen, u32 taglen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, SM4_CMD_FINAL); + slot->hwc.sm4.cmd_final.input =3D input_dma; + slot->hwc.sm4.cmd_final.output =3D output_dma; + slot->hwc.sm4.cmd_final.tag =3D tag_dma; + slot->hwc.sm4.cmd_final.iolen =3D iolen; + slot->hwc.sm4.cmd_final.taglen =3D taglen; +} + +/* setkey */ +static int cmh_sm4_aead_setkey(struct crypto_aead *tfm, const u8 *key, + unsigned int keylen) +{ + struct cmh_sm4_aead_tfm_ctx *tctx =3D crypto_aead_ctx(tfm); + /* SM4 always uses 128-bit keys */ + if (keylen !=3D CMH_SM4_KEY_SIZE) + return -EINVAL; + + if (tctx->sw_cipher) { + int ret; + + ret =3D crypto_cipher_setkey(tctx->sw_cipher, key, keylen); + if (ret) + return ret; + } + + return cmh_key_setkey_raw(&tctx->key, key, keylen, CORE_ID_SM4); +} + +static int cmh_sm4_aead_setauthsize(struct crypto_aead *tfm, + unsigned int authsize) +{ + struct cmh_sm4_aead_tfm_ctx *tctx =3D crypto_aead_ctx(tfm); + const struct cmh_sm4_aead_info *info =3D cmh_sm4_aead_get_info(tfm); + + if (info->type =3D=3D CMH_SM4_AEAD_GCM) { + /* eSW enforces taglen =3D=3D 16 for SM4 GCM (EIP40_SM4_TAG_SIZE) */ + if (authsize !=3D 16) + return -EINVAL; + } else { + /* CCM: accept 4, 6, 8, 10, 12, 14, 16 per RFC 3610 */ + if (authsize < 4 || authsize > 16 || (authsize & 1)) + return -EINVAL; + } + + tctx->authsize =3D authsize; + return 0; +} + +static int cmh_sm4_aead_init_tfm(struct crypto_aead *tfm) +{ + struct cmh_sm4_aead_tfm_ctx *tctx =3D crypto_aead_ctx(tfm); + const struct cmh_sm4_aead_info *info =3D cmh_sm4_aead_get_info(tfm); + + memset(tctx, 0, sizeof(*tctx)); + tctx->authsize =3D info->maxauthsize; + + if (info->type =3D=3D CMH_SM4_AEAD_CCM) { + struct crypto_cipher *ci; + + ci =3D crypto_alloc_cipher("sm4", 0, 0); + if (IS_ERR(ci)) + return PTR_ERR(ci); + tctx->sw_cipher =3D ci; + } + + crypto_aead_set_reqsize(tfm, sizeof(struct cmh_sm4_aead_reqctx)); + return 0; +} + +static void cmh_sm4_aead_exit_tfm(struct crypto_aead *tfm) +{ + struct cmh_sm4_aead_tfm_ctx *tctx =3D crypto_aead_ctx(tfm); + + if (tctx->sw_cipher) + crypto_free_cipher(tctx->sw_cipher); + cmh_key_destroy(&tctx->key); +} + +/* DMA unmap helper */ +static void cmh_sm4_aead_unmap_dma(struct cmh_sm4_aead_reqctx *rctx) +{ + u32 tag_map_len; + + cmh_dma_unmap_single(rctx->iv_dma, rctx->iv_map_len, DMA_TO_DEVICE); + tag_map_len =3D rctx->empty_gcm_fallback ? + SM4_GCM_HW_IV_SIZE : rctx->authsize; + cmh_dma_unmap_single(rctx->tag_dma, tag_map_len, + (rctx->encrypting || rctx->empty_gcm_fallback) ? + DMA_FROM_DEVICE : DMA_TO_DEVICE); + if (rctx->cryptlen > 0) { + cmh_dma_unmap_single(rctx->out_dma, rctx->cryptlen, + DMA_FROM_DEVICE); + cmh_dma_unmap_single(rctx->in_dma, rctx->cryptlen, + DMA_TO_DEVICE); + } + if (rctx->assoclen > 0) + cmh_dma_unmap_single(rctx->aad_dma, rctx->assoclen, + DMA_TO_DEVICE); +} + +static void cmh_sm4_aead_free_bufs(struct cmh_sm4_aead_reqctx *rctx) +{ + kfree(rctx->iv_buf); + rctx->iv_buf =3D NULL; + kfree(rctx->tag_buf); + rctx->tag_buf =3D NULL; + kfree_sensitive(rctx->out_buf); + rctx->out_buf =3D NULL; + kfree_sensitive(rctx->in_buf); + rctx->in_buf =3D NULL; + kfree(rctx->aad_buf); + rctx->aad_buf =3D NULL; +} + +static void cmh_sm4_aead_complete(void *data, int error) +{ + struct aead_request *req =3D data; + struct cmh_sm4_aead_reqctx *rctx =3D aead_request_ctx(req); + + if (error =3D=3D -EINPROGRESS) { + cmh_complete(&req->base, error); + return; + } + + cmh_sm4_aead_unmap_dma(rctx); + + /* + * Map HW error on decrypt to -EBADMSG. The eSW SM4 core uses a + * single error code (-EIO) for both authentication failures and + * other core errors (e.g. DMA timeout), so we cannot distinguish + * them from the MBX_STATUS alone. In practice the only error + * during a well-formed AEAD decrypt is auth-tag mismatch; a DMA + * timeout would indicate a fatal HW problem where -EBADMSG vs + * -EIO is moot. The kernel crypto API requires -EBADMSG for + * AEAD authentication failures. + */ + if (error =3D=3D -EIO && !rctx->encrypting) + error =3D -EBADMSG; + + if (!error) { + if (rctx->empty_gcm_fallback && !rctx->encrypting) { + if (crypto_memneq(rctx->tag_buf, rctx->in_buf, + rctx->authsize)) + error =3D -EBADMSG; + } + if (!error && rctx->cryptlen > 0) + scatterwalk_map_and_copy(rctx->out_buf, req->dst, + req->assoclen, + rctx->cryptlen, 1); + /* + * Out-of-place (src !=3D dst): the AAD is authenticated but not + * transformed, so the crypto API expects it copied verbatim + * into the dst AAD region (in-place dst already holds it). + */ + if (!error && req->src !=3D req->dst && req->assoclen > 0) + scatterwalk_map_and_copy(rctx->aad_buf, req->dst, + 0, req->assoclen, 1); + if (!error && rctx->encrypting) + scatterwalk_map_and_copy(rctx->tag_buf, req->dst, + req->assoclen + + rctx->cryptlen, + rctx->authsize, 1); + } + + cmh_sm4_aead_free_bufs(rctx); + cmh_complete(&req->base, error); +} + +/* + * GCM empty-input fallback (SM4). + * + * When both AAD and plaintext are empty, GCM reduces to: + * tag =3D E(K, J0) where J0 =3D nonce || 0x00000001 + * + * The eSW GCM engine rejects this degenerate case, so we compute it + * via a single ECB block encryption of J0. + * + * VCQ: [SYS_CMD_WRITE] + SM4_CMD_INIT(ECB) + SM4_CMD_FINAL + FLUSH + */ +static int cmh_sm4_gcm_empty(struct aead_request *req, u32 sm4_op) +{ + struct crypto_aead *tfm =3D crypto_aead_reqtfm(req); + struct cmh_sm4_aead_tfm_ctx *tctx =3D crypto_aead_ctx(tfm); + struct cmh_sm4_aead_reqctx *rctx =3D aead_request_ctx(req); + struct vcq_cmd cmds[CMH_SM4_AEAD_MAX_PAYLOAD]; + u64 key_ref; + u32 keylen, authsize; + struct core_dispatch d; + s32 target_mbx; + u32 core_id; + u32 idx; + int ret; + gfp_t gfp; + + authsize =3D tctx->authsize; + + gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + + memset(rctx, 0, sizeof(*rctx)); + rctx->cryptlen =3D 0; + rctx->assoclen =3D 0; + rctx->authsize =3D authsize; + rctx->encrypting =3D (sm4_op =3D=3D SM4_OP_ENCRYPT); + rctx->empty_gcm_fallback =3D true; + + /* Build J0 =3D nonce || 0x00000001 in iv_buf */ + rctx->iv_buf =3D kzalloc(SM4_GCM_HW_IV_SIZE, gfp); + if (!rctx->iv_buf) + return -ENOMEM; + memcpy(rctx->iv_buf, req->iv, SM4_GCM_IV_SIZE); + rctx->iv_buf[15] =3D 0x01; + rctx->iv_map_len =3D SM4_GCM_HW_IV_SIZE; + + rctx->iv_dma =3D cmh_dma_map_single(rctx->iv_buf, SM4_GCM_HW_IV_SIZE, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->iv_dma)) { + ret =3D -ENOMEM; + goto out_free_iv; + } + + /* Tag buffer -- receives E(K, J0) output */ + rctx->tag_buf =3D kzalloc(SM4_GCM_HW_IV_SIZE, gfp); + if (!rctx->tag_buf) { + ret =3D -ENOMEM; + goto out_unmap_iv; + } + rctx->tag_dma =3D cmh_dma_map_single(rctx->tag_buf, SM4_GCM_HW_IV_SIZE, + DMA_FROM_DEVICE); + if (cmh_dma_map_error(rctx->tag_dma)) { + ret =3D -ENOMEM; + goto out_free_tag; + } + + /* For decrypt: read expected tag from request */ + if (!rctx->encrypting) { + rctx->in_buf =3D kmalloc(authsize, gfp); + if (!rctx->in_buf) { + ret =3D -ENOMEM; + goto out_unmap_tag; + } + scatterwalk_map_and_copy(rctx->in_buf, req->src, 0, + authsize, 0); + } + + /* Resolve key */ + idx =3D 0; + rctx->key_dma =3D tctx->key.raw.dma; + vcq_add_sys_write(&cmds[idx++], SYS_REF_TEMP, + (u64)rctx->key_dma, SYS_REF_NONE, + tctx->key.raw.len, + tctx->key.raw.sys_type); + key_ref =3D SYS_REF_TEMP; + keylen =3D tctx->key.raw.len; + d =3D cmh_core_select_instance(CMH_CORE_SM4); + target_mbx =3D d.mbx_idx; + core_id =3D d.core_id; + + /* ECB INIT: single block encryption of J0 */ + vcq_add_sm4_aead_init(&cmds[idx++], core_id, key_ref, + 0, keylen, 0, SM4_MODE_ECB, SM4_OP_ENCRYPT, + 0, SM4_GCM_HW_IV_SIZE); + + /* FINAL: J0 in, E(K,J0) out */ + vcq_add_sm4_aead_final(&cmds[idx++], core_id, + (u64)rctx->iv_dma, (u64)rctx->tag_dma, + 0, SM4_GCM_HW_IV_SIZE, 0); + + vcq_add_flush(&cmds[idx++], core_id); + + ret =3D cmh_vcq_pack_and_submit_async(cmds, idx, rctx->packed, + CMH_SM4_AEAD_MAX_PACKED, + target_mbx, + cmh_sm4_aead_complete, req, + !!(req->base.flags & + CRYPTO_TFM_REQ_MAY_BACKLOG), + cmh_tm_async_timeout_jiffies()); + if (ret =3D=3D -EBUSY) + return -EBUSY; + if (ret) + goto out_free_in; + + return -EINPROGRESS; + +out_free_in: + kfree_sensitive(rctx->in_buf); +out_unmap_tag: + cmh_dma_unmap_single(rctx->tag_dma, SM4_GCM_HW_IV_SIZE, + DMA_FROM_DEVICE); +out_free_tag: + kfree(rctx->tag_buf); +out_unmap_iv: + cmh_dma_unmap_single(rctx->iv_dma, SM4_GCM_HW_IV_SIZE, DMA_TO_DEVICE); +out_free_iv: + kfree(rctx->iv_buf); + return ret; +} + +/* + * CCM empty-input fallback (SM4). + * + * When both AAD and plaintext are empty, CCM reduces to: + * T =3D E(K, B0) -- CBC-MAC of the single formatting block + * S0 =3D E(K, A0) -- CTR block zero + * tag =3D (T XOR S0)[0..authsize-1] + * + * The eSW rejects this degenerate case, so the driver computes it + * synchronously via two crypto_cipher single-block encryptions. + */ +static int cmh_sm4_ccm_empty(struct aead_request *req, u32 sm4_op) +{ + struct crypto_aead *tfm =3D crypto_aead_reqtfm(req); + struct cmh_sm4_aead_tfm_ctx *tctx =3D crypto_aead_ctx(tfm); + u32 authsize =3D tctx->authsize; + u8 b0[CMH_SM4_BLOCK_SIZE], a0[CMH_SM4_BLOCK_SIZE]; + u8 t[CMH_SM4_BLOCK_SIZE], s0[CMH_SM4_BLOCK_SIZE]; + u8 tag[CMH_SM4_BLOCK_SIZE]; + u8 L; + u32 i; + + /* Defense-in-depth: iv[0] =3D L-1, valid L is 2..8 per RFC 3610 S2.1 */ + if (WARN_ON_ONCE(req->iv[0] < 1 || req->iv[0] > 7)) + return -EINVAL; + + L =3D req->iv[0] + 1; + + if (tctx->key.mode !=3D CMH_KEY_RAW) + return -EOPNOTSUPP; + + /* B0: flags || nonce || Q(=3D0). Adata=3D0, t=3Dauthsize, q=3DL. */ + memset(b0, 0, CMH_SM4_BLOCK_SIZE); + b0[0] =3D (u8)(8 * ((authsize - 2) / 2) + (L - 1)); + memcpy(&b0[1], &req->iv[1], 15 - L); + + /* A0: (L-1) || nonce || counter(=3D0) */ + memset(a0, 0, CMH_SM4_BLOCK_SIZE); + a0[0] =3D (u8)(L - 1); + memcpy(&a0[1], &req->iv[1], 15 - L); + + crypto_cipher_encrypt_one(tctx->sw_cipher, t, b0); + crypto_cipher_encrypt_one(tctx->sw_cipher, s0, a0); + + for (i =3D 0; i < authsize; i++) + tag[i] =3D t[i] ^ s0[i]; + + if (sm4_op =3D=3D SM4_OP_ENCRYPT) { + scatterwalk_map_and_copy(tag, req->dst, + req->assoclen, authsize, 1); + } else { + u8 expected[CMH_SM4_BLOCK_SIZE]; + + scatterwalk_map_and_copy(expected, req->src, + req->assoclen, authsize, 0); + if (crypto_memneq(tag, expected, authsize)) + return -EBADMSG; + } + + return 0; +} + +static int cmh_sm4_aead_crypt(struct aead_request *req, u32 sm4_op) +{ + struct crypto_aead *tfm =3D crypto_aead_reqtfm(req); + struct cmh_sm4_aead_tfm_ctx *tctx =3D crypto_aead_ctx(tfm); + const struct cmh_sm4_aead_info *info =3D cmh_sm4_aead_get_info(tfm); + struct cmh_sm4_aead_reqctx *rctx =3D aead_request_ctx(req); + struct vcq_cmd cmds[CMH_SM4_AEAD_MAX_PAYLOAD]; + u64 key_ref; + u32 keylen, authsize, cryptlen; + struct core_dispatch d; + s32 target_mbx; + u32 core_id; + u32 idx; + int ret; + gfp_t gfp; + + if (tctx->key.mode =3D=3D CMH_KEY_NONE) + return -ENOKEY; + + authsize =3D tctx->authsize; + + if (sm4_op =3D=3D SM4_OP_ENCRYPT) { + cryptlen =3D req->cryptlen; + } else { + if (req->cryptlen < authsize) + return -EINVAL; + cryptlen =3D req->cryptlen - authsize; + } + + /* + * Validate CCM IV format early -- the empty-input fallback and + * nonce extraction both depend on iv[0] being in range [1,7]. + */ + if (info->type =3D=3D CMH_SM4_AEAD_CCM) { + if (req->iv[0] < 1 || req->iv[0] > 7) + return -EINVAL; + } + + /* + * The CMH eSW rejects SM4 GCM/CCM when both aadlen and iolen + * are zero. For GCM, the tag is simply E(K, J0) -- use ECB + * fallback. For CCM, compute tag =3D E(K,B0) XOR E(K,A0) in SW. + */ + if (cryptlen =3D=3D 0 && req->assoclen =3D=3D 0) { + if (info->type =3D=3D CMH_SM4_AEAD_GCM) + return cmh_sm4_gcm_empty(req, sm4_op); + return cmh_sm4_ccm_empty(req, sm4_op); + } + + /* + * HW uses a proprietary LLI scatter-gather format that is + * incompatible with struct scatterlist, so the payload is + * linearised into contiguous buffers for DMA. Cap total + * size to prevent excessive memory consumption. + */ + if ((u64)cryptlen + req->assoclen > SZ_1M) + return -EINVAL; + + gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + + memset(rctx, 0, sizeof(*rctx)); + rctx->cryptlen =3D cryptlen; + rctx->assoclen =3D req->assoclen; + rctx->authsize =3D authsize; + rctx->encrypting =3D (sm4_op =3D=3D SM4_OP_ENCRYPT); + + /* Linearise AAD */ + if (req->assoclen > 0) { + rctx->aad_buf =3D kmalloc(req->assoclen, gfp | __GFP_NOWARN); + if (!rctx->aad_buf) + return -ENOMEM; + scatterwalk_map_and_copy(rctx->aad_buf, req->src, + 0, req->assoclen, 0); + rctx->aad_dma =3D cmh_dma_map_single(rctx->aad_buf, + req->assoclen, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->aad_dma)) { + ret =3D -ENOMEM; + goto out_free_aad; + } + } + + /* Linearise input */ + if (cryptlen > 0) { + rctx->in_buf =3D kmalloc(cryptlen, gfp | __GFP_NOWARN); + if (!rctx->in_buf) { + ret =3D -ENOMEM; + goto out_unmap_aad; + } + scatterwalk_map_and_copy(rctx->in_buf, req->src, + req->assoclen, cryptlen, 0); + rctx->in_dma =3D cmh_dma_map_single(rctx->in_buf, cryptlen, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->in_dma)) { + ret =3D -ENOMEM; + goto out_free_in; + } + } + + /* Allocate output buffer */ + if (cryptlen > 0) { + rctx->out_buf =3D kmalloc(cryptlen, gfp | __GFP_NOWARN); + if (!rctx->out_buf) { + ret =3D -ENOMEM; + goto out_unmap_in; + } + rctx->out_dma =3D cmh_dma_map_single(rctx->out_buf, cryptlen, + DMA_FROM_DEVICE); + if (cmh_dma_map_error(rctx->out_dma)) { + ret =3D -ENOMEM; + goto out_free_out; + } + } + + /* Tag buffer */ + rctx->tag_buf =3D kmalloc(authsize, gfp); + if (!rctx->tag_buf) { + ret =3D -ENOMEM; + goto out_unmap_out; + } + + if (!rctx->encrypting) { + scatterwalk_map_and_copy(rctx->tag_buf, req->src, + req->assoclen + cryptlen, + authsize, 0); + } else { + memset(rctx->tag_buf, 0, authsize); + } + + rctx->tag_dma =3D cmh_dma_map_single(rctx->tag_buf, authsize, + rctx->encrypting ? + DMA_FROM_DEVICE : DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->tag_dma)) { + ret =3D -ENOMEM; + goto out_free_tag; + } + + /* Map IV/nonce */ + if (info->type =3D=3D CMH_SM4_AEAD_GCM) { + rctx->iv_buf =3D kzalloc(SM4_GCM_HW_IV_SIZE, gfp); + if (!rctx->iv_buf) { + ret =3D -ENOMEM; + goto out_unmap_tag; + } + memcpy(rctx->iv_buf, req->iv, SM4_GCM_IV_SIZE); + rctx->iv_map_len =3D SM4_GCM_HW_IV_SIZE; + rctx->iv_dma =3D cmh_dma_map_single(rctx->iv_buf, + rctx->iv_map_len, + DMA_TO_DEVICE); + } else { + u32 noncelen; + + if (req->iv[0] < 1 || req->iv[0] > 7) { + ret =3D -EINVAL; + goto out_unmap_tag; + } + noncelen =3D 14 - req->iv[0]; + + rctx->iv_buf =3D kmemdup(req->iv + 1, noncelen, gfp); + if (!rctx->iv_buf) { + ret =3D -ENOMEM; + goto out_unmap_tag; + } + rctx->iv_map_len =3D noncelen; + rctx->iv_dma =3D cmh_dma_map_single(rctx->iv_buf, + rctx->iv_map_len, + DMA_TO_DEVICE); + } + if (cmh_dma_map_error(rctx->iv_dma)) { + ret =3D -ENOMEM; + goto out_free_iv; + } + + /* Resolve key reference */ + idx =3D 0; + + rctx->key_dma =3D tctx->key.raw.dma; + rctx->keylen =3D tctx->key.raw.len; + vcq_add_sys_write(&cmds[idx++], SYS_REF_TEMP, + (u64)rctx->key_dma, SYS_REF_NONE, + tctx->key.raw.len, + tctx->key.raw.sys_type); + key_ref =3D SYS_REF_TEMP; + keylen =3D tctx->key.raw.len; + d =3D cmh_core_select_instance(CMH_CORE_SM4); + target_mbx =3D d.mbx_idx; + core_id =3D d.core_id; + + /* Build INIT command */ + if (info->type =3D=3D CMH_SM4_AEAD_CCM) { + vcq_add_sm4_ccm_init(&cmds[idx++], core_id, key_ref, + (u64)rctx->iv_dma, keylen, + rctx->iv_map_len, sm4_op, + req->assoclen, cryptlen, authsize); + } else { + vcq_add_sm4_aead_init(&cmds[idx++], core_id, key_ref, + (u64)rctx->iv_dma, keylen, + SM4_GCM_HW_IV_SIZE, info->sm4_mode, + sm4_op, req->assoclen, cryptlen); + } + + if (req->assoclen > 0) + vcq_add_sm4_aad_final(&cmds[idx++], core_id, + (u64)rctx->aad_dma, req->assoclen); + + vcq_add_sm4_aead_final(&cmds[idx++], core_id, + cryptlen > 0 ? (u64)rctx->in_dma : 0, + cryptlen > 0 ? (u64)rctx->out_dma : 0, + (u64)rctx->tag_dma, cryptlen, authsize); + + vcq_add_flush(&cmds[idx++], core_id); + + ret =3D cmh_vcq_pack_and_submit_async(cmds, idx, rctx->packed, + CMH_SM4_AEAD_MAX_PACKED, + target_mbx, + cmh_sm4_aead_complete, req, + !!(req->base.flags & + CRYPTO_TFM_REQ_MAY_BACKLOG), + cmh_tm_async_timeout_jiffies()); + if (ret =3D=3D -EBUSY) + return -EBUSY; + if (ret) + goto out_cleanup_all; + + return -EINPROGRESS; + +out_cleanup_all: + cmh_dma_unmap_single(rctx->iv_dma, rctx->iv_map_len, DMA_TO_DEVICE); +out_free_iv: + kfree(rctx->iv_buf); +out_unmap_tag: + cmh_dma_unmap_single(rctx->tag_dma, authsize, + rctx->encrypting ? DMA_FROM_DEVICE : + DMA_TO_DEVICE); +out_free_tag: + kfree(rctx->tag_buf); +out_unmap_out: + if (cryptlen > 0) + cmh_dma_unmap_single(rctx->out_dma, cryptlen, DMA_FROM_DEVICE); +out_free_out: + kfree_sensitive(rctx->out_buf); +out_unmap_in: + if (cryptlen > 0) + cmh_dma_unmap_single(rctx->in_dma, cryptlen, DMA_TO_DEVICE); +out_free_in: + kfree_sensitive(rctx->in_buf); +out_unmap_aad: + if (req->assoclen > 0) + cmh_dma_unmap_single(rctx->aad_dma, req->assoclen, + DMA_TO_DEVICE); +out_free_aad: + kfree(rctx->aad_buf); + return ret; +} + +static int cmh_sm4_aead_encrypt(struct aead_request *req) +{ + return cmh_sm4_aead_crypt(req, SM4_OP_ENCRYPT); +} + +static int cmh_sm4_aead_decrypt(struct aead_request *req) +{ + return cmh_sm4_aead_crypt(req, SM4_OP_DECRYPT); +} + +/* Registration */ + +static struct cmh_sm4_aead_drv sm4_aead_drv_algs[ARRAY_SIZE(sm4_aead_algs)= ]; + +/** + * cmh_sm4_aead_register() - Register SM4-GCM/CCM AEAD algorithms with the= crypto framework + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_sm4_aead_register(void) +{ + unsigned int i; + int ret; + + if (!cmh_core_present(CMH_CORE_SM4)) + return 0; + + for (i =3D 0; i < ARRAY_SIZE(sm4_aead_algs); i++) { + const struct cmh_sm4_aead_info *info =3D &sm4_aead_algs[i]; + struct cmh_sm4_aead_drv *drv =3D &sm4_aead_drv_algs[i]; + struct aead_alg *alg =3D &drv->alg; + + drv->info =3D info; + + memset(alg, 0, sizeof(*alg)); + + alg->setkey =3D cmh_sm4_aead_setkey; + alg->setauthsize =3D cmh_sm4_aead_setauthsize; + alg->encrypt =3D cmh_sm4_aead_encrypt; + alg->decrypt =3D cmh_sm4_aead_decrypt; + alg->init =3D cmh_sm4_aead_init_tfm; + alg->exit =3D cmh_sm4_aead_exit_tfm; + alg->ivsize =3D info->ivsize; + alg->maxauthsize =3D info->maxauthsize; + + strscpy(alg->base.cra_name, info->alg_name, + CRYPTO_MAX_ALG_NAME); + strscpy(alg->base.cra_driver_name, info->drv_name, + CRYPTO_MAX_ALG_NAME); + alg->base.cra_priority =3D 300; + alg->base.cra_flags =3D CRYPTO_ALG_KERN_DRIVER_ONLY | + CRYPTO_ALG_ASYNC; + alg->base.cra_blocksize =3D 1; + alg->base.cra_ctxsize =3D sizeof(struct cmh_sm4_aead_tfm_ctx); + alg->base.cra_module =3D THIS_MODULE; + + ret =3D crypto_register_aead(alg); + if (ret) { + dev_err(cmh_dev(), "cmh_sm4_aead: failed to register %s (rc=3D%d)\n", + info->alg_name, ret); + goto err_unregister; + } + + dev_dbg(cmh_dev(), "cmh_sm4_aead: registered %s\n", info->alg_name); + } + + return 0; + +err_unregister: + while (i--) + crypto_unregister_aead(&sm4_aead_drv_algs[i].alg); + return ret; +} + +/** + * cmh_sm4_aead_unregister() - Unregister SM4 AEAD algorithms from the cry= pto framework + */ +void cmh_sm4_aead_unregister(void) +{ + unsigned int i; + + if (!cmh_core_present(CMH_CORE_SM4)) + return; + + for (i =3D 0; i < ARRAY_SIZE(sm4_aead_algs); i++) { + crypto_unregister_aead(&sm4_aead_drv_algs[i].alg); + dev_dbg(cmh_dev(), "cmh_sm4_aead: unregistered %s\n", + sm4_aead_algs[i].alg_name); + } +} diff --git a/drivers/crypto/cmh/cmh_sm4_cmac.c b/drivers/crypto/cmh/cmh_sm4= _cmac.c new file mode 100644 index 000000000000..abeadce312c0 --- /dev/null +++ b/drivers/crypto/cmh/cmh_sm4_cmac.c @@ -0,0 +1,968 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Kernel Crypto API SM4-CMAC / SM4-XCBC (ahash) Driver + * + * Registers cmac(sm4) and xcbc(sm4) as ahash algorithms. + * + * Both produce a 16-byte tag (MAC) from a key and message. + * VCQ sequence: [SYS_CMD_WRITE] + SM4_CMD_INIT(CMAC/XCBC) + + * SM4_CMD_AAD_FINAL + SM4_CMD_FINAL + FLUSH + */ + +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "cmh_sm4.h" +#include "cmh_vcq.h" +#include "cmh_sm4_abi.h" +#include "cmh_sys_abi.h" +#include "cmh_sys.h" +#include "cmh_txn.h" +#include "cmh_dma.h" +#include "cmh_key.h" + +#define SM4_MAC_DIGEST_SIZE 16U +#define SM4_MAC_BLOCK_SIZE 16U +/* + * Maximum accumulated data for SM4 MAC -- driver-imposed, not HW. + * + * The SM4 core does not expose external save/restore VCQ commands, + * so the driver must accumulate all data in kernel memory via + * .update() and submit it atomically in .final(). This cap limits + * the per-request kernel allocation. + */ +#define SM4_MAC_MAX_DATA (64 * 1024) + +struct cmh_sm4_mac_alg_info { + u32 sm4_mode; /* SM4_MODE_CMAC or SM4_MODE_XCBC */ + const char *alg_name; + const char *drv_name; +}; + +static const struct cmh_sm4_mac_alg_info sm4_mac_algs[] =3D { + { SM4_MODE_CMAC, "cmac(sm4)", "rambus-cmh-cmac-sm4" }, + { SM4_MODE_XCBC, "xcbc(sm4)", "rambus-cmh-xcbc-sm4" }, +}; + +struct cmh_sm4_mac_tfm_ctx { + struct cmh_key_ctx key; + u32 sm4_mode; + struct crypto_cipher *sw_cipher; /* empty-input fallback (CMAC/XCBC) */ + struct crypto_ahash *fb; /* generic SW fallback (oversized ops) */ + /* Cached subkeys (derived at setkey time for concurrency safety) */ + u8 xcbc_k1[CMH_SM4_BLOCK_SIZE]; /* K1 =3D E(K, 0x01..01) */ + u8 xcbc_k3[CMH_SM4_BLOCK_SIZE]; /* K3 =3D E(K, 0x03..03) */ + u8 cmac_k2[CMH_SM4_BLOCK_SIZE]; /* K2 =3D dbl(dbl(E(K, 0))) */ + bool subkeys_valid; + spinlock_t chunk_lock; /* protects all_chunks + tfm_buffered */ + struct list_head all_chunks; /* orphan-safe chunk tracking */ + size_t tfm_buffered; /* bytes on all_chunks; DoS cap */ +}; + +/* + * Per-transform cap on total bytes buffered across all_chunks. Bounds + * memory an AF_ALG client can pin via repeated open/update/abandon of + * request sockets (the crypto API has no per-request destructor). + */ +#define CMH_SM4_MAC_TFM_MAX_BUFFERED (16 * 1024 * 1024) + +/* Per-request context (lives in ahash_request::__ctx) */ +/* Chunk node for O(1) update() appends */ +struct cmh_sm4_mac_chunk { + struct list_head list; + struct list_head tfm_node; /* per-tfm orphan tracking */ + u32 len; + u8 data[]; +}; + +/* Per-request context (lives in ahash_request::__ctx) */ + +#define CMH_SM4_MAC_MAX_PAYLOAD 5 +#define CMH_SM4_MAC_MAX_PACKED (CMH_SM4_MAC_MAX_PAYLOAD * 2) + +struct cmh_sm4_mac_reqctx { + struct list_head chunks; + u32 total_len; + bool switched; /* handed off to SW fallback */ + u8 *buf; /* linearised in final() */ + /* DMA state for async final */ + dma_addr_t key_dma; + dma_addr_t in_dma; + dma_addr_t tag_dma; + u8 *tag_buf; + u32 keylen; + struct vcq_cmd packed[CMH_SM4_MAC_MAX_PACKED]; +}; + +/* + * Flat state for export/import, tagged by the leading @format byte: + * CMH_SM4_MAC_FMT_RAW is the accumulated input bytes (flat window); + * CMH_SM4_MAC_FMT_FB is the software fallback's own exported state, + * used once the request has switched to the fallback so export/import + * (transform clone) works at any input length. + */ +#define CMH_SM4_MAC_FMT_RAW 0 +#define CMH_SM4_MAC_FMT_FB 1 + +struct cmh_sm4_mac_export_state { + u8 format; + u8 __pad[3]; + u32 total_len; + u8 data[]; +}; + +/* + * The crypto subsystem pre-allocates statesize bytes per request. + * CMH_SM4_MAC_STATE_SIZE (4096) sizes both the CMH_SM4_MAC_FMT_RAW + * window (CMH_SM4_MAC_EXPORT_MAX accumulated bytes) and the + * CMH_SM4_MAC_FMT_FB software state. A RAW export past + * CMH_SM4_MAC_EXPORT_MAX transparently switches to the fallback and + * emits CMH_SM4_MAC_FMT_FB instead, so export/import is not capped. + */ +#define CMH_SM4_MAC_STATE_SIZE 4096 +#define CMH_SM4_MAC_EXPORT_MAX \ + (CMH_SM4_MAC_STATE_SIZE - sizeof(struct cmh_sm4_mac_export_state)) + +struct cmh_sm4_mac_drv { + struct ahash_alg alg; + const struct cmh_sm4_mac_alg_info *info; +}; + +/* + * GF(2^128) doubling used to derive the CMAC subkeys (NIST SP 800-38B). + * Shift the 128-bit big-endian value left by one bit and, if the top bit + * was set, reduce with Rb =3D 0x87. + */ +static void cmh_sm4_cmac_dbl(u8 out[CMH_SM4_BLOCK_SIZE], + const u8 in[CMH_SM4_BLOCK_SIZE]) +{ + u8 carry =3D in[0] >> 7; + unsigned int i; + + for (i =3D 0; i < CMH_SM4_BLOCK_SIZE - 1; i++) + out[i] =3D (in[i] << 1) | (in[i + 1] >> 7); + out[CMH_SM4_BLOCK_SIZE - 1] =3D (in[CMH_SM4_BLOCK_SIZE - 1] << 1) ^ + (carry ? 0x87 : 0x00); +} + +static int cmh_sm4_mac_setkey(struct crypto_ahash *tfm, const u8 *key, + unsigned int keylen) +{ + struct cmh_sm4_mac_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + int ret; + + if (keylen !=3D CMH_SM4_KEY_SIZE) + return -EINVAL; + + /* + * Program the HW key first; only (re)derive the software subkeys on + * success so a failed HW step cannot leave the SW subkeys (new) + * inconsistent with the HW key (old). + */ + if (tctx->sw_cipher) + tctx->subkeys_valid =3D false; + + ret =3D cmh_key_setkey_raw(&tctx->key, key, keylen, CORE_ID_SM4); + if (ret) + return ret; + + if (tctx->sw_cipher && tctx->sm4_mode =3D=3D SM4_MODE_XCBC) { + u8 const1[CMH_SM4_BLOCK_SIZE], const3[CMH_SM4_BLOCK_SIZE]; + + ret =3D crypto_cipher_setkey(tctx->sw_cipher, key, keylen); + if (ret) + return ret; + + /* Pre-derive XCBC subkeys for concurrent-safe final() */ + memset(const1, 0x01, CMH_SM4_BLOCK_SIZE); + memset(const3, 0x03, CMH_SM4_BLOCK_SIZE); + crypto_cipher_encrypt_one(tctx->sw_cipher, tctx->xcbc_k1, + const1); + crypto_cipher_encrypt_one(tctx->sw_cipher, tctx->xcbc_k3, + const3); + + /* + * Leave sw_cipher keyed with K1 permanently. + * final() only needs E(K1, block) and never touches the + * original key again, so no re-keying in the hot path + * eliminates the per-tfm concurrency race entirely. + */ + ret =3D crypto_cipher_setkey(tctx->sw_cipher, tctx->xcbc_k1, + CMH_SM4_BLOCK_SIZE); + if (ret) + return ret; + } else if (tctx->sw_cipher && tctx->sm4_mode =3D=3D SM4_MODE_CMAC) { + u8 zero[CMH_SM4_BLOCK_SIZE] =3D { 0 }; + u8 l[CMH_SM4_BLOCK_SIZE], k1[CMH_SM4_BLOCK_SIZE]; + + ret =3D crypto_cipher_setkey(tctx->sw_cipher, key, keylen); + if (ret) + return ret; + + /* + * Pre-derive the CMAC subkey K2 for the empty-message + * fallback (NIST SP 800-38B): + * L =3D E(K, 0^128); K1 =3D dbl(L); K2 =3D dbl(K1) + * sw_cipher is left keyed with the original K, so final() + * computes E(K, K2 ^ pad) with no hot-path re-keying. + */ + crypto_cipher_encrypt_one(tctx->sw_cipher, l, zero); + cmh_sm4_cmac_dbl(k1, l); + cmh_sm4_cmac_dbl(tctx->cmac_k2, k1); + memzero_explicit(l, sizeof(l)); + memzero_explicit(k1, sizeof(k1)); + } + + if (tctx->sw_cipher) + tctx->subkeys_valid =3D true; + + /* Keep the software fallback keyed in lock-step for oversized ops. */ + return crypto_ahash_setkey(tctx->fb, key, keylen); +} + +static void cmh_sm4_mac_free_chunks(struct cmh_sm4_mac_reqctx *rctx, + struct cmh_sm4_mac_tfm_ctx *tctx) +{ + struct cmh_sm4_mac_chunk *c, *tmp; + + spin_lock_bh(&tctx->chunk_lock); + list_for_each_entry_safe(c, tmp, &rctx->chunks, list) { + list_del(&c->list); + list_del(&c->tfm_node); + tctx->tfm_buffered -=3D c->len; + kfree_sensitive(c); + } + spin_unlock_bh(&tctx->chunk_lock); + rctx->total_len =3D 0; +} + +/* Software-fallback helpers (arbitrary-length + transform-clone support) = */ + +/* + * The fallback ahash_request lives immediately after the reqctx; + * cmh_sm4_mac_init_tfm() reserves crypto_ahash_reqsize(fb) bytes for it. + */ +static struct ahash_request * +cmh_sm4_mac_fb_req(struct cmh_sm4_mac_reqctx *rctx) +{ + return PTR_ALIGN((void *)(rctx + 1), crypto_tfm_ctx_alignment()); +} + +static int cmh_sm4_mac_fb_update_virt(struct ahash_request *fb_req, + const u8 *data, u32 len) +{ + ahash_request_set_virt(fb_req, data, NULL, len); + return crypto_ahash_update(fb_req); +} + +/* + * Switch a request from the HW-buffered path to the software fallback + * (the generic cmac(sm4)/xcbc(sm4) matching this transform): initialise + * the fallback request, replay every accumulated chunk through it, then + * drop the chunks. The fallback transform was keyed by + * cmh_sm4_mac_setkey() when the caller installed the MAC key. + */ +static int cmh_sm4_mac_switch_to_fb(struct ahash_request *req) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_sm4_mac_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cmh_sm4_mac_reqctx *rctx =3D ahash_request_ctx(req); + struct ahash_request *fb_req =3D cmh_sm4_mac_fb_req(rctx); + struct cmh_sm4_mac_chunk *c; + int ret; + + ahash_request_set_tfm(fb_req, tctx->fb); + ahash_request_set_callback(fb_req, 0, NULL, NULL); + + ret =3D crypto_ahash_init(fb_req); + if (ret) + return ret; + + list_for_each_entry(c, &rctx->chunks, list) { + ret =3D cmh_sm4_mac_fb_update_virt(fb_req, c->data, c->len); + if (ret) + return ret; + } + + cmh_sm4_mac_free_chunks(rctx, tctx); + rctx->switched =3D true; + return 0; +} + +/* Forward the current update() payload to the fallback. */ +static int cmh_sm4_mac_fb_forward(struct ahash_request *req, + struct cmh_sm4_mac_reqctx *rctx) +{ + struct ahash_request *fb_req =3D cmh_sm4_mac_fb_req(rctx); + + if (req->base.flags & CRYPTO_AHASH_REQ_VIRT) + return cmh_sm4_mac_fb_update_virt(fb_req, req->svirt, + req->nbytes); + ahash_request_set_crypt(fb_req, req->src, NULL, req->nbytes); + return crypto_ahash_update(fb_req); +} + +static int cmh_sm4_mac_init(struct ahash_request *req) +{ + struct cmh_sm4_mac_reqctx *rctx =3D ahash_request_ctx(req); + + memset(rctx, 0, sizeof(*rctx)); + INIT_LIST_HEAD(&rctx->chunks); + return 0; +} + +static int cmh_sm4_mac_update(struct ahash_request *req) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_sm4_mac_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cmh_sm4_mac_reqctx *rctx =3D ahash_request_ctx(req); + struct cmh_sm4_mac_chunk *chunk; + gfp_t gfp; + int ret; + + if (!req->nbytes) + return 0; + + /* Already handed off to the fallback: forward directly (O(1) mem). */ + if (rctx->switched) + return cmh_sm4_mac_fb_forward(req, rctx); + + /* + * Exceeding the HW input cap: switch to the software fallback + * (replaying the buffered chunks) rather than failing, then + * forward this update. + */ + if (req->nbytes > SM4_MAC_MAX_DATA - rctx->total_len) { + ret =3D cmh_sm4_mac_switch_to_fb(req); + if (ret) + goto err_free_chunks; + return cmh_sm4_mac_fb_forward(req, rctx); + } + + gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + chunk =3D kmalloc(sizeof(*chunk) + req->nbytes, gfp); + if (!chunk) { + ret =3D -ENOMEM; + goto err_free_chunks; + } + + chunk->len =3D req->nbytes; + if (req->base.flags & CRYPTO_AHASH_REQ_VIRT) + memcpy(chunk->data, req->svirt, req->nbytes); + else + scatterwalk_map_and_copy(chunk->data, req->src, + 0, req->nbytes, 0); + list_add_tail(&chunk->list, &rctx->chunks); + spin_lock_bh(&tctx->chunk_lock); + if (tctx->tfm_buffered + chunk->len > CMH_SM4_MAC_TFM_MAX_BUFFERED) { + spin_unlock_bh(&tctx->chunk_lock); + list_del(&chunk->list); + kfree_sensitive(chunk); + ret =3D -ENOMEM; + goto err_free_chunks; + } + list_add_tail(&chunk->tfm_node, &tctx->all_chunks); + tctx->tfm_buffered +=3D chunk->len; + spin_unlock_bh(&tctx->chunk_lock); + rctx->total_len +=3D req->nbytes; + return 0; + +err_free_chunks: + /* + * Terminal error -- free all previously accumulated chunks. + * callers may not call .final() on error, so they would leak. + */ + cmh_sm4_mac_free_chunks(rctx, tctx); + return ret; +} + +static void cmh_sm4_mac_complete(void *data, int error) +{ + struct ahash_request *req =3D data; + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_sm4_mac_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cmh_sm4_mac_reqctx *rctx =3D ahash_request_ctx(req); + + if (error =3D=3D -EINPROGRESS) { + cmh_complete(&req->base, error); + return; + } + + if (rctx->total_len > 0) + cmh_dma_unmap_single(rctx->in_dma, rctx->total_len, + DMA_TO_DEVICE); + cmh_dma_unmap_single(rctx->tag_dma, SM4_MAC_DIGEST_SIZE, + DMA_FROM_DEVICE); + + if (!error) + memcpy(req->result, rctx->tag_buf, SM4_MAC_DIGEST_SIZE); + + kfree(rctx->tag_buf); + rctx->tag_buf =3D NULL; + cmh_sm4_mac_free_chunks(rctx, tctx); + kfree_sensitive(rctx->buf); + rctx->buf =3D NULL; + rctx->total_len =3D 0; + cmh_complete(&req->base, error); +} + +static int cmh_sm4_mac_final(struct ahash_request *req) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_sm4_mac_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cmh_sm4_mac_reqctx *rctx =3D ahash_request_ctx(req); + struct vcq_cmd cmds[CMH_SM4_MAC_MAX_PAYLOAD]; + u64 key_ref; + u32 keylen; + struct core_dispatch d; + s32 target_mbx; + u32 core_id; + u32 idx; + int ret; + gfp_t gfp; + + if (tctx->key.mode =3D=3D CMH_KEY_NONE) { + ret =3D -ENOKEY; + goto out_free_chunks; + } + + /* + * Switched to the generic software fallback (input over the HW cap, + * or an oversized clone was imported): complete there. Must precede + * the empty-input paths below -- a switch frees the chunks, so + * total_len is 0 even though data was streamed to the fallback. + */ + if (rctx->switched) { + struct ahash_request *fb_req =3D cmh_sm4_mac_fb_req(rctx); + + ahash_request_set_crypt(fb_req, NULL, req->result, 0); + return crypto_ahash_final(fb_req); + } + + /* + * XCBC empty-input SW fallback (RFC 3566). + * + * For a zero-length message: + * K1 =3D E(K, 0x01010101...) -- encryption subkey + * K3 =3D E(K, 0x03030303...) -- incomplete-block subkey + * pad =3D 0x80 00...00 -- single 1 bit + 127 zero bits + * tag =3D E(K1, pad XOR K3) + * + * The eSW produces incorrect output for this case, so the driver + * computes it synchronously using crypto_cipher. + * + * For DS keys we cannot derive subkeys (no raw key material), + * and the HW also cannot handle empty XCBC correctly, so + * return -EOPNOTSUPP. + */ + if (rctx->total_len =3D=3D 0 && tctx->sm4_mode =3D=3D SM4_MODE_XCBC) { + u8 block[CMH_SM4_BLOCK_SIZE]; + u32 i; + + if (tctx->key.mode !=3D CMH_KEY_RAW || + !tctx->subkeys_valid) { + cmh_sm4_mac_free_chunks(rctx, tctx); + return -EOPNOTSUPP; + } + + /* block =3D pad XOR K3 */ + memset(block, 0, CMH_SM4_BLOCK_SIZE); + block[0] =3D 0x80; + for (i =3D 0; i < CMH_SM4_BLOCK_SIZE; i++) + block[i] ^=3D tctx->xcbc_k3[i]; + + /* + * tag =3D E(K1, block) + * + * sw_cipher is permanently keyed with K1 (set at setkey + * time), so this is safe for concurrent requests sharing + * the same tfm -- no re-keying, no race. + */ + crypto_cipher_encrypt_one(tctx->sw_cipher, req->result, + block); + + cmh_sm4_mac_free_chunks(rctx, tctx); + return 0; + } + + /* + * CMAC empty-input SW fallback (NIST SP 800-38B). + * + * For a zero-length message the sole block is incomplete, so the + * K2 subkey is used: + * pad =3D 0x80 00...00 -- single 1 bit + 127 zero bits + * tag =3D E(K, pad XOR K2) + * + * The eSW produces incorrect output for this case, so the driver + * computes it synchronously using crypto_cipher. + * + * For DS keys we cannot derive subkeys (no raw key material), + * and the HW also cannot handle empty CMAC correctly, so + * return -EOPNOTSUPP. + */ + if (rctx->total_len =3D=3D 0 && tctx->sm4_mode =3D=3D SM4_MODE_CMAC) { + u8 block[CMH_SM4_BLOCK_SIZE]; + u32 i; + + if (tctx->key.mode !=3D CMH_KEY_RAW || !tctx->subkeys_valid) { + cmh_sm4_mac_free_chunks(rctx, tctx); + return -EOPNOTSUPP; + } + + /* block =3D pad XOR K2 */ + memset(block, 0, CMH_SM4_BLOCK_SIZE); + block[0] =3D 0x80; + for (i =3D 0; i < CMH_SM4_BLOCK_SIZE; i++) + block[i] ^=3D tctx->cmac_k2[i]; + + /* + * tag =3D E(K, block). sw_cipher is keyed with the original + * key K (set at setkey time, never re-keyed), so this is + * safe for concurrent requests sharing the same tfm. + */ + crypto_cipher_encrypt_one(tctx->sw_cipher, req->result, + block); + + cmh_sm4_mac_free_chunks(rctx, tctx); + return 0; + } + + gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + + /* Linearise chunks into a single contiguous buffer for DMA */ + if (rctx->total_len > 0) { + struct cmh_sm4_mac_chunk *c; + u32 off =3D 0; + + rctx->buf =3D kmalloc(rctx->total_len, gfp); + if (!rctx->buf) { + ret =3D -ENOMEM; + goto out_free_chunks; + } + list_for_each_entry(c, &rctx->chunks, list) { + memcpy(rctx->buf + off, c->data, c->len); + off +=3D c->len; + } + } + + rctx->tag_buf =3D kzalloc(SM4_MAC_DIGEST_SIZE, gfp); + if (!rctx->tag_buf) { + ret =3D -ENOMEM; + goto out_free_buf; + } + + rctx->tag_dma =3D cmh_dma_map_single(rctx->tag_buf, + SM4_MAC_DIGEST_SIZE, + DMA_FROM_DEVICE); + if (cmh_dma_map_error(rctx->tag_dma)) { + ret =3D -ENOMEM; + goto out_free_tag; + } + + if (rctx->total_len > 0) { + rctx->in_dma =3D cmh_dma_map_single(rctx->buf, rctx->total_len, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->in_dma)) { + ret =3D -ENOMEM; + goto out_unmap_tag; + } + } + + idx =3D 0; + + rctx->key_dma =3D tctx->key.raw.dma; + rctx->keylen =3D tctx->key.raw.len; + vcq_add_sys_write(&cmds[idx++], SYS_REF_TEMP, + (u64)rctx->key_dma, SYS_REF_NONE, + tctx->key.raw.len, + tctx->key.raw.sys_type); + key_ref =3D SYS_REF_TEMP; + keylen =3D tctx->key.raw.len; + d =3D cmh_core_select_instance(CMH_CORE_SM4); + target_mbx =3D d.mbx_idx; + core_id =3D d.core_id; + + /* + * INIT: mode=3DCMAC or XCBC + * CMAC/XCBC data goes through the AAD path: + * aadlen =3D total data length, iolen =3D 0 + */ + { + struct vcq_cmd *slot =3D &cmds[idx++]; + + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, SM4_CMD_INIT); + slot->hwc.sm4.cmd_init.key =3D key_ref; + slot->hwc.sm4.cmd_init.iv =3D 0; + slot->hwc.sm4.cmd_init.keylen =3D keylen; + slot->hwc.sm4.cmd_init.ivlen =3D 0; + slot->hwc.sm4.cmd_init.mode =3D tctx->sm4_mode; + slot->hwc.sm4.cmd_init.op =3D SM4_OP_ENCRYPT; + slot->hwc.sm4.cmd_init.aadlen =3D rctx->total_len; + slot->hwc.sm4.cmd_init.iolen =3D 0; + } + + /* AAD_FINAL: send data through the AAD path */ + if (rctx->total_len > 0) { + struct vcq_cmd *slot =3D &cmds[idx++]; + + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, SM4_CMD_AAD_FINAL); + slot->hwc.sm4.cmd_aad_final.data =3D (u64)rctx->in_dma; + slot->hwc.sm4.cmd_aad_final.datalen =3D rctx->total_len; + } + + /* FINAL: tag extraction only (no data) */ + { + struct vcq_cmd *slot =3D &cmds[idx++]; + + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, SM4_CMD_FINAL); + slot->hwc.sm4.cmd_final.input =3D 0; + slot->hwc.sm4.cmd_final.output =3D 0; + slot->hwc.sm4.cmd_final.tag =3D (u64)rctx->tag_dma; + slot->hwc.sm4.cmd_final.iolen =3D 0; + slot->hwc.sm4.cmd_final.taglen =3D SM4_MAC_DIGEST_SIZE; + } + + vcq_add_flush(&cmds[idx++], core_id); + + ret =3D cmh_vcq_pack_and_submit_async(cmds, idx, rctx->packed, + CMH_SM4_MAC_MAX_PACKED, + target_mbx, + cmh_sm4_mac_complete, req, + !!(req->base.flags & + CRYPTO_TFM_REQ_MAY_BACKLOG), + cmh_tm_async_timeout_jiffies()); + if (ret =3D=3D -EBUSY) + return -EBUSY; + if (ret) { + /* + * Synchronous rejection (e.g. -EAGAIN: CMQ full, no backlog). + * Free only the per-submit transients and keep the accumulated + * chunks intact so the caller can retry the identical final(). + * If no retry comes, cra_exit reclaims the orphaned chunks; the + * per-tfm buffered-byte cap bounds how much stays pinned. + */ + if (rctx->total_len > 0 && !cmh_dma_map_error(rctx->in_dma)) + cmh_dma_unmap_single(rctx->in_dma, rctx->total_len, + DMA_TO_DEVICE); + cmh_dma_unmap_single(rctx->tag_dma, SM4_MAC_DIGEST_SIZE, + DMA_FROM_DEVICE); + kfree(rctx->tag_buf); + rctx->tag_buf =3D NULL; + kfree_sensitive(rctx->buf); + rctx->buf =3D NULL; + /* Keep chunks + total_len so the retry rebuilds the buffer. */ + return ret; + } + + return -EINPROGRESS; + +out_unmap_tag: + cmh_dma_unmap_single(rctx->tag_dma, SM4_MAC_DIGEST_SIZE, + DMA_FROM_DEVICE); +out_free_tag: + kfree(rctx->tag_buf); +out_free_buf: + kfree_sensitive(rctx->buf); + rctx->buf =3D NULL; +out_free_chunks: + cmh_sm4_mac_free_chunks(rctx, tctx); + rctx->total_len =3D 0; + return ret; +} + +/* + * ahash .export()/.import(): serialise/deserialise the software + * accumulation buffer, or the generic fallback's own state once the + * request has switched to it (oversized input or clone). + */ + +static int cmh_sm4_mac_export(struct ahash_request *req, void *out) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_sm4_mac_reqctx *rctx =3D ahash_request_ctx(req); + struct cmh_sm4_mac_export_state *state =3D out; + struct cmh_sm4_mac_chunk *chunk; + u32 offset =3D 0; + int ret; + + /* + * If more data is buffered than the flat window holds, switch to + * the software fallback so a bounded, fixed-size state can be + * exported -- making export/import (clone) work at any length. + */ + if (!rctx->switched && rctx->total_len > CMH_SM4_MAC_EXPORT_MAX) { + ret =3D cmh_sm4_mac_switch_to_fb(req); + if (ret) + return ret; + } + + /* Zero the whole state buffer so no kernel memory leaks out. */ + memset(state, 0, crypto_ahash_statesize(tfm)); + + if (rctx->switched) { + state->format =3D CMH_SM4_MAC_FMT_FB; + return crypto_ahash_export(cmh_sm4_mac_fb_req(rctx), + state->data); + } + + state->format =3D CMH_SM4_MAC_FMT_RAW; + state->total_len =3D rctx->total_len; + list_for_each_entry(chunk, &rctx->chunks, list) { + memcpy(state->data + offset, chunk->data, chunk->len); + offset +=3D chunk->len; + } + return 0; +} + +static int cmh_sm4_mac_import(struct ahash_request *req, const void *in) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_sm4_mac_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cmh_sm4_mac_reqctx *rctx =3D ahash_request_ctx(req); + const struct cmh_sm4_mac_export_state *state =3D in; + struct cmh_sm4_mac_chunk *chunk; + + /* + * Do NOT call free_chunks() here: the crypto API does not + * guarantee the request context is in a valid state before + * import(), so the list pointers may be stale or invalid. + * Re-initialize from scratch instead. Any pre-existing chunks + * are tracked on tctx->all_chunks and freed in exit_tfm. + */ + memset(rctx, 0, sizeof(*rctx)); + INIT_LIST_HEAD(&rctx->chunks); + + /* Fallback-format state: replay it into a fallback request. */ + if (state->format =3D=3D CMH_SM4_MAC_FMT_FB) { + struct ahash_request *fb_req =3D cmh_sm4_mac_fb_req(rctx); + int ret; + + ahash_request_set_tfm(fb_req, tctx->fb); + ahash_request_set_callback(fb_req, 0, NULL, NULL); + ret =3D crypto_ahash_import(fb_req, state->data); + if (ret) + return ret; + rctx->switched =3D true; + return 0; + } + + if (state->format !=3D CMH_SM4_MAC_FMT_RAW) + return -EINVAL; + + if (state->total_len > CMH_SM4_MAC_EXPORT_MAX) + return -EINVAL; + + if (state->total_len) { + chunk =3D kmalloc(sizeof(*chunk) + state->total_len, + req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC); + if (!chunk) + return -ENOMEM; + chunk->len =3D state->total_len; + memcpy(chunk->data, state->data, state->total_len); + spin_lock_bh(&tctx->chunk_lock); + list_add_tail(&chunk->list, &rctx->chunks); + list_add_tail(&chunk->tfm_node, &tctx->all_chunks); + tctx->tfm_buffered +=3D chunk->len; + spin_unlock_bh(&tctx->chunk_lock); + rctx->total_len =3D state->total_len; + } + return 0; +} + +static int cmh_sm4_mac_finup(struct ahash_request *req) +{ + int err; + + err =3D cmh_sm4_mac_update(req); + if (err) + return err; + return cmh_sm4_mac_final(req); +} + +static int cmh_sm4_mac_digest(struct ahash_request *req) +{ + int err; + + err =3D cmh_sm4_mac_init(req); + if (err) + return err; + return cmh_sm4_mac_finup(req); +} + +/* Registration */ + +static struct cmh_sm4_mac_drv sm4_mac_drv_algs[ARRAY_SIZE(sm4_mac_algs)]; + +static int cmh_sm4_mac_init_tfm(struct crypto_ahash *tfm) +{ + struct cmh_sm4_mac_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct ahash_alg *alg =3D crypto_ahash_alg(tfm); + struct cmh_sm4_mac_drv *drv =3D + container_of(alg, struct cmh_sm4_mac_drv, alg); + struct crypto_ahash *fb; + + memset(tctx, 0, sizeof(*tctx)); + tctx->sm4_mode =3D drv->info->sm4_mode; + spin_lock_init(&tctx->chunk_lock); + INIT_LIST_HEAD(&tctx->all_chunks); + + /* Allocate SW cipher for the CMAC/XCBC empty-input fallback */ + if (tctx->sm4_mode =3D=3D SM4_MODE_XCBC || + tctx->sm4_mode =3D=3D SM4_MODE_CMAC) { + struct crypto_cipher *ci; + + ci =3D crypto_alloc_cipher("sm4", 0, 0); + if (IS_ERR(ci)) + return PTR_ERR(ci); + tctx->sw_cipher =3D ci; + } + + /* + * Generic software fallback for oversized input / clone. The + * generic cmac(sm4)/xcbc(sm4) has no core export/import state, so it + * cannot serve as a CRYPTO_ALG_NEED_FALLBACK fallback -- allocate it + * explicitly (masking out CRYPTO_ALG_ASYNC to exclude this driver). + */ + fb =3D crypto_alloc_ahash(crypto_ahash_alg_name(tfm), 0, + CRYPTO_ALG_ASYNC); + if (IS_ERR(fb)) { + if (tctx->sw_cipher) + crypto_free_cipher(tctx->sw_cipher); + tctx->sw_cipher =3D NULL; + return PTR_ERR(fb); + } + tctx->fb =3D fb; + + crypto_ahash_set_reqsize(tfm, + sizeof(struct cmh_sm4_mac_reqctx) + + crypto_tfm_ctx_alignment() + + sizeof(struct ahash_request) + + crypto_ahash_reqsize(fb)); + return 0; +} + +static void cmh_sm4_mac_exit_tfm(struct crypto_ahash *tfm) +{ + struct cmh_sm4_mac_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cmh_sm4_mac_chunk *c, *tmp; + + /* Free any orphaned chunks (e.g. testmgr export/reimport poison) */ + spin_lock_bh(&tctx->chunk_lock); + list_for_each_entry_safe(c, tmp, &tctx->all_chunks, tfm_node) { + list_del(&c->tfm_node); + tctx->tfm_buffered -=3D c->len; + kfree_sensitive(c); + } + spin_unlock_bh(&tctx->chunk_lock); + + if (tctx->sw_cipher) + crypto_free_cipher(tctx->sw_cipher); + if (tctx->fb) + crypto_free_ahash(tctx->fb); + memzero_explicit(tctx->xcbc_k1, sizeof(tctx->xcbc_k1)); + memzero_explicit(tctx->xcbc_k3, sizeof(tctx->xcbc_k3)); + memzero_explicit(tctx->cmac_k2, sizeof(tctx->cmac_k2)); + cmh_key_destroy(&tctx->key); +} + +/** + * cmh_sm4_cmac_register() - Register SM4-CMAC/XCBC hash algorithms with t= he crypto framework + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_sm4_cmac_register(void) +{ + unsigned int i; + int ret; + + if (!cmh_core_present(CMH_CORE_SM4)) + return 0; + + for (i =3D 0; i < ARRAY_SIZE(sm4_mac_algs); i++) { + const struct cmh_sm4_mac_alg_info *info =3D &sm4_mac_algs[i]; + struct cmh_sm4_mac_drv *drv =3D &sm4_mac_drv_algs[i]; + struct ahash_alg *alg =3D &drv->alg; + + drv->info =3D info; + + memset(alg, 0, sizeof(*alg)); + + alg->init =3D cmh_sm4_mac_init; + alg->update =3D cmh_sm4_mac_update; + alg->final =3D cmh_sm4_mac_final; + alg->finup =3D cmh_sm4_mac_finup; + alg->digest =3D cmh_sm4_mac_digest; + alg->export =3D cmh_sm4_mac_export; + alg->import =3D cmh_sm4_mac_import; + alg->setkey =3D cmh_sm4_mac_setkey; + alg->init_tfm =3D cmh_sm4_mac_init_tfm; + alg->exit_tfm =3D cmh_sm4_mac_exit_tfm; + + alg->halg.digestsize =3D SM4_MAC_DIGEST_SIZE; + alg->halg.statesize =3D CMH_SM4_MAC_STATE_SIZE; + + strscpy(alg->halg.base.cra_name, info->alg_name, + CRYPTO_MAX_ALG_NAME); + strscpy(alg->halg.base.cra_driver_name, info->drv_name, + CRYPTO_MAX_ALG_NAME); + alg->halg.base.cra_priority =3D 300; + alg->halg.base.cra_flags =3D CRYPTO_ALG_KERN_DRIVER_ONLY | + CRYPTO_ALG_NO_FALLBACK | + CRYPTO_ALG_ASYNC | + CRYPTO_ALG_REQ_VIRT; + alg->halg.base.cra_blocksize =3D SM4_MAC_BLOCK_SIZE; + alg->halg.base.cra_ctxsize =3D sizeof(struct cmh_sm4_mac_tfm_ctx); + alg->halg.base.cra_module =3D THIS_MODULE; + + ret =3D crypto_register_ahash(alg); + if (ret) { + dev_err(cmh_dev(), "cmh_sm4_mac: failed to register %s (rc=3D%d)\n", + info->alg_name, ret); + goto err_unregister; + } + + dev_dbg(cmh_dev(), "cmh_sm4_mac: registered %s\n", + info->alg_name); + } + + return 0; + +err_unregister: + while (i--) + crypto_unregister_ahash(&sm4_mac_drv_algs[i].alg); + return ret; +} + +/** + * cmh_sm4_cmac_unregister() - Unregister SM4 MAC hash algorithms from the= crypto framework + */ +void cmh_sm4_cmac_unregister(void) +{ + unsigned int i; + + if (!cmh_core_present(CMH_CORE_SM4)) + return; + + for (i =3D 0; i < ARRAY_SIZE(sm4_mac_algs); i++) { + crypto_unregister_ahash(&sm4_mac_drv_algs[i].alg); + dev_dbg(cmh_dev(), "cmh_sm4_mac: unregistered %s\n", + sm4_mac_algs[i].alg_name); + } +} diff --git a/drivers/crypto/cmh/cmh_sm4_skcipher.c b/drivers/crypto/cmh/cmh= _sm4_skcipher.c new file mode 100644 index 000000000000..3063eff7857f --- /dev/null +++ b/drivers/crypto/cmh/cmh_sm4_skcipher.c @@ -0,0 +1,709 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Kernel Crypto API SM4 (skcipher) Driver + * + * Registers skcipher algorithms with the Linux crypto subsystem: + * ecb(sm4), cbc(sm4), ctr(sm4), cfb(sm4), xts(sm4) + * + * Uses the CMH SM4 Core via VCQ commands: + * [SYS_CMD_WRITE] + SM4_CMD_INIT + SM4_CMD_FINAL + VCQ_CMD_FLUSH + * + * The SM4 core requires bidirectional DMA -- both input and output + * buffers are mapped and passed in a single SM4_CMD_FINAL command. + * + * Raw-key atomicity: SYS_CMD_WRITE to SYS_REF_TEMP is packed into + * the same VCQ as SM4 commands (see cmh_key.h for details). + * + */ + +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "cmh_sm4.h" +#include "cmh_vcq.h" +#include "cmh_sm4_abi.h" +#include "cmh_sys_abi.h" +#include "cmh_sys.h" +#include "cmh_txn.h" +#include "cmh_dma.h" +#include "cmh_key.h" + +/* Algorithm Table */ + +struct cmh_sm4_alg_info { + u32 sm4_mode; /* SM4_MODE_* */ + u32 ivsize; /* bytes (0 for ECB) */ + u32 min_keysize; + u32 max_keysize; + const char *alg_name; /* Linux crypto name: "ecb(sm4)" */ + const char *drv_name; /* driver name: "rambus-cmh-ecb-sm4" */ +}; + +static const struct cmh_sm4_alg_info sm4_algs[] =3D { + { SM4_MODE_ECB, 0, CMH_SM4_KEY_SIZE, CMH_SM4_KEY_SIZE, + "ecb(sm4)", "rambus-cmh-ecb-sm4" }, + { SM4_MODE_CBC, CMH_SM4_IV_SIZE, CMH_SM4_KEY_SIZE, CMH_SM4_KEY_SIZE, + "cbc(sm4)", "rambus-cmh-cbc-sm4" }, + { SM4_MODE_CTR, CMH_SM4_IV_SIZE, CMH_SM4_KEY_SIZE, CMH_SM4_KEY_SIZE, + "ctr(sm4)", "rambus-cmh-ctr-sm4" }, + { SM4_MODE_CFB, CMH_SM4_IV_SIZE, CMH_SM4_KEY_SIZE, CMH_SM4_KEY_SIZE, + "cfb(sm4)", "rambus-cmh-cfb-sm4" }, + { SM4_MODE_XTS, CMH_SM4_IV_SIZE, CMH_SM4_KEY_SIZE * 2, + CMH_SM4_KEY_SIZE * 2, + "xts(sm4)", "rambus-cmh-xts-sm4" }, +}; + +/* Per-transform context (allocated by crypto framework) */ + +struct cmh_sm4_tfm_ctx { + struct cmh_key_ctx key; +}; + +/* Per-request context (lives in skcipher_request::__ctx) */ + +/* + * Maximum payload commands: + * [SYS_CMD_WRITE] + SM4_CMD_INIT + [SM4_CMD_UPDATE] + SM4_CMD_FINAL + * + VCQ_CMD_FLUSH =3D 5 + * UPDATE is used for XTS data > 2 blocks (see cmh_sm4_crypt). + */ +#define CMH_SM4_MAX_PAYLOAD 5 +#define CMH_SM4_MAX_PACKED (CMH_SM4_MAX_PAYLOAD * 2) + +struct cmh_sm4_reqctx { + dma_addr_t in_dma; + dma_addr_t out_dma; + dma_addr_t iv_dma; + dma_addr_t iv2_dma; + dma_addr_t key_dma; + u8 *in_buf; + u8 *out_buf; + u8 *iv_buf; + u8 *iv2_buf; + u32 cryptlen; + u32 ivsize; + u32 keylen; + u32 sm4_mode; + u32 sm4_op; + /* CTR counter-wrap split state */ + u32 ctr_chunk1_len; + u32 core_id; + s32 target_mbx; + u64 key_ref; + struct vcq_cmd packed[CMH_SM4_MAX_PACKED]; +}; + +/* VCQ Builders -- SM4-specific */ + +static void vcq_add_sm4_init(struct vcq_cmd *slot, u32 core_id, u64 key_re= f, u64 iv_dma, + u32 keylen, u32 ivlen, u32 mode, u32 op, + u32 iolen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, SM4_CMD_INIT); + slot->hwc.sm4.cmd_init.key =3D key_ref; + slot->hwc.sm4.cmd_init.iv =3D iv_dma; + slot->hwc.sm4.cmd_init.keylen =3D keylen; + slot->hwc.sm4.cmd_init.ivlen =3D ivlen; + slot->hwc.sm4.cmd_init.mode =3D mode; + slot->hwc.sm4.cmd_init.op =3D op; + slot->hwc.sm4.cmd_init.aadlen =3D 0; + slot->hwc.sm4.cmd_init.iolen =3D iolen; +} + +static void vcq_add_sm4_update(struct vcq_cmd *slot, u32 core_id, u64 inpu= t_dma, + u64 output_dma, u32 iolen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, SM4_CMD_UPDATE); + slot->hwc.sm4.cmd_update.input =3D input_dma; + slot->hwc.sm4.cmd_update.output =3D output_dma; + slot->hwc.sm4.cmd_update.iolen =3D iolen; +} + +static void vcq_add_sm4_final(struct vcq_cmd *slot, u32 core_id, u64 input= _dma, + u64 output_dma, u32 iolen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, SM4_CMD_FINAL); + slot->hwc.sm4.cmd_final.input =3D input_dma; + slot->hwc.sm4.cmd_final.output =3D output_dma; + slot->hwc.sm4.cmd_final.iolen =3D iolen; + slot->hwc.sm4.cmd_final.tag =3D 0; + slot->hwc.sm4.cmd_final.taglen =3D 0; +} + +/* + * We wrap each skcipher_alg with its info pointer in a compound struct, + * then use container_of() in cmh_sm4_get_info() to recover it. + */ +struct cmh_sm4_alg_drv { + struct skcipher_alg alg; + const struct cmh_sm4_alg_info *info; +}; + +static bool sm4_is_stream_mode(u32 mode) +{ + return mode =3D=3D SM4_MODE_CTR || mode =3D=3D SM4_MODE_CFB; +} + +/* + * Update req->iv after a successful encrypt/decrypt. + * Same semantics as cmh_aes_update_iv -- see cmh_aes.c. + */ +static void cmh_sm4_update_iv(struct skcipher_request *req, u32 mode, + u32 op, const u8 *in_buf, const u8 *out_buf) +{ + u32 bs =3D CMH_SM4_BLOCK_SIZE; + u32 nblocks; + + switch (mode) { + case SM4_MODE_CBC: + if (op =3D=3D SM4_OP_ENCRYPT) + memcpy(req->iv, out_buf + req->cryptlen - bs, bs); + else + memcpy(req->iv, in_buf + req->cryptlen - bs, bs); + break; + case SM4_MODE_CTR: + /* Arithmetic big-endian 128-bit counter increment */ + nblocks =3D DIV_ROUND_UP(req->cryptlen, bs); + { + u8 *iv =3D req->iv; + int i; + + for (i =3D bs - 1; i >=3D 0 && nblocks; i--) { + u32 sum =3D (u32)iv[i] + (nblocks & 0xff); + + iv[i] =3D (u8)sum; + nblocks =3D (nblocks >> 8) + (sum >> 8); + } + } + break; + case SM4_MODE_CFB: + /* + * For sub-block requests (cryptlen < 16), there is no + * complete ciphertext block to chain, so the IV is left + * unchanged -- CFB-128 has no defined chaining semantic + * for partial blocks (shift-register CFB-n is a different + * mode). Without this guard the pointer arithmetic + * underflows and reads before the buffer. + */ + if (req->cryptlen >=3D bs) { + if (op =3D=3D SM4_OP_ENCRYPT) + memcpy(req->iv, out_buf + req->cryptlen - bs, + bs); + else + memcpy(req->iv, in_buf + req->cryptlen - bs, + bs); + } + break; + default: + break; + } +} + +/* skcipher Operations */ + +static const struct cmh_sm4_alg_info * +cmh_sm4_get_info(struct crypto_skcipher *tfm) +{ + struct skcipher_alg *alg =3D crypto_skcipher_alg(tfm); + + return container_of(alg, struct cmh_sm4_alg_drv, alg)->info; +} + +static int cmh_sm4_setkey(struct crypto_skcipher *tfm, const u8 *key, + unsigned int keylen) +{ + struct cmh_sm4_tfm_ctx *tctx =3D crypto_skcipher_ctx(tfm); + const struct cmh_sm4_alg_info *info =3D cmh_sm4_get_info(tfm); + + if (info->sm4_mode =3D=3D SM4_MODE_XTS) { + int err; + + /* XTS: double key (32 bytes) */ + if (keylen !=3D CMH_SM4_KEY_SIZE * 2) + return -EINVAL; + err =3D xts_verify_key(tfm, key, keylen); + if (err) + return err; + } else { + /* SM4 always uses 128-bit (16-byte) keys */ + if (keylen !=3D CMH_SM4_KEY_SIZE) + return -EINVAL; + } + + return cmh_key_setkey_raw(&tctx->key, key, keylen, CORE_ID_SM4); +} + +static int cmh_sm4_init_tfm(struct crypto_skcipher *tfm) +{ + struct cmh_sm4_tfm_ctx *tctx =3D crypto_skcipher_ctx(tfm); + + memset(tctx, 0, sizeof(*tctx)); + crypto_skcipher_set_reqsize(tfm, sizeof(struct cmh_sm4_reqctx)); + return 0; +} + +static void cmh_sm4_exit_tfm(struct crypto_skcipher *tfm) +{ + struct cmh_sm4_tfm_ctx *tctx =3D crypto_skcipher_ctx(tfm); + + cmh_key_destroy(&tctx->key); +} + +#define CMH_SM4_MAX_CRYPTLEN SZ_32M + +/* DMA unmap helper */ +static void cmh_sm4_unmap_dma(struct cmh_sm4_reqctx *rctx) +{ + if (rctx->iv2_buf) + cmh_dma_unmap_single(rctx->iv2_dma, rctx->ivsize, + DMA_TO_DEVICE); + if (rctx->ivsize > 0) + cmh_dma_unmap_single(rctx->iv_dma, rctx->ivsize, + DMA_TO_DEVICE); + cmh_dma_unmap_single(rctx->out_dma, rctx->cryptlen, DMA_FROM_DEVICE); + cmh_dma_unmap_single(rctx->in_dma, rctx->cryptlen, DMA_TO_DEVICE); +} + +static void cmh_sm4_free_bufs(struct cmh_sm4_reqctx *rctx) +{ + kfree(rctx->iv2_buf); + rctx->iv2_buf =3D NULL; + kfree(rctx->iv_buf); + rctx->iv_buf =3D NULL; + kfree_sensitive(rctx->out_buf); + rctx->out_buf =3D NULL; + kfree_sensitive(rctx->in_buf); + rctx->in_buf =3D NULL; +} + +/* + * Submit the second CTR chunk after the first completes. + * Called from cmh_sm4_complete when ctr_chunk1_len > 0. + */ +static int cmh_sm4_ctr_submit_chunk2(struct skcipher_request *req); + +static void cmh_sm4_complete(void *data, int error) +{ + struct skcipher_request *req =3D data; + struct cmh_sm4_reqctx *rctx =3D skcipher_request_ctx(req); + + if (error =3D=3D -EINPROGRESS) { + cmh_complete(&req->base, error); + return; + } + + /* + * CTR counter-wrap: first chunk completed, submit second. + * DMA mappings remain valid (they cover the full buffer). + * + * Recursion depth bounded: chunk2 clears ctr_chunk1_len before + * submission, so the second cmh_sm4_complete invocation sees 0 + * and finalizes (max depth =3D 2). + */ + if (rctx->ctr_chunk1_len && !error) { + int ret =3D cmh_sm4_ctr_submit_chunk2(req); + + if (!ret || ret =3D=3D -EBUSY) + return; + /* Submission failed; clean up below */ + error =3D ret; + } + + cmh_sm4_unmap_dma(rctx); + + if (!error) { + scatterwalk_map_and_copy(rctx->out_buf, req->dst, + 0, rctx->cryptlen, 1); + cmh_sm4_update_iv(req, rctx->sm4_mode, rctx->sm4_op, + rctx->in_buf, rctx->out_buf); + } + + cmh_sm4_free_bufs(rctx); + cmh_complete(&req->base, error); +} + +static int cmh_sm4_ctr_submit_chunk2(struct skcipher_request *req) +{ + struct crypto_skcipher *tfm =3D crypto_skcipher_reqtfm(req); + struct cmh_sm4_tfm_ctx *tctx =3D crypto_skcipher_ctx(tfm); + struct cmh_sm4_reqctx *rctx =3D skcipher_request_ctx(req); + struct vcq_cmd cmds[CMH_SM4_MAX_PAYLOAD]; + u32 chunk1 =3D rctx->ctr_chunk1_len; + u32 chunk2 =3D rctx->cryptlen - chunk1; + u64 key_ref; + u32 keylen; + u32 idx =3D 0; + + /* Clear split flag so next completion is final */ + rctx->ctr_chunk1_len =3D 0; + + vcq_add_sys_write(&cmds[idx++], SYS_REF_TEMP, + (u64)rctx->key_dma, SYS_REF_NONE, + tctx->key.raw.len, + tctx->key.raw.sys_type); + key_ref =3D SYS_REF_TEMP; + keylen =3D tctx->key.raw.len; + + vcq_add_sm4_init(&cmds[idx++], rctx->core_id, key_ref, + (u64)rctx->iv2_dma, keylen, rctx->ivsize, + rctx->sm4_mode, rctx->sm4_op, chunk2); + vcq_add_sm4_final(&cmds[idx++], rctx->core_id, + (u64)(rctx->in_dma + chunk1), + (u64)(rctx->out_dma + chunk1), chunk2); + vcq_add_flush(&cmds[idx++], rctx->core_id); + + return cmh_vcq_pack_and_submit_async(cmds, idx, rctx->packed, + CMH_SM4_MAX_PACKED, + rctx->target_mbx, + cmh_sm4_complete, req, + !!(req->base.flags & + CRYPTO_TFM_REQ_MAY_BACKLOG), + cmh_tm_async_timeout_jiffies()); +} + +static int cmh_sm4_crypt(struct skcipher_request *req, u32 sm4_op) +{ + struct crypto_skcipher *tfm =3D crypto_skcipher_reqtfm(req); + struct cmh_sm4_tfm_ctx *tctx =3D crypto_skcipher_ctx(tfm); + const struct cmh_sm4_alg_info *info =3D cmh_sm4_get_info(tfm); + struct cmh_sm4_reqctx *rctx =3D skcipher_request_ctx(req); + struct vcq_cmd cmds[CMH_SM4_MAX_PAYLOAD]; + u64 key_ref; + u32 keylen; + struct core_dispatch d; + s32 target_mbx; + u32 core_id; + u32 idx; + int ret; + gfp_t gfp; + + if (tctx->key.mode =3D=3D CMH_KEY_NONE) + return -ENOKEY; + + if (!req->cryptlen) + return 0; + + if (req->cryptlen > CMH_SM4_MAX_CRYPTLEN) + return -EINVAL; + + switch (info->sm4_mode) { + case SM4_MODE_CTR: + case SM4_MODE_CFB: + break; + case SM4_MODE_XTS: + if (req->cryptlen < CMH_SM4_BLOCK_SIZE) + return -EINVAL; + break; + default: + if (req->cryptlen & (CMH_SM4_BLOCK_SIZE - 1)) + return -EINVAL; + break; + } + + gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + + memset(rctx, 0, sizeof(*rctx)); + rctx->cryptlen =3D req->cryptlen; + rctx->ivsize =3D info->ivsize; + rctx->sm4_mode =3D info->sm4_mode; + rctx->sm4_op =3D sm4_op; + rctx->iv2_buf =3D NULL; + + /* + * cryptlen is user-controlled up to CMH_SM4_MAX_CRYPTLEN (well above + * KMALLOC_MAX_SIZE), so use __GFP_NOWARN: an oversized request fails + * cleanly with -ENOMEM instead of splatting the page allocator. + */ + rctx->in_buf =3D kmalloc(req->cryptlen, gfp | __GFP_NOWARN); + if (!rctx->in_buf) + return -ENOMEM; + + scatterwalk_map_and_copy(rctx->in_buf, req->src, 0, req->cryptlen, 0); + + rctx->in_dma =3D cmh_dma_map_single(rctx->in_buf, req->cryptlen, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->in_dma)) { + ret =3D -ENOMEM; + goto out_free_in; + } + + rctx->out_buf =3D kmalloc(req->cryptlen, gfp | __GFP_NOWARN); + if (!rctx->out_buf) { + ret =3D -ENOMEM; + goto out_unmap_in; + } + + rctx->out_dma =3D cmh_dma_map_single(rctx->out_buf, req->cryptlen, + DMA_FROM_DEVICE); + if (cmh_dma_map_error(rctx->out_dma)) { + ret =3D -ENOMEM; + goto out_free_out; + } + + if (info->ivsize > 0) { + rctx->iv_buf =3D kmemdup(req->iv, info->ivsize, gfp); + if (!rctx->iv_buf) { + ret =3D -ENOMEM; + goto out_unmap_out; + } + rctx->iv_dma =3D cmh_dma_map_single(rctx->iv_buf, info->ivsize, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->iv_dma)) { + ret =3D -ENOMEM; + goto out_free_iv; + } + } + + idx =3D 0; + + rctx->key_dma =3D tctx->key.raw.dma; + rctx->keylen =3D tctx->key.raw.len; + vcq_add_sys_write(&cmds[idx++], SYS_REF_TEMP, + (u64)rctx->key_dma, SYS_REF_NONE, + tctx->key.raw.len, + tctx->key.raw.sys_type); + key_ref =3D SYS_REF_TEMP; + keylen =3D tctx->key.raw.len; + d =3D cmh_core_select_instance(CMH_CORE_SM4); + target_mbx =3D d.mbx_idx; + core_id =3D d.core_id; + + /* + * iolen in INIT: passed for all modes. The EIP-40 eSW ignores + * it for CTR (stream cipher), but uses it for XTS/CBC/ECB to + * know the total data length. Pass cryptlen unconditionally. + */ + vcq_add_sm4_init(&cmds[idx++], core_id, key_ref, (u64)rctx->iv_dma, + keylen, info->ivsize, info->sm4_mode, sm4_op, + req->cryptlen); + + if (info->sm4_mode =3D=3D SM4_MODE_XTS && + req->cryptlen > 2 * CMH_SM4_BLOCK_SIZE) { + u32 final_len, update_len; + + if (req->cryptlen & (CMH_SM4_BLOCK_SIZE - 1)) + final_len =3D CMH_SM4_BLOCK_SIZE + + (req->cryptlen & (CMH_SM4_BLOCK_SIZE - 1)); + else + final_len =3D 2 * CMH_SM4_BLOCK_SIZE; + + update_len =3D req->cryptlen - final_len; + + vcq_add_sm4_update(&cmds[idx++], core_id, + (u64)rctx->in_dma, + (u64)rctx->out_dma, update_len); + vcq_add_sm4_final(&cmds[idx++], core_id, + (u64)(rctx->in_dma + update_len), + (u64)(rctx->out_dma + update_len), + final_len); + } else if (info->sm4_mode =3D=3D SM4_MODE_CTR) { + /* + * CTR counter-wrap: split at the 64-bit boundary, + * consistent with the AES-SCA driver. The completion + * callback submits chunk2 with IV =3D {upper64+1, 0}. + */ + u64 lower64 =3D get_unaligned_be64(rctx->iv_buf + 8); + u32 nblocks =3D DIV_ROUND_UP(req->cryptlen, + CMH_SM4_BLOCK_SIZE); + u64 bwrap =3D lower64 ? (~lower64 + 1ULL) : U64_MAX; + + if (nblocks > bwrap) { + u32 chunk1 =3D (u32)bwrap * CMH_SM4_BLOCK_SIZE; + u64 upper64; + + /* Prepare second IV for chained submission */ + rctx->iv2_buf =3D kmalloc(info->ivsize, gfp); + if (!rctx->iv2_buf) { + ret =3D -ENOMEM; + goto out_unmap_iv; + } + upper64 =3D get_unaligned_be64(rctx->iv_buf); + put_unaligned_be64(upper64 + 1, rctx->iv2_buf); + put_unaligned_be64(0, rctx->iv2_buf + 8); + + rctx->iv2_dma =3D + cmh_dma_map_single(rctx->iv2_buf, + info->ivsize, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->iv2_dma)) { + ret =3D -ENOMEM; + goto out_free_iv2; + } + + /* Store state for the chained second submission */ + rctx->ctr_chunk1_len =3D chunk1; + rctx->core_id =3D core_id; + rctx->target_mbx =3D target_mbx; + rctx->key_ref =3D key_ref; + + /* First transaction: only chunk1 */ + vcq_add_sm4_final(&cmds[idx++], core_id, + (u64)rctx->in_dma, + (u64)rctx->out_dma, chunk1); + } else { + /* No wrap: single FINAL with all data */ + vcq_add_sm4_final(&cmds[idx++], core_id, + (u64)rctx->in_dma, + (u64)rctx->out_dma, + req->cryptlen); + } + } else { + vcq_add_sm4_final(&cmds[idx++], core_id, + (u64)rctx->in_dma, + (u64)rctx->out_dma, req->cryptlen); + } + + vcq_add_flush(&cmds[idx++], core_id); + + ret =3D cmh_vcq_pack_and_submit_async(cmds, idx, rctx->packed, + CMH_SM4_MAX_PACKED, target_mbx, + cmh_sm4_complete, req, + !!(req->base.flags & + CRYPTO_TFM_REQ_MAY_BACKLOG), + cmh_tm_async_timeout_jiffies()); + if (ret =3D=3D -EBUSY) + return -EBUSY; + if (ret) + goto out_cleanup_all; + + return -EINPROGRESS; + +out_cleanup_all: + if (rctx->iv2_buf) { + cmh_dma_unmap_single(rctx->iv2_dma, info->ivsize, + DMA_TO_DEVICE); + } +out_free_iv2: + kfree(rctx->iv2_buf); +out_unmap_iv: + if (info->ivsize > 0) + cmh_dma_unmap_single(rctx->iv_dma, info->ivsize, + DMA_TO_DEVICE); +out_free_iv: + kfree(rctx->iv_buf); +out_unmap_out: + cmh_dma_unmap_single(rctx->out_dma, req->cryptlen, DMA_FROM_DEVICE); +out_free_out: + kfree_sensitive(rctx->out_buf); +out_unmap_in: + cmh_dma_unmap_single(rctx->in_dma, req->cryptlen, DMA_TO_DEVICE); +out_free_in: + kfree_sensitive(rctx->in_buf); + return ret; +} + +static int cmh_sm4_encrypt(struct skcipher_request *req) +{ + return cmh_sm4_crypt(req, SM4_OP_ENCRYPT); +} + +static int cmh_sm4_decrypt(struct skcipher_request *req) +{ + return cmh_sm4_crypt(req, SM4_OP_DECRYPT); +} + +/* Registration */ + +static struct cmh_sm4_alg_drv sm4_drv_algs[ARRAY_SIZE(sm4_algs)]; + +/** + * cmh_sm4_register() - Register SM4-CBC/CTR/ECB/XTS skcipher algorithms + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_sm4_register(void) +{ + unsigned int i; + int ret; + + if (!cmh_core_present(CMH_CORE_SM4)) + return 0; + + for (i =3D 0; i < ARRAY_SIZE(sm4_algs); i++) { + const struct cmh_sm4_alg_info *info =3D &sm4_algs[i]; + struct cmh_sm4_alg_drv *drv =3D &sm4_drv_algs[i]; + struct skcipher_alg *alg =3D &drv->alg; + + drv->info =3D info; + + memset(alg, 0, sizeof(*alg)); + + alg->setkey =3D cmh_sm4_setkey; + alg->encrypt =3D cmh_sm4_encrypt; + alg->decrypt =3D cmh_sm4_decrypt; + alg->init =3D cmh_sm4_init_tfm; + alg->exit =3D cmh_sm4_exit_tfm; + alg->min_keysize =3D info->min_keysize; + alg->max_keysize =3D info->max_keysize; + alg->ivsize =3D info->ivsize; + + strscpy(alg->base.cra_name, info->alg_name, + CRYPTO_MAX_ALG_NAME); + strscpy(alg->base.cra_driver_name, info->drv_name, + CRYPTO_MAX_ALG_NAME); + alg->base.cra_priority =3D 300; + alg->base.cra_flags =3D CRYPTO_ALG_KERN_DRIVER_ONLY | + CRYPTO_ALG_ASYNC; + alg->base.cra_blocksize =3D sm4_is_stream_mode(info->sm4_mode) + ? 1 : CMH_SM4_BLOCK_SIZE; + /* + * Stream modes (CTR/CFB) use cra_blocksize=3D1 but consume a + * full cipher block per counter/feedback step; declare + * chunksize so the crypto layer buffers sub-block data instead + * of splitting the keystream across requests. + */ + if (sm4_is_stream_mode(info->sm4_mode)) + alg->chunksize =3D CMH_SM4_BLOCK_SIZE; + alg->base.cra_ctxsize =3D sizeof(struct cmh_sm4_tfm_ctx); + alg->base.cra_module =3D THIS_MODULE; + + ret =3D crypto_register_skcipher(alg); + if (ret) { + dev_err(cmh_dev(), "cmh_sm4: failed to register %s (rc=3D%d)\n", + info->alg_name, ret); + goto err_unregister; + } + + dev_dbg(cmh_dev(), "cmh_sm4: registered %s\n", info->alg_name); + } + + return 0; + +err_unregister: + while (i--) + crypto_unregister_skcipher(&sm4_drv_algs[i].alg); + return ret; +} + +/** + * cmh_sm4_unregister() - Unregister SM4 skcipher algorithms from the cryp= to framework + */ +void cmh_sm4_unregister(void) +{ + unsigned int i; + + if (!cmh_core_present(CMH_CORE_SM4)) + return; + + for (i =3D 0; i < ARRAY_SIZE(sm4_algs); i++) { + crypto_unregister_skcipher(&sm4_drv_algs[i].alg); + dev_dbg(cmh_dev(), "cmh_sm4: unregistered %s\n", sm4_algs[i].alg_name); + } +} diff --git a/drivers/crypto/cmh/include/cmh_sm4.h b/drivers/crypto/cmh/incl= ude/cmh_sm4.h new file mode 100644 index 000000000000..9f4b0fb918db --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_sm4.h @@ -0,0 +1,24 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- SM4 Crypto API Drivers + * + * Registers SM4 algorithms with the Linux crypto subsystem: + * skcipher: ecb/cbc/ctr/cfb/xts(sm4) + * aead: gcm/ccm(sm4) + * shash: cmac/xcbc(sm4) + */ + +#ifndef CMH_SM4_H +#define CMH_SM4_H + +int cmh_sm4_register(void); +void cmh_sm4_unregister(void); + +int cmh_sm4_aead_register(void); +void cmh_sm4_aead_unregister(void); + +int cmh_sm4_cmac_register(void); +void cmh_sm4_cmac_unregister(void); + +#endif /* CMH_SM4_H */ --=20 2.43.7 From nobody Fri Sep 25 01:20:34 2026 Received: from BL2PR02CU003.outbound.protection.outlook.com (mail-eastusazon11021134.outbound.protection.outlook.com [52.101.52.134]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CA0D24BCAA9; Thu, 17 Sep 2026 22:59:46 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.52.134 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789685996; cv=fail; b=XIT2UIdSKNoumpkozRmDmvnI8TZhEG0pT1GAP8Cqm7IkIhg27stig05xo+a0SngH86e/CjSzVXNZ6ovViSxTPouXf7R15ys2HrsvY0iePZxuMj9t64YcB7GcouQbwWVY8z5W3wSuTndqX7idia0YylZGl9k91YKg3BTWYzi0Gr4= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789685996; c=relaxed/simple; bh=xYixultXaWu171fW16v/Dn+U6bRb7RvugdYqVh72Y+c=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=iFFYBuWZqJO3lMrtkU/X+UIDY7stiU901ARF+bpQ3enDZWHkqa0FdD0jaIr4bd/AIZTlP9khXi2eVL/12U3ImYUUspAM0yUjZ8fsi09naV8k9rB68pN9Yv9e7h3fWFnlqG1cEYmT7fT1hkO0JmL2k/uSX1dKXwUwIYn0MpBckDY= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com; spf=fail smtp.mailfrom=rambus.com; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b=PGI3PiF7; arc=fail smtp.client-ip=52.101.52.134 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=rambus.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b="PGI3PiF7" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=wvWZtxQIJldCzCL1xanQgmiH2KcdG9n4o8wDrAtziIKUbGATWi6MD7IQSaYywu5tmkAbe4xLUEuD1qe0XgQezjrf0YbpzkgeVVl6XyDFT5TiaulmcVelWWRJDS1JONDgb7u3c0b1RsUwjlOA6GIt6tolPTcnvtt6vHMiV7I+EYFBcDf0PleWyKISui/wMOUW0wlZubOpQzbCYncqsDeAs2EZx916fM0Z44juGUscHQcvXjOxubobUDJcb0RllEl9hjtKiFXVdMhGNBP10ikLJMP0ZCf1e/KU00XwZ30uf0nVWPDC4t3GAbyASR30EodeOFbnBJG4DieAYscPcSCxcg== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=w9fhIszzESWxv7kKanVzMrRsB7KClZmKn6+SPVzFQds=; b=L3rL2UERPEmjOg8ktTCi7u6H9+KSudEZy6MFr0l4oR6OfKV+PwxqWt1Pzt4xLIyfbr9w0WqHCAWRjb0VLKg1C3qwf0OxXqDv0PeN93uPH2aYP1OPDeOE8tqE4PqXBaGfT3ZFlhO6cxWcztji18dfZ70Svv/FUjIVfFZtEbbFckolk0aTR3QVh1Rq9PTCwvANgmuejQjRF2bmBiaHA7NPdSkKgmge866yCMcfFc9IXKO42seNX9hDkKMkc8ZZhn7W6vMo8scaJ+Hpm6E2iemsMOUq/bqYw5znDRIudCuP6VLtSCwlRdzfkeaf5LSluW7lY17YKTbXhgLlJW8PdcbNEw== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 192.86.86.210) smtp.rcpttodomain=cryptography.com smtp.mailfrom=rambus.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=rambus.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=rambus.com; s=selector1; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=w9fhIszzESWxv7kKanVzMrRsB7KClZmKn6+SPVzFQds=; b=PGI3PiF7baaWkglKJ2zTZslAJIi29QXXFbym1uccYPqRC57PhN0h27RavBMJ+PX2T6UvkLRZU4+nbmSmzJ0Qsuj0/UsU8MCBqahIocEl3aAUntcEMofs4rLkBJKncyhLH/ZhKj1DgxKL1xo21/E35RqwQMLFzPeHXeuiXKd2D+fFKt3OaHymXlymhSkF41xu37j4a2nu1fR7KqVd1hBFP8BrdOtn4ApVor16UCOqKVJC/cagqD52P+kFg/ynl1tuIQUJ39GTs2B24lVze0VEIAyhItGT7d3vrmRWg/O2uJlMkxfoNQkTn6lBysMiLar0aOYwqin86F9gWlXWDwzwXA== Received: from BN0PR10CA0020.namprd10.prod.outlook.com (2603:10b6:408:143::11) by DS0PR04MB9872.namprd04.prod.outlook.com (2603:10b6:8:302::12) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.428.13; Thu, 17 Sep 2026 22:59:34 +0000 Received: from BN2PEPF0000A800.namprd02.prod.outlook.com (2603:10b6:408:143:cafe::27) by BN0PR10CA0020.outlook.office365.com (2603:10b6:408:143::11) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.428.12 via Frontend Transport; Thu, 17 Sep 2026 22:59:34 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 192.86.86.210) smtp.mailfrom=rambus.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=rambus.com; Received-SPF: Pass (protection.outlook.com: domain of rambus.com designates 192.86.86.210 as permitted sender) receiver=protection.outlook.com; client-ip=192.86.86.210; helo=hqxsv-psmtppxy02.rambus.com; pr=C Received: from hqxsv-psmtppxy02.rambus.com (192.86.86.210) by BN2PEPF0000A800.mail.protection.outlook.com (10.167.245.167) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.451.8 via Frontend Transport; Thu, 17 Sep 2026 22:59:34 +0000 Received: from hqxsv-cmdev3-aousherovitch.rambus.com (hqn-lb-int-float.rambus.com [10.12.20.20]) by hqxsv-psmtppxy02.rambus.com (Postfix) with ESMTP id 40CEE1801767; Thu, 17 Sep 2026 22:59:32 +0000 (UTC) From: Alex Ousherovitch To: Albert Ou , Alex Ousherovitch , Conor Dooley , "David S. Miller" , Herbert Xu , Jonathan Corbet , Krzysztof Kozlowski , Palmer Dabbelt , Paul Walmsley , Rob Herring , Saravanakrishnan Krishnamoorthy , Shuah Khan Cc: Alexandre Ghiti , devicetree@vger.kernel.org, Joel Wittenauer , linux-api@vger.kernel.org, linux-crypto@vger.kernel.org, linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-riscv@lists.infradead.org, Shuah Khan , Thi Nguyen Subject: [PATCH v5 10/19] crypto: cmh - add ChaCha20-Poly1305 Date: Thu, 17 Sep 2026 15:59:19 -0700 Message-ID: <20260917225929.2494111-11-aousherovitch@rambus.com> X-Mailer: git-send-email 2.43.7 In-Reply-To: <20260917225929.2494111-1-aousherovitch@rambus.com> References: <20260917225929.2494111-1-aousherovitch@rambus.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: BN2PEPF0000A800:EE_|DS0PR04MB9872:EE_ X-MS-Office365-Filtering-Correlation-Id: 368951f1-204a-4ee1-cb37-08df150f575b X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|1800799024|36860700016|7416014|82310400026|23010399003|376014|10067099003|3023799007|6133799003|18002099003|22082099003|11063799006|56012099006|921020; X-Microsoft-Antispam-Message-Info: T26cwzyhRdorUupZmFKhqtFuTFU2uKQuwo2/dRGfPjbSxAErn3aJGtM8N/yeKFuFk8DyNpG589Lr8erO9yXQ5OGsdeSxaLh1L54i34V45xjg/YUpAy+PvtVWiHPvBP5ctqDB7vrO+tyPKIdSYlnGXk5BdeeJDJ79B0S0hTNYpH1NXlrGuNw9goHnKTiCMz24Bfw3Xi0hOiEZZb1XiCH4bhvAKxTyCVLUNY1u1J/0UCCaIlY9atkrDd4ve4rs8zZUtD6uyl2jOMTAsnWmuHTm9umczPMnhFxJorKoV6Sf6QQrQXRsPVAP236vKdpmU2ZbaMelFJlBdfkVGV8r5XVUM+n5XFG4hg/varuJiBj06FXoQ8f2BKkB3lMX80fB3p1wRggoHn4vNYI/8fm/vnIcTA8fmFZY88g2MqYMBR6JhhMq8MTVhmJKCSN6smnbEDwTy55sdBuVqpRKAPwdvSgbWfhjUZa1x9GiQYaqc+6T8u/o/zqLgoMj6zHidUNvB5StbTSA2YYSC1rk8DSRa6FODK7O1xge4sj+an7YHiA9vm/m8zx7hzMgTRGvXpnQJUkXWBoNCglP7vpqmWQ2i942exYPcNKxpHxNL4FDVq/dQRA+SrOEekQrO8Eh3sktNwTwLcSnPzgqdLsjvGFAJbGbklepDQRJrxJSzUN23nw4fQJNiUJrV0sOEekgNRhXPCis/+v8Z++DeoA0kISiBB9HscTMreiwiAn6TKJ5cfHOVOofgbh9ZhsjnD3iOx52UZDK X-Forefront-Antispam-Report: CIP:192.86.86.210;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:hqxsv-psmtppxy02.rambus.com;PTR:InfoDomainNonexistent;CAT:NONE;SFS:(13230040)(1800799024)(36860700016)(7416014)(82310400026)(23010399003)(376014)(10067099003)(3023799007)(6133799003)(18002099003)(22082099003)(11063799006)(56012099006)(921020);DIR:OUT;SFP:1102; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: NBG7wkWoAdmZnmeLL8gUoA818yf+CN+UNVQMnxRRjHaKOa76YQn5kAl8q1QwhQ9mYbB+VV0IThzdpyQXp3rm2doDuVAbe1mvTzDPCBceBEIaG4UVUivqgbhxeE8sW/pC8okv+ZHeynkoX9xQA/ftPPm6SohJOO7mrxMTJ1NOU/pFJT6zMGXoAmJcjEo4vjfOlmZmQ2lSfQWoQ5NLcY6t5IYN4yuwyOTFUgVufGTPOt4HZSQiE+/oV/rTWTjnPc6G7Ng+PU3RzHXv+h7R8zvaIfEoprA8U4PM43zxrnjUfm019lh+e4eCayELBgCsyodg6RaDq7HofCSmFaGpHYMZtpyQP5ZQlRCpHL0VBKfyUsjCpI6WyHAGb9BYGoyaNuQmSVoKtQx36zOR6dBTePle8gBvZ63WekpnJ56jOZxUxC1eY26QgBn6pHIwi/zeaDBS X-OriginatorOrg: rambus.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 17 Sep 2026 22:59:34.1913 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: 368951f1-204a-4ee1-cb37-08df150f575b X-MS-Exchange-CrossTenant-Id: bd0ba799-c2b9-413c-9c56-5d1731c4827c X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=bd0ba799-c2b9-413c-9c56-5d1731c4827c;Ip=[192.86.86.210];Helo=[hqxsv-psmtppxy02.rambus.com] X-MS-Exchange-CrossTenant-AuthSource: BN2PEPF0000A800.namprd02.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: DS0PR04MB9872 Content-Type: text/plain; charset="utf-8" Register ChaCha20-Poly1305 AEAD and ChaCha20 skcipher algorithms using the CMH CCP core (core ID 0x18). Also registers the Poly1305 ahash for standalone use. Signed-off-by: Alex Ousherovitch Co-developed-by: Saravanakrishnan Krishnamoorthy Signed-off-by: Saravanakrishnan Krishnamoorthy --- drivers/crypto/cmh/Makefile | 5 +- drivers/crypto/cmh/cmh_ccp.c | 362 +++++++++++++ drivers/crypto/cmh/cmh_ccp_aead.c | 600 +++++++++++++++++++++ drivers/crypto/cmh/cmh_ccp_poly.c | 743 +++++++++++++++++++++++++++ drivers/crypto/cmh/cmh_main.c | 25 + drivers/crypto/cmh/include/cmh_ccp.h | 24 + 6 files changed, 1758 insertions(+), 1 deletion(-) create mode 100644 drivers/crypto/cmh/cmh_ccp.c create mode 100644 drivers/crypto/cmh/cmh_ccp_aead.c create mode 100644 drivers/crypto/cmh/cmh_ccp_poly.c create mode 100644 drivers/crypto/cmh/include/cmh_ccp.h diff --git a/drivers/crypto/cmh/Makefile b/drivers/crypto/cmh/Makefile index 032c863d856e..ef24879ba1f1 100644 --- a/drivers/crypto/cmh/Makefile +++ b/drivers/crypto/cmh/Makefile @@ -25,7 +25,10 @@ cmh-y :=3D \ cmh_aes_cmac.o \ cmh_sm4_skcipher.o \ cmh_sm4_aead.o \ - cmh_sm4_cmac.o + cmh_sm4_cmac.o \ + cmh_ccp.o \ + cmh_ccp_aead.o \ + cmh_ccp_poly.o =20 # Management ioctl device (/dev/cmh_mgmt): key lifecycle, PKE, PQC ioctls. cmh-$(CONFIG_CRYPTO_DEV_CMH_MGMT) +=3D \ diff --git a/drivers/crypto/cmh/cmh_ccp.c b/drivers/crypto/cmh/cmh_ccp.c new file mode 100644 index 000000000000..f8daa0c7c433 --- /dev/null +++ b/drivers/crypto/cmh/cmh_ccp.c @@ -0,0 +1,362 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Kernel Crypto API ChaCha20 (skcipher) Driver + * + * Registers the "chacha20" skcipher algorithm with the Linux crypto + * subsystem, backed by the CMH CCP core. + * + * VCQ sequence: + * [SYS_CMD_WRITE] + CCP_CMD_CHACHA20_INIT + CCP_CMD_FINAL + CCP_CMD_FLU= SH + * + * The CCP core expects a 16-byte counter+nonce (ctrnonce): + * bytes [0..3] =3D 32-bit LE counter + * bytes [4..15] =3D 12-byte nonce + * + * The Linux chacha20 skcipher interface passes a 16-byte IV in the + * same format, so we forward it directly. + * + * ChaCha20 is a stream cipher -- arbitrary plaintext lengths are + * supported (no block-alignment requirement). + */ + +#include +#include +#include +#include +#include +#include +#include +#include + +#include "cmh_ccp.h" +#include "cmh_vcq.h" +#include "cmh_ccp_abi.h" +#include "cmh_sys_abi.h" +#include "cmh_sys.h" +#include "cmh_txn.h" +#include "cmh_dma.h" +#include "cmh_key.h" + +/* Per-transform context */ + +struct cmh_ccp_tfm_ctx { + struct cmh_key_ctx key; +}; + +/* Per-request context (lives in skcipher_request::__ctx) */ + +/* + * Maximum payload commands: + * [SYS_CMD_WRITE] + CCP_CMD_CHACHA20_INIT + CCP_CMD_FINAL + FLUSH =3D 4 + */ +#define CMH_CCP_MAX_PAYLOAD 4 +#define CMH_CCP_MAX_PACKED (CMH_CCP_MAX_PAYLOAD * 2) + +struct cmh_ccp_reqctx { + dma_addr_t in_dma; + dma_addr_t out_dma; + dma_addr_t iv_dma; + dma_addr_t key_dma; + u8 *in_buf; + u8 *out_buf; + u8 *iv_buf; + u32 cryptlen; + u32 keylen; + struct vcq_cmd packed[CMH_CCP_MAX_PACKED]; +}; + +/* VCQ Builders -- ChaCha20-specific */ + +static void vcq_add_ccp_chacha_init(struct vcq_cmd *slot, u32 core_id, u64= key_ref, + u64 ctrnonce_dma, u32 keylen, u32 op) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, CCP_CMD_CHACHA20_INIT); + slot->hwc.ccp.cmd_chacha.key =3D key_ref; + slot->hwc.ccp.cmd_chacha.ctrnonce =3D ctrnonce_dma; + slot->hwc.ccp.cmd_chacha.keylen =3D keylen; + slot->hwc.ccp.cmd_chacha.ctrnoncelen =3D CCP_CTRNONCE_SIZE; + slot->hwc.ccp.cmd_chacha.ctrlen =3D CCP_CHACHA_CTR_LEN; + slot->hwc.ccp.cmd_chacha.op =3D op; +} + +static void vcq_add_ccp_final(struct vcq_cmd *slot, u32 core_id, u64 input= _dma, + u64 output_dma, u32 iolen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, CCP_CMD_FINAL); + slot->hwc.ccp.cmd_final.input =3D input_dma; + slot->hwc.ccp.cmd_final.output =3D output_dma; + slot->hwc.ccp.cmd_final.tag =3D 0; + slot->hwc.ccp.cmd_final.iolen =3D iolen; + slot->hwc.ccp.cmd_final.taglen =3D 0; +} + +/* skcipher Operations */ +static int cmh_ccp_setkey(struct crypto_skcipher *tfm, const u8 *key, + unsigned int keylen) +{ + struct cmh_ccp_tfm_ctx *tctx =3D crypto_skcipher_ctx(tfm); + /* ChaCha20 requires 32-byte key per RFC 8439 */ + if (keylen !=3D 32) + return -EINVAL; + + return cmh_key_setkey_raw(&tctx->key, key, keylen, CORE_ID_CCP); +} + +static int cmh_ccp_init_tfm(struct crypto_skcipher *tfm) +{ + struct cmh_ccp_tfm_ctx *tctx =3D crypto_skcipher_ctx(tfm); + + memset(tctx, 0, sizeof(*tctx)); + crypto_skcipher_set_reqsize(tfm, sizeof(struct cmh_ccp_reqctx)); + return 0; +} + +static void cmh_ccp_exit_tfm(struct crypto_skcipher *tfm) +{ + struct cmh_ccp_tfm_ctx *tctx =3D crypto_skcipher_ctx(tfm); + + cmh_key_destroy(&tctx->key); +} + +/* DMA unmap helper */ +static void cmh_ccp_unmap_dma(struct cmh_ccp_reqctx *rctx) +{ + cmh_dma_unmap_single(rctx->iv_dma, CCP_CTRNONCE_SIZE, DMA_TO_DEVICE); + cmh_dma_unmap_single(rctx->out_dma, rctx->cryptlen, DMA_FROM_DEVICE); + cmh_dma_unmap_single(rctx->in_dma, rctx->cryptlen, DMA_TO_DEVICE); +} + +static void cmh_ccp_free_bufs(struct cmh_ccp_reqctx *rctx) +{ + kfree(rctx->iv_buf); + rctx->iv_buf =3D NULL; + kfree_sensitive(rctx->out_buf); + rctx->out_buf =3D NULL; + kfree_sensitive(rctx->in_buf); + rctx->in_buf =3D NULL; +} + +static void cmh_ccp_complete(void *data, int error) +{ + struct skcipher_request *req =3D data; + struct cmh_ccp_reqctx *rctx =3D skcipher_request_ctx(req); + + if (error =3D=3D -EINPROGRESS) { + cmh_complete(&req->base, error); + return; + } + + cmh_ccp_unmap_dma(rctx); + + /* + * ChaCha20 is a one-shot stream cipher. Like the generic + * chacha20 skcipher we leave req->iv unchanged (no counter + * write-back), matching testmgr's output-IV expectation. + */ + if (!error) + scatterwalk_map_and_copy(rctx->out_buf, req->dst, + 0, rctx->cryptlen, 1); + + cmh_ccp_free_bufs(rctx); + cmh_complete(&req->base, error); +} + +/* + * Core encrypt/decrypt -- builds a VCQ transaction and submits async. + * + * ChaCha20 is a stream cipher: encrypt and decrypt use the same + * underlying XOR operation. + */ +static int cmh_ccp_crypt(struct skcipher_request *req, u32 ccp_op) +{ + struct crypto_skcipher *tfm =3D crypto_skcipher_reqtfm(req); + struct cmh_ccp_tfm_ctx *tctx =3D crypto_skcipher_ctx(tfm); + struct cmh_ccp_reqctx *rctx =3D skcipher_request_ctx(req); + struct vcq_cmd cmds[CMH_CCP_MAX_PAYLOAD]; + u64 key_ref; + u32 keylen; + struct core_dispatch d; + s32 target_mbx; + u32 core_id; + u32 idx; + int ret; + gfp_t gfp; + + if (tctx->key.mode =3D=3D CMH_KEY_NONE) + return -ENOKEY; + + if (!req->cryptlen) + return 0; + + /* Limit linearisation buffers to avoid large allocations. */ + if (req->cryptlen > SZ_1M) + return -EINVAL; + + gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + + memset(rctx, 0, sizeof(*rctx)); + rctx->cryptlen =3D req->cryptlen; + + /* Linearise input from scatterlist */ + rctx->in_buf =3D kmalloc(req->cryptlen, gfp | __GFP_NOWARN); + if (!rctx->in_buf) + return -ENOMEM; + + scatterwalk_map_and_copy(rctx->in_buf, req->src, 0, req->cryptlen, 0); + + rctx->in_dma =3D cmh_dma_map_single(rctx->in_buf, req->cryptlen, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->in_dma)) { + ret =3D -ENOMEM; + goto out_free_in; + } + + rctx->out_buf =3D kmalloc(req->cryptlen, gfp | __GFP_NOWARN); + if (!rctx->out_buf) { + ret =3D -ENOMEM; + goto out_unmap_in; + } + + rctx->out_dma =3D cmh_dma_map_single(rctx->out_buf, req->cryptlen, + DMA_FROM_DEVICE); + if (cmh_dma_map_error(rctx->out_dma)) { + ret =3D -ENOMEM; + goto out_free_out; + } + + rctx->iv_buf =3D kmemdup(req->iv, CCP_CTRNONCE_SIZE, gfp); + if (!rctx->iv_buf) { + ret =3D -ENOMEM; + goto out_unmap_out; + } + + rctx->iv_dma =3D cmh_dma_map_single(rctx->iv_buf, CCP_CTRNONCE_SIZE, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->iv_dma)) { + ret =3D -ENOMEM; + goto out_free_iv; + } + + /* Resolve key reference */ + idx =3D 0; + + rctx->key_dma =3D tctx->key.raw.dma; + rctx->keylen =3D tctx->key.raw.len; + vcq_add_sys_write(&cmds[idx++], SYS_REF_TEMP, + (u64)rctx->key_dma, SYS_REF_NONE, + tctx->key.raw.len, + tctx->key.raw.sys_type); + key_ref =3D SYS_REF_TEMP; + keylen =3D tctx->key.raw.len; + d =3D cmh_core_select_instance(CMH_CORE_CCP); + target_mbx =3D d.mbx_idx; + core_id =3D d.core_id; + + vcq_add_ccp_chacha_init(&cmds[idx++], core_id, key_ref, + (u64)rctx->iv_dma, keylen, ccp_op); + + vcq_add_ccp_final(&cmds[idx++], core_id, (u64)rctx->in_dma, + (u64)rctx->out_dma, req->cryptlen); + + vcq_add_flush(&cmds[idx++], core_id); + + ret =3D cmh_vcq_pack_and_submit_async(cmds, idx, rctx->packed, + CMH_CCP_MAX_PACKED, target_mbx, + cmh_ccp_complete, req, + !!(req->base.flags & + CRYPTO_TFM_REQ_MAY_BACKLOG), + cmh_tm_async_timeout_jiffies()); + /* -EBUSY =3D backlogged; ownership transferred to callback. */ + if (ret =3D=3D -EBUSY) + return -EBUSY; + if (ret) + goto out_cleanup_all; + + return -EINPROGRESS; + +out_cleanup_all: + cmh_dma_unmap_single(rctx->iv_dma, CCP_CTRNONCE_SIZE, DMA_TO_DEVICE); +out_free_iv: + kfree(rctx->iv_buf); +out_unmap_out: + cmh_dma_unmap_single(rctx->out_dma, req->cryptlen, DMA_FROM_DEVICE); +out_free_out: + kfree_sensitive(rctx->out_buf); +out_unmap_in: + cmh_dma_unmap_single(rctx->in_dma, req->cryptlen, DMA_TO_DEVICE); +out_free_in: + kfree_sensitive(rctx->in_buf); + return ret; +} + +static int cmh_ccp_encrypt(struct skcipher_request *req) +{ + return cmh_ccp_crypt(req, CCP_OP_ENCRYPT); +} + +static int cmh_ccp_decrypt(struct skcipher_request *req) +{ + return cmh_ccp_crypt(req, CCP_OP_DECRYPT); +} + +/* Registration */ + +static struct skcipher_alg cmh_chacha20_alg =3D { + .setkey =3D cmh_ccp_setkey, + .encrypt =3D cmh_ccp_encrypt, + .decrypt =3D cmh_ccp_decrypt, + .init =3D cmh_ccp_init_tfm, + .exit =3D cmh_ccp_exit_tfm, + .min_keysize =3D 32, + .max_keysize =3D 32, + .ivsize =3D CCP_CTRNONCE_SIZE, + .base =3D { + .cra_name =3D "chacha20", + .cra_driver_name =3D "rambus-cmh-chacha20", + .cra_priority =3D 300, + .cra_flags =3D CRYPTO_ALG_KERN_DRIVER_ONLY | + CRYPTO_ALG_ASYNC, + .cra_blocksize =3D 1, /* stream cipher */ + .cra_ctxsize =3D sizeof(struct cmh_ccp_tfm_ctx), + .cra_module =3D THIS_MODULE, + }, +}; + +/** + * cmh_ccp_register() - Register ChaCha20 skcipher algorithm with the cryp= to framework + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_ccp_register(void) +{ + int ret; + + if (!cmh_core_present(CMH_CORE_CCP)) + return 0; + + ret =3D crypto_register_skcipher(&cmh_chacha20_alg); + if (ret) + dev_err(cmh_dev(), "cmh_ccp: failed to register chacha20 (rc=3D%d)\n", r= et); + else + dev_dbg(cmh_dev(), "cmh_ccp: registered chacha20\n"); + + return ret; +} + +/** + * cmh_ccp_unregister() - Unregister ChaCha20 skcipher algorithm from the = crypto framework + */ +void cmh_ccp_unregister(void) +{ + if (!cmh_core_present(CMH_CORE_CCP)) + return; + + crypto_unregister_skcipher(&cmh_chacha20_alg); + dev_dbg(cmh_dev(), "cmh_ccp: unregistered chacha20\n"); +} diff --git a/drivers/crypto/cmh/cmh_ccp_aead.c b/drivers/crypto/cmh/cmh_ccp= _aead.c new file mode 100644 index 000000000000..817eaa0fc6e3 --- /dev/null +++ b/drivers/crypto/cmh/cmh_ccp_aead.c @@ -0,0 +1,600 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Kernel Crypto API ChaCha20-Poly1305 AEAD Driver (RFC 7539) + * + * Registers "rfc7539(chacha20,poly1305)" as an AEAD algorithm with the + * Linux crypto subsystem, backed by the CMH CCP core. + * + * VCQ sequence: + * [SYS_CMD_WRITE] + CCP_CMD_AEAD_INIT + CCP_CMD_AAD_FINAL + * + CCP_CMD_FINAL + CCP_CMD_FLUSH + * + * CCP_CMD_AAD_FINAL is always issued -- even with no associated data -- + * because it transitions the CCP core out of its AAD phase into the + * state CCP_CMD_FINAL requires. + * + * The RFC 7539 AEAD interface passes a 12-byte nonce via req->iv. + * The CCP core expects a 16-byte ctrnonce (4-byte LE counter + 12-byte + * nonce). We prepend a zero counter (per RFC 7539 S2.8: counter 0 + * generates the Poly1305 key, counter 1 starts encryption -- the + * CMH eSW handles this internally from the initial counter value of 0). + * + * Tag is always 16 bytes (Poly1305 authenticator). + */ + +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "cmh_ccp.h" +#include "cmh_vcq.h" +#include "cmh_ccp_abi.h" +#include "cmh_sys_abi.h" +#include "cmh_sys.h" +#include "cmh_txn.h" +#include "cmh_dma.h" +#include "cmh_key.h" + +#define CCP_AEAD_IV_SIZE 12U /* RFC 7539 nonce */ +#define CCP_ESP_IV_SIZE 8U /* RFC 7539 ESP nonce (4-byte salt at setkey) = */ +#define CCP_ESP_SALT_SIZE 4U +#define CCP_AEAD_TAG_SIZE 16U /* Poly1305 tag */ + +struct cmh_ccp_aead_tfm_ctx { + struct cmh_key_ctx key; + u32 authsize; + u8 salt[CCP_ESP_SALT_SIZE]; /* ESP salt (unused for rfc7539) */ +}; + +/* Per-request context (lives in aead_request::__ctx) */ + +/* + * Maximum payload commands: + * [SYS_CMD_WRITE] + CCP_CMD_AEAD_INIT + CCP_CMD_AAD_FINAL + * + CCP_CMD_FINAL + FLUSH =3D 5 + */ +#define CMH_CCP_AEAD_MAX_PAYLOAD 5 +#define CMH_CCP_AEAD_MAX_PACKED (CMH_CCP_AEAD_MAX_PAYLOAD * 2) + +struct cmh_ccp_aead_reqctx { + dma_addr_t in_dma; + dma_addr_t out_dma; + dma_addr_t iv_dma; + dma_addr_t key_dma; + dma_addr_t aad_dma; + dma_addr_t tag_dma; + u8 *in_buf; + u8 *out_buf; + u8 *iv_buf; + u8 *aad_buf; + u8 *tag_buf; + u32 cryptlen; + u32 assoclen; + u32 authsize; + u32 keylen; + bool encrypting; + struct vcq_cmd packed[CMH_CCP_AEAD_MAX_PACKED]; +}; + +/* VCQ Builders -- CCP AEAD-specific */ + +static void vcq_add_ccp_aead_init(struct vcq_cmd *slot, u32 core_id, u64 k= ey_ref, + u64 ctrnonce_dma, u32 keylen, u32 op) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, CCP_CMD_AEAD_INIT); + slot->hwc.ccp.cmd_aead.key =3D key_ref; + slot->hwc.ccp.cmd_aead.ctrnonce =3D ctrnonce_dma; + slot->hwc.ccp.cmd_aead.keylen =3D keylen; + slot->hwc.ccp.cmd_aead.ctrnoncelen =3D CCP_CTRNONCE_SIZE; + slot->hwc.ccp.cmd_aead.op =3D op; +} + +static void vcq_add_ccp_aad_final(struct vcq_cmd *slot, u32 core_id, u64 a= ad_dma, + u32 aadlen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, CCP_CMD_AAD_FINAL); + slot->hwc.ccp.cmd_aad_final.aad =3D aad_dma; + slot->hwc.ccp.cmd_aad_final.aadlen =3D aadlen; +} + +static void vcq_add_ccp_aead_final(struct vcq_cmd *slot, u32 core_id, u64 = input_dma, + u64 output_dma, u64 tag_dma, + u32 iolen, u32 taglen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, CCP_CMD_FINAL); + slot->hwc.ccp.cmd_final.input =3D input_dma; + slot->hwc.ccp.cmd_final.output =3D output_dma; + slot->hwc.ccp.cmd_final.tag =3D tag_dma; + slot->hwc.ccp.cmd_final.iolen =3D iolen; + slot->hwc.ccp.cmd_final.taglen =3D taglen; +} + +/* setkey */ +static int cmh_ccp_aead_setkey(struct crypto_aead *tfm, const u8 *key, + unsigned int keylen) +{ + struct cmh_ccp_aead_tfm_ctx *tctx =3D crypto_aead_ctx(tfm); + /* RFC 7539 AEAD requires 32-byte key */ + if (keylen !=3D CHACHA_KEY_SIZE) + return -EINVAL; + + return cmh_key_setkey_raw(&tctx->key, key, keylen, CORE_ID_CCP); +} + +static int cmh_ccp_aead_setauthsize(struct crypto_aead *tfm, + unsigned int authsize) +{ + struct cmh_ccp_aead_tfm_ctx *tctx =3D crypto_aead_ctx(tfm); + + /* Poly1305 tag is always 16 bytes */ + if (authsize !=3D CCP_AEAD_TAG_SIZE) + return -EINVAL; + + tctx->authsize =3D authsize; + return 0; +} + +static int cmh_ccp_aead_init_tfm(struct crypto_aead *tfm) +{ + struct cmh_ccp_aead_tfm_ctx *tctx =3D crypto_aead_ctx(tfm); + + memset(tctx, 0, sizeof(*tctx)); + tctx->authsize =3D CCP_AEAD_TAG_SIZE; + crypto_aead_set_reqsize(tfm, sizeof(struct cmh_ccp_aead_reqctx)); + return 0; +} + +static void cmh_ccp_aead_exit_tfm(struct crypto_aead *tfm) +{ + struct cmh_ccp_aead_tfm_ctx *tctx =3D crypto_aead_ctx(tfm); + + cmh_key_destroy(&tctx->key); +} + +/* DMA unmap helper */ +static void cmh_ccp_aead_unmap_dma(struct cmh_ccp_aead_reqctx *rctx) +{ + cmh_dma_unmap_single(rctx->iv_dma, CCP_CTRNONCE_SIZE, DMA_TO_DEVICE); + cmh_dma_unmap_single(rctx->tag_dma, rctx->authsize, + rctx->encrypting ? DMA_FROM_DEVICE : + DMA_TO_DEVICE); + if (rctx->cryptlen > 0) { + cmh_dma_unmap_single(rctx->out_dma, rctx->cryptlen, + DMA_FROM_DEVICE); + cmh_dma_unmap_single(rctx->in_dma, rctx->cryptlen, + DMA_TO_DEVICE); + } + if (rctx->assoclen > 0) + cmh_dma_unmap_single(rctx->aad_dma, rctx->assoclen, + DMA_TO_DEVICE); +} + +static void cmh_ccp_aead_free_bufs(struct cmh_ccp_aead_reqctx *rctx) +{ + kfree(rctx->iv_buf); + rctx->iv_buf =3D NULL; + kfree(rctx->tag_buf); + rctx->tag_buf =3D NULL; + kfree_sensitive(rctx->out_buf); + rctx->out_buf =3D NULL; + kfree_sensitive(rctx->in_buf); + rctx->in_buf =3D NULL; + kfree(rctx->aad_buf); + rctx->aad_buf =3D NULL; +} + +static void cmh_ccp_aead_complete(void *data, int error) +{ + struct aead_request *req =3D data; + struct cmh_ccp_aead_reqctx *rctx =3D aead_request_ctx(req); + + if (error =3D=3D -EINPROGRESS) { + cmh_complete(&req->base, error); + return; + } + + cmh_ccp_aead_unmap_dma(rctx); + + /* + * Map HW error on decrypt to -EBADMSG. The eSW CCP core uses a + * single error code (-EIO) for both authentication failures and + * other core errors (e.g. DMA timeout), so we cannot distinguish + * them from the MBX_STATUS alone. In practice the only error + * during a well-formed AEAD decrypt is auth-tag mismatch; a DMA + * timeout would indicate a fatal HW problem where -EBADMSG vs + * -EIO is moot. The kernel crypto API requires -EBADMSG for + * AEAD authentication failures. + */ + if (error =3D=3D -EIO && !rctx->encrypting) + error =3D -EBADMSG; + + if (!error) { + if (rctx->cryptlen > 0) + scatterwalk_map_and_copy(rctx->out_buf, req->dst, + req->assoclen, + rctx->cryptlen, 1); + if (rctx->encrypting) + scatterwalk_map_and_copy(rctx->tag_buf, req->dst, + req->assoclen + + rctx->cryptlen, + rctx->authsize, 1); + } + + cmh_ccp_aead_free_bufs(rctx); + cmh_complete(&req->base, error); +} + +/* + * Core AEAD encrypt/decrypt -- async path. + * + * Encrypt: plaintext -> ciphertext + 16-byte tag + * Decrypt: ciphertext + tag -> plaintext (tag verified by CMH eSW) + * + * VCQ: [SYS_CMD_WRITE] + AEAD_INIT + [AAD_FINAL] + FINAL + FLUSH + */ +static int cmh_ccp_aead_crypt(struct aead_request *req, u32 ccp_op) +{ + struct crypto_aead *tfm =3D crypto_aead_reqtfm(req); + struct cmh_ccp_aead_tfm_ctx *tctx =3D crypto_aead_ctx(tfm); + struct cmh_ccp_aead_reqctx *rctx =3D aead_request_ctx(req); + struct vcq_cmd cmds[CMH_CCP_AEAD_MAX_PAYLOAD]; + u64 key_ref; + u32 keylen, authsize, cryptlen; + struct core_dispatch d; + s32 target_mbx; + u32 core_id; + u32 idx; + int ret; + gfp_t gfp; + + if (tctx->key.mode =3D=3D CMH_KEY_NONE) + return -ENOKEY; + + authsize =3D tctx->authsize; + + if (ccp_op =3D=3D CCP_OP_ENCRYPT) { + cryptlen =3D req->cryptlen; + } else { + if (req->cryptlen < authsize) + return -EINVAL; + cryptlen =3D req->cryptlen - authsize; + } + + /* + * HW uses a proprietary LLI scatter-gather format that is + * incompatible with struct scatterlist, so the payload is + * linearised into contiguous buffers for DMA. Cap total + * size to prevent excessive memory consumption. + */ + if ((u64)cryptlen + req->assoclen > SZ_1M) + return -EINVAL; + + gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + + memset(rctx, 0, sizeof(*rctx)); + rctx->cryptlen =3D cryptlen; + rctx->assoclen =3D req->assoclen; + rctx->authsize =3D authsize; + rctx->encrypting =3D (ccp_op =3D=3D CCP_OP_ENCRYPT); + + /* + * rfc7539esp: the last ivsize (8) bytes of the AAD region are the + * IV/nonce, not actual associated data. Subtract them so HW only + * authenticates the real AAD. + */ + if (crypto_aead_ivsize(tfm) =3D=3D CCP_ESP_IV_SIZE) { + if (rctx->assoclen < CCP_ESP_IV_SIZE) + return -EINVAL; + rctx->assoclen -=3D CCP_ESP_IV_SIZE; + } + + /* Linearise AAD */ + if (rctx->assoclen > 0) { + rctx->aad_buf =3D kmalloc(rctx->assoclen, gfp | __GFP_NOWARN); + if (!rctx->aad_buf) + return -ENOMEM; + scatterwalk_map_and_copy(rctx->aad_buf, req->src, + 0, rctx->assoclen, 0); + rctx->aad_dma =3D cmh_dma_map_single(rctx->aad_buf, + rctx->assoclen, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->aad_dma)) { + ret =3D -ENOMEM; + goto out_free_aad; + } + } + + /* Linearise input */ + if (cryptlen > 0) { + rctx->in_buf =3D kmalloc(cryptlen, gfp | __GFP_NOWARN); + if (!rctx->in_buf) { + ret =3D -ENOMEM; + goto out_unmap_aad; + } + scatterwalk_map_and_copy(rctx->in_buf, req->src, + req->assoclen, cryptlen, 0); + rctx->in_dma =3D cmh_dma_map_single(rctx->in_buf, cryptlen, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->in_dma)) { + ret =3D -ENOMEM; + goto out_free_in; + } + } + + /* Allocate output buffer */ + if (cryptlen > 0) { + rctx->out_buf =3D kmalloc(cryptlen, gfp | __GFP_NOWARN); + if (!rctx->out_buf) { + ret =3D -ENOMEM; + goto out_unmap_in; + } + rctx->out_dma =3D cmh_dma_map_single(rctx->out_buf, cryptlen, + DMA_FROM_DEVICE); + if (cmh_dma_map_error(rctx->out_dma)) { + ret =3D -ENOMEM; + goto out_free_out; + } + } + + /* Tag buffer */ + rctx->tag_buf =3D kmalloc(authsize, gfp); + if (!rctx->tag_buf) { + ret =3D -ENOMEM; + goto out_unmap_out; + } + + if (!rctx->encrypting) { + scatterwalk_map_and_copy(rctx->tag_buf, req->src, + req->assoclen + cryptlen, + authsize, 0); + } else { + memset(rctx->tag_buf, 0, authsize); + } + + rctx->tag_dma =3D cmh_dma_map_single(rctx->tag_buf, authsize, + rctx->encrypting ? + DMA_FROM_DEVICE : DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->tag_dma)) { + ret =3D -ENOMEM; + goto out_free_tag; + } + + /* Build 16-byte ctrnonce: 4-byte zero counter + 12-byte nonce. + * rfc7539: counter(4) | req->iv(12) + * rfc7539esp: counter(4) | salt(4) | req->iv(8) + */ + rctx->iv_buf =3D kzalloc(CCP_CTRNONCE_SIZE, gfp); + if (!rctx->iv_buf) { + ret =3D -ENOMEM; + goto out_unmap_tag; + } + if (crypto_aead_ivsize(tfm) =3D=3D CCP_ESP_IV_SIZE) { + memcpy(rctx->iv_buf + CCP_CHACHA_CTR_LEN, + tctx->salt, CCP_ESP_SALT_SIZE); + memcpy(rctx->iv_buf + CCP_CHACHA_CTR_LEN + CCP_ESP_SALT_SIZE, + req->iv, CCP_ESP_IV_SIZE); + } else { + memcpy(rctx->iv_buf + CCP_CHACHA_CTR_LEN, + req->iv, CCP_AEAD_IV_SIZE); + } + + rctx->iv_dma =3D cmh_dma_map_single(rctx->iv_buf, CCP_CTRNONCE_SIZE, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->iv_dma)) { + ret =3D -ENOMEM; + goto out_free_iv; + } + + /* Resolve key reference */ + idx =3D 0; + + rctx->key_dma =3D tctx->key.raw.dma; + rctx->keylen =3D tctx->key.raw.len; + vcq_add_sys_write(&cmds[idx++], SYS_REF_TEMP, + (u64)rctx->key_dma, SYS_REF_NONE, + tctx->key.raw.len, + tctx->key.raw.sys_type); + key_ref =3D SYS_REF_TEMP; + keylen =3D tctx->key.raw.len; + d =3D cmh_core_select_instance(CMH_CORE_CCP); + target_mbx =3D d.mbx_idx; + core_id =3D d.core_id; + + /* AEAD_INIT */ + vcq_add_ccp_aead_init(&cmds[idx++], core_id, key_ref, + (u64)rctx->iv_dma, keylen, ccp_op); + + /* + * AAD_FINAL closes the AAD phase and moves the CCP core from + * STATE_AAD to STATE_UPDATE_WAIT, which CCP_CMD_FINAL requires. + * It must be issued even with no associated data (assoclen =3D=3D 0): + * otherwise the core stays in STATE_AAD and FINAL fails with -EIO + * (e.g. testmgr rfc7539 fuzz vector "alen=3D0 plen=3D11"). With + * assoclen 0 the eSW skips the DMA and only records AAD_LEN =3D 0. + */ + vcq_add_ccp_aad_final(&cmds[idx++], core_id, + rctx->assoclen > 0 ? (u64)rctx->aad_dma : 0, + rctx->assoclen); + + /* FINAL with tag */ + vcq_add_ccp_aead_final(&cmds[idx++], core_id, + cryptlen > 0 ? (u64)rctx->in_dma : 0, + cryptlen > 0 ? (u64)rctx->out_dma : 0, + (u64)rctx->tag_dma, cryptlen, authsize); + + vcq_add_flush(&cmds[idx++], core_id); + + ret =3D cmh_vcq_pack_and_submit_async(cmds, idx, rctx->packed, + CMH_CCP_AEAD_MAX_PACKED, + target_mbx, + cmh_ccp_aead_complete, req, + !!(req->base.flags & + CRYPTO_TFM_REQ_MAY_BACKLOG), + cmh_tm_async_timeout_jiffies()); + if (ret =3D=3D -EBUSY) + return -EBUSY; + if (ret) + goto out_cleanup_all; + + return -EINPROGRESS; + +out_cleanup_all: + cmh_dma_unmap_single(rctx->iv_dma, CCP_CTRNONCE_SIZE, DMA_TO_DEVICE); +out_free_iv: + kfree(rctx->iv_buf); +out_unmap_tag: + cmh_dma_unmap_single(rctx->tag_dma, authsize, + rctx->encrypting ? DMA_FROM_DEVICE : + DMA_TO_DEVICE); +out_free_tag: + kfree(rctx->tag_buf); +out_unmap_out: + if (cryptlen > 0) + cmh_dma_unmap_single(rctx->out_dma, cryptlen, DMA_FROM_DEVICE); +out_free_out: + kfree_sensitive(rctx->out_buf); +out_unmap_in: + if (cryptlen > 0) + cmh_dma_unmap_single(rctx->in_dma, cryptlen, DMA_TO_DEVICE); +out_free_in: + kfree_sensitive(rctx->in_buf); +out_unmap_aad: + if (rctx->assoclen > 0) + cmh_dma_unmap_single(rctx->aad_dma, rctx->assoclen, + DMA_TO_DEVICE); +out_free_aad: + kfree(rctx->aad_buf); + return ret; +} + +static int cmh_ccp_aead_encrypt(struct aead_request *req) +{ + return cmh_ccp_aead_crypt(req, CCP_OP_ENCRYPT); +} + +static int cmh_ccp_aead_decrypt(struct aead_request *req) +{ + return cmh_ccp_aead_crypt(req, CCP_OP_DECRYPT); +} + +/* -- rfc7539esp: ESP variant with 4-byte salt + 8-byte IV ---------------= */ + +/* + * ESP setkey: 36 bytes =3D 32-byte ChaCha20 key + 4-byte salt. + * The salt is prepended to the 8-byte per-packet IV from the ESP header + * to form the 12-byte RFC 7539 nonce. + */ +static int cmh_ccp_esp_setkey(struct crypto_aead *tfm, const u8 *key, + unsigned int keylen) +{ + struct cmh_ccp_aead_tfm_ctx *tctx =3D crypto_aead_ctx(tfm); + + if (keylen !=3D CHACHA_KEY_SIZE + CCP_ESP_SALT_SIZE) + return -EINVAL; + + memcpy(tctx->salt, key + CHACHA_KEY_SIZE, CCP_ESP_SALT_SIZE); + return cmh_key_setkey_raw(&tctx->key, key, CHACHA_KEY_SIZE, CORE_ID_CCP); +} + +/* Registration */ + +static struct aead_alg cmh_rfc7539_alg =3D { + .setkey =3D cmh_ccp_aead_setkey, + .setauthsize =3D cmh_ccp_aead_setauthsize, + .encrypt =3D cmh_ccp_aead_encrypt, + .decrypt =3D cmh_ccp_aead_decrypt, + .init =3D cmh_ccp_aead_init_tfm, + .exit =3D cmh_ccp_aead_exit_tfm, + .ivsize =3D CCP_AEAD_IV_SIZE, + .maxauthsize =3D CCP_AEAD_TAG_SIZE, + .base =3D { + .cra_name =3D "rfc7539(chacha20,poly1305)", + .cra_driver_name =3D "rambus-cmh-rfc7539-chacha20-poly1305", + .cra_priority =3D 300, + .cra_flags =3D CRYPTO_ALG_KERN_DRIVER_ONLY | + CRYPTO_ALG_ASYNC, + .cra_blocksize =3D 1, + .cra_ctxsize =3D sizeof(struct cmh_ccp_aead_tfm_ctx), + .cra_module =3D THIS_MODULE, + }, +}; + +static struct aead_alg cmh_rfc7539esp_alg =3D { + .setkey =3D cmh_ccp_esp_setkey, + .setauthsize =3D cmh_ccp_aead_setauthsize, + .encrypt =3D cmh_ccp_aead_encrypt, + .decrypt =3D cmh_ccp_aead_decrypt, + .init =3D cmh_ccp_aead_init_tfm, + .exit =3D cmh_ccp_aead_exit_tfm, + .ivsize =3D CCP_ESP_IV_SIZE, + .maxauthsize =3D CCP_AEAD_TAG_SIZE, + .base =3D { + .cra_name =3D "rfc7539esp(chacha20,poly1305)", + .cra_driver_name =3D "rambus-cmh-rfc7539esp-chacha20-poly1305", + .cra_priority =3D 300, + .cra_flags =3D CRYPTO_ALG_KERN_DRIVER_ONLY | + CRYPTO_ALG_ASYNC, + .cra_blocksize =3D 1, + .cra_ctxsize =3D sizeof(struct cmh_ccp_aead_tfm_ctx), + .cra_module =3D THIS_MODULE, + }, +}; + +/** + * cmh_ccp_aead_register() - Register ChaCha20-Poly1305 AEAD algorithm wit= h the crypto framework + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_ccp_aead_register(void) +{ + int ret; + + if (!cmh_core_present(CMH_CORE_CCP)) + return 0; + + ret =3D crypto_register_aead(&cmh_rfc7539_alg); + if (ret) { + dev_err(cmh_dev(), "cmh_ccp_aead: failed to register rfc7539 (rc=3D%d)\n= ", + ret); + return ret; + } + dev_dbg(cmh_dev(), "cmh_ccp_aead: registered rfc7539(chacha20,poly1305)\n= "); + + ret =3D crypto_register_aead(&cmh_rfc7539esp_alg); + if (ret) { + dev_err(cmh_dev(), "cmh_ccp_aead: failed to register rfc7539esp (rc=3D%d= )\n", + ret); + crypto_unregister_aead(&cmh_rfc7539_alg); + return ret; + } + dev_dbg(cmh_dev(), "cmh_ccp_aead: registered rfc7539esp(chacha20,poly1305= )\n"); + + return 0; +} + +/** + * cmh_ccp_aead_unregister() - Unregister ChaCha20-Poly1305 AEAD algorithms + */ +void cmh_ccp_aead_unregister(void) +{ + if (!cmh_core_present(CMH_CORE_CCP)) + return; + + crypto_unregister_aead(&cmh_rfc7539esp_alg); + crypto_unregister_aead(&cmh_rfc7539_alg); + dev_dbg(cmh_dev(), "cmh_ccp_aead: unregistered rfc7539/rfc7539esp\n"); +} diff --git a/drivers/crypto/cmh/cmh_ccp_poly.c b/drivers/crypto/cmh/cmh_ccp= _poly.c new file mode 100644 index 000000000000..d8482dd96d12 --- /dev/null +++ b/drivers/crypto/cmh/cmh_ccp_poly.c @@ -0,0 +1,743 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Kernel Crypto API Poly1305 (ahash) Driver + * + * Registers "poly1305" as an ahash algorithm with the Linux crypto + * subsystem, backed by the CMH CCP core. + * + * Poly1305 is a one-time authenticator that produces a 16-byte MAC. + * It requires two 16-byte keys: r (clamped multiplier) and s (nonce). + * + * Key format: 32 bytes =3D r_key[0..15] || s_key[16..31] + * This matches the Poly1305 key layout in RFC 7539 S2.5. + * + * VCQ sequence: + * SYS_CMD_WRITE(s_key) + SYS_CMD_WRITE(r_key) + * + CCP_CMD_POLY1305_INIT + CCP_CMD_FINAL + CCP_CMD_FLUSH + * + * Both keys are written to SYS_REF_TEMP; the CMH eSW stacks them + * so that POLY1305_INIT finds r_key (most recent) as rkey and + * s_key (previous) as skey. + * + * The ahash interface accumulates data via .update() and submits the + * full VCQ asynchronously in .final(). Because the CCP core exposes no + * external save/restore, input is buffered in kernel memory and capped + * at POLY_MAX_DATA (64 KB). Past that cap -- or when the flat export + * window is exceeded -- the request transparently switches to the + * Poly1305 library (), which keeps arbitrary-length + * MACs and transform clone (export/import) conformant with O(1) memory. + */ + +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "cmh_ccp.h" +#include "cmh_vcq.h" +#include "cmh_ccp_abi.h" +#include "cmh_sys_abi.h" +#include "cmh_sys.h" +#include "cmh_txn.h" +#include "cmh_dma.h" +#include "cmh_key.h" + +/* + * Maximum accumulated data for Poly1305 -- driver-imposed, not HW. + * + * The CCP core does not expose external save/restore VCQ commands, + * so the driver must accumulate all data in kernel memory via + * .update() and submit it atomically in .final(). This cap limits + * the per-request kernel allocation. + */ +#define POLY_MAX_DATA (64 * 1024) + +/* + * Per-transform cap on total bytes buffered across all_chunks. Bounds + * the memory an AF_ALG client can pin by repeatedly opening request + * sockets, updating, and abandoning them (the crypto API has no + * per-request destructor, so their chunks live until the TFM is freed). + */ +#define CMH_POLY_TFM_MAX_BUFFERED (16 * 1024 * 1024) + +/* + * Per-transform context -- stores the raw 32-byte key (r || s). + * + * Only the raw-key path is supported for standalone Poly1305. + */ +struct cmh_poly_tfm_ctx { + u8 *key; /* kmalloc'd (r || s); DMA-safe */ + dma_addr_t rkey_dma; + dma_addr_t skey_dma; + u32 keylen; + bool has_key; + spinlock_t chunk_lock; /* protects all_chunks + tfm_buffered */ + struct list_head all_chunks; /* orphan-safe chunk tracking */ + size_t tfm_buffered; /* bytes on all_chunks; DoS cap */ +}; + +/* Chunk node for O(1) update() appends */ +struct cmh_poly_chunk { + struct list_head list; + struct list_head tfm_node; /* per-tfm orphan tracking */ + u32 len; + u8 data[]; +}; + +/* Per-request context (lives in ahash_request::__ctx) */ + +/* + * Maximum payload commands: + * SYS_CMD_WRITE(s) + SYS_CMD_WRITE(r) + POLY1305_INIT + * + CCP_CMD_FINAL + FLUSH =3D 5 + */ +#define CMH_POLY_MAX_PAYLOAD 5 +#define CMH_POLY_MAX_PACKED (CMH_POLY_MAX_PAYLOAD * 2) + +struct cmh_poly_reqctx { + struct list_head chunks; + u32 total_len; + bool switched; /* handed off to the Poly1305 library */ + struct poly1305_desc_ctx fb_state; /* SW fallback running state */ + u8 *buf; /* linearised in final() */ + /* DMA state for async final */ + dma_addr_t in_dma; + dma_addr_t tag_dma; + u8 *tag_buf; + struct vcq_cmd packed[CMH_POLY_MAX_PACKED]; +}; + +/* + * Export/import (transform clone): the CCP core lacks external + * save/restore, so the driver serialises its accumulated input for the + * common (bounded) case (CMH_POLY_FMT_RAW). When the input exceeds the + * HW cap (POLY_MAX_DATA, 64 KB) or the flat export window, the request + * switches to the Poly1305 library and serialises the library's + * fixed-size running state (CMH_POLY_FMT_FB), so arbitrary-length MACs + * and transform clone both stay conformant with O(1) driver memory. + */ +#define CMH_POLY_FMT_RAW 0 +#define CMH_POLY_FMT_FB 1 + +struct cmh_poly_export_state { + u8 format; + u8 __pad[3]; + u32 total_len; + u8 data[]; +}; + +#define CMH_POLY_STATE_SIZE 4096 +#define CMH_POLY_EXPORT_MAX \ + (CMH_POLY_STATE_SIZE - sizeof(struct cmh_poly_export_state)) + +static void vcq_add_ccp_poly_init(struct vcq_cmd *slot, u32 core_id, + u64 rkey_ref, u32 rkeylen, + u64 skey_ref, u32 skeylen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, CCP_CMD_POLY1305_INIT); + slot->hwc.ccp.cmd_poly.rkey =3D rkey_ref; + slot->hwc.ccp.cmd_poly.rkeylen =3D rkeylen; + slot->hwc.ccp.cmd_poly.skey =3D skey_ref; + slot->hwc.ccp.cmd_poly.skeylen =3D skeylen; +} + +static void vcq_add_ccp_poly_final(struct vcq_cmd *slot, u32 core_id, + u64 input_dma, u64 tag_dma, + u32 iolen, u32 taglen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, CCP_CMD_FINAL); + slot->hwc.ccp.cmd_final.input =3D input_dma; + slot->hwc.ccp.cmd_final.output =3D 0; + slot->hwc.ccp.cmd_final.tag =3D tag_dma; + slot->hwc.ccp.cmd_final.iolen =3D iolen; + slot->hwc.ccp.cmd_final.taglen =3D taglen; +} + +static int cmh_poly_setkey(struct crypto_ahash *tfm, const u8 *key, + unsigned int keylen) +{ + struct cmh_poly_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + + /* Poly1305: exactly 32 bytes (r[16] + s[16]) */ + if (keylen !=3D POLY1305_KEY_SIZE) + return -EINVAL; + + /* Unmap old key DMA if re-keying */ + if (tctx->has_key) { + cmh_dma_unmap_single(tctx->rkey_dma, CCP_POLY_KEY_SIZE, + DMA_TO_DEVICE); + cmh_dma_unmap_single(tctx->skey_dma, CCP_POLY_KEY_SIZE, + DMA_TO_DEVICE); + } + + /* + * DMA the key from its own kmalloc'd buffer, not an inline tfm-ctx + * field: a standalone allocation is DMA-safe and cannot share a + * cacheline with CPU-written ctx fields (chunk_lock, all_chunks). + */ + if (!tctx->key) { + tctx->key =3D kmalloc(POLY1305_KEY_SIZE, GFP_KERNEL); + if (!tctx->key) + return -ENOMEM; + } + memcpy(tctx->key, key, POLY1305_KEY_SIZE); + tctx->keylen =3D POLY1305_KEY_SIZE; + + /* + * Pre-map both key halves for DMA. The key buffer lives in + * the tfm context and is stable until exit_tfm() or re-setkey. + */ + tctx->skey_dma =3D cmh_dma_map_single(tctx->key + CCP_POLY_KEY_SIZE, + CCP_POLY_KEY_SIZE, + DMA_TO_DEVICE); + if (cmh_dma_map_error(tctx->skey_dma)) { + tctx->has_key =3D false; + return -ENOMEM; + } + + tctx->rkey_dma =3D cmh_dma_map_single(tctx->key, CCP_POLY_KEY_SIZE, + DMA_TO_DEVICE); + if (cmh_dma_map_error(tctx->rkey_dma)) { + cmh_dma_unmap_single(tctx->skey_dma, CCP_POLY_KEY_SIZE, + DMA_TO_DEVICE); + tctx->has_key =3D false; + return -ENOMEM; + } + + tctx->has_key =3D true; + return 0; +} + +static void cmh_poly_free_chunks(struct cmh_poly_reqctx *rctx, + struct cmh_poly_tfm_ctx *tctx) +{ + struct cmh_poly_chunk *c, *tmp; + + spin_lock_bh(&tctx->chunk_lock); + list_for_each_entry_safe(c, tmp, &rctx->chunks, list) { + list_del(&c->list); + list_del(&c->tfm_node); + tctx->tfm_buffered -=3D c->len; + kfree_sensitive(c); + } + spin_unlock_bh(&tctx->chunk_lock); + rctx->total_len =3D 0; +} + +/* Software-fallback helpers (arbitrary-length + transform-clone support) = */ + +/* Feed a scatterlist to the Poly1305 library, segment by segment. */ +static void cmh_poly_fb_feed_sg(struct poly1305_desc_ctx *desc, + struct scatterlist *sg, u32 len) +{ + struct sg_mapping_iter miter; + u32 remaining =3D len; + + sg_miter_start(&miter, sg, sg_nents(sg), + SG_MITER_FROM_SG | SG_MITER_ATOMIC); + while (remaining && sg_miter_next(&miter)) { + u32 n =3D min_t(u32, miter.length, remaining); + + poly1305_update(desc, miter.addr, n); + remaining -=3D n; + } + sg_miter_stop(&miter); +} + +/* + * Switch a request from the HW-buffered path to the Poly1305 library: + * initialise the library state with the transform key, replay every + * accumulated chunk through it, then drop the chunks. Afterwards the + * request is O(1) in memory and no longer input-capped. + */ +static int cmh_poly_switch_to_fb(struct ahash_request *req) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_poly_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cmh_poly_reqctx *rctx =3D ahash_request_ctx(req); + struct cmh_poly_chunk *c; + + if (!tctx->has_key) + return -ENOKEY; + + poly1305_init(&rctx->fb_state, tctx->key); + list_for_each_entry(c, &rctx->chunks, list) + poly1305_update(&rctx->fb_state, c->data, c->len); + cmh_poly_free_chunks(rctx, tctx); + rctx->switched =3D true; + return 0; +} + +/* Forward the current update() payload to the library. */ +static void cmh_poly_fb_forward(struct ahash_request *req, + struct cmh_poly_reqctx *rctx) +{ + if (req->base.flags & CRYPTO_AHASH_REQ_VIRT) + poly1305_update(&rctx->fb_state, req->svirt, req->nbytes); + else + cmh_poly_fb_feed_sg(&rctx->fb_state, req->src, req->nbytes); +} + +static int cmh_poly_init(struct ahash_request *req) +{ + struct cmh_poly_reqctx *rctx =3D ahash_request_ctx(req); + + memset(rctx, 0, sizeof(*rctx)); + INIT_LIST_HEAD(&rctx->chunks); + return 0; +} + +static int cmh_poly_update(struct ahash_request *req) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_poly_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cmh_poly_reqctx *rctx =3D ahash_request_ctx(req); + struct cmh_poly_chunk *chunk; + gfp_t gfp; + int ret; + + if (!req->nbytes) + return 0; + + /* Already handed off to the library: forward directly (O(1) mem). */ + if (rctx->switched) { + cmh_poly_fb_forward(req, rctx); + return 0; + } + + /* + * Exceeding the HW input cap: switch to the Poly1305 library + * (replaying the buffered chunks) rather than failing, then + * forward this update. + */ + if (req->nbytes > POLY_MAX_DATA - rctx->total_len) { + ret =3D cmh_poly_switch_to_fb(req); + if (ret) + goto err_free_chunks; + cmh_poly_fb_forward(req, rctx); + return 0; + } + + gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + chunk =3D kmalloc(sizeof(*chunk) + req->nbytes, gfp); + if (!chunk) { + ret =3D -ENOMEM; + goto err_free_chunks; + } + + chunk->len =3D req->nbytes; + if (req->base.flags & CRYPTO_AHASH_REQ_VIRT) + memcpy(chunk->data, req->svirt, req->nbytes); + else + scatterwalk_map_and_copy(chunk->data, req->src, + 0, req->nbytes, 0); + spin_lock_bh(&tctx->chunk_lock); + if (tctx->tfm_buffered + chunk->len > CMH_POLY_TFM_MAX_BUFFERED) { + spin_unlock_bh(&tctx->chunk_lock); + kfree_sensitive(chunk); + ret =3D -ENOMEM; + goto err_free_chunks; + } + list_add_tail(&chunk->list, &rctx->chunks); + list_add_tail(&chunk->tfm_node, &tctx->all_chunks); + tctx->tfm_buffered +=3D chunk->len; + spin_unlock_bh(&tctx->chunk_lock); + rctx->total_len +=3D req->nbytes; + return 0; + +err_free_chunks: + /* + * Terminal error -- free all previously accumulated chunks. + * Callers may not call .final() on error, so they would leak. + */ + cmh_poly_free_chunks(rctx, tctx); + return ret; +} + +static void cmh_poly_complete(void *data, int error) +{ + struct ahash_request *req =3D data; + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_poly_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cmh_poly_reqctx *rctx =3D ahash_request_ctx(req); + + if (error =3D=3D -EINPROGRESS) { + cmh_complete(&req->base, error); + return; + } + + if (rctx->total_len > 0) + cmh_dma_unmap_single(rctx->in_dma, rctx->total_len, + DMA_TO_DEVICE); + cmh_dma_unmap_single(rctx->tag_dma, POLY1305_DIGEST_SIZE, + DMA_FROM_DEVICE); + + if (!error) + memcpy(req->result, rctx->tag_buf, POLY1305_DIGEST_SIZE); + + kfree(rctx->tag_buf); + rctx->tag_buf =3D NULL; + cmh_poly_free_chunks(rctx, tctx); + kfree_sensitive(rctx->buf); + rctx->buf =3D NULL; + rctx->total_len =3D 0; + cmh_complete(&req->base, error); +} + +static int cmh_poly_final(struct ahash_request *req) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_poly_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cmh_poly_reqctx *rctx =3D ahash_request_ctx(req); + struct vcq_cmd cmds[CMH_POLY_MAX_PAYLOAD]; + struct core_dispatch d; + s32 target_mbx; + u32 core_id; + u32 idx; + int ret; + gfp_t gfp; + + /* Switched to the Poly1305 library: complete there (synchronous). */ + if (rctx->switched) { + poly1305_final(&rctx->fb_state, req->result); + return 0; + } + + if (!tctx->has_key) { + ret =3D -ENOKEY; + goto out_free_chunks; + } + + gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + + /* Linearise chunks into a single contiguous buffer for DMA */ + if (rctx->total_len > 0) { + struct cmh_poly_chunk *c; + u32 off =3D 0; + + rctx->buf =3D kmalloc(rctx->total_len, gfp); + if (!rctx->buf) { + ret =3D -ENOMEM; + goto out_free_chunks; + } + list_for_each_entry(c, &rctx->chunks, list) { + memcpy(rctx->buf + off, c->data, c->len); + off +=3D c->len; + } + } + + /* Tag output buffer */ + rctx->tag_buf =3D kzalloc(POLY1305_DIGEST_SIZE, gfp); + if (!rctx->tag_buf) { + ret =3D -ENOMEM; + goto out_free_buf; + } + + rctx->tag_dma =3D cmh_dma_map_single(rctx->tag_buf, + POLY1305_DIGEST_SIZE, + DMA_FROM_DEVICE); + if (cmh_dma_map_error(rctx->tag_dma)) { + ret =3D -ENOMEM; + goto out_free_tag; + } + + /* Map input data */ + if (rctx->total_len > 0) { + rctx->in_dma =3D cmh_dma_map_single(rctx->buf, rctx->total_len, + DMA_TO_DEVICE); + if (cmh_dma_map_error(rctx->in_dma)) { + ret =3D -ENOMEM; + goto out_unmap_tag; + } + } + + /* + * Key DMA handles are pre-mapped in setkey() and live in + * the tfm context. Use them directly for the VCQ writes. + */ + + d =3D cmh_core_select_instance(CMH_CORE_CCP); + target_mbx =3D d.mbx_idx; + core_id =3D d.core_id; + idx =3D 0; + + /* Write s_key to SYS_REF_TEMP first (bottom of stack) */ + vcq_add_sys_write(&cmds[idx++], SYS_REF_TEMP, + (u64)tctx->skey_dma, SYS_REF_NONE, + CCP_POLY_KEY_SIZE, + SYS_TYPE_SET(SYS_TYPE_FLAG_PT, CORE_ID_CCP)); + + /* Write r_key to SYS_REF_TEMP second (top of stack) */ + vcq_add_sys_write(&cmds[idx++], SYS_REF_TEMP, + (u64)tctx->rkey_dma, SYS_REF_NONE, + CCP_POLY_KEY_SIZE, + SYS_TYPE_SET(SYS_TYPE_FLAG_PT, CORE_ID_CCP)); + + /* POLY1305_INIT: rkey=3DTEMP (top), skey=3DTEMP (next) */ + vcq_add_ccp_poly_init(&cmds[idx++], core_id, SYS_REF_TEMP, + CCP_POLY_KEY_SIZE, SYS_REF_TEMP, + CCP_POLY_KEY_SIZE); + + /* FINAL: data -> tag */ + vcq_add_ccp_poly_final(&cmds[idx++], core_id, + rctx->total_len > 0 ? (u64)rctx->in_dma : 0, + (u64)rctx->tag_dma, rctx->total_len, + POLY1305_DIGEST_SIZE); + + vcq_add_flush(&cmds[idx++], core_id); + + ret =3D cmh_vcq_pack_and_submit_async(cmds, idx, rctx->packed, + CMH_POLY_MAX_PACKED, target_mbx, + cmh_poly_complete, req, + !!(req->base.flags & + CRYPTO_TFM_REQ_MAY_BACKLOG), + cmh_tm_async_timeout_jiffies()); + if (ret =3D=3D -EBUSY) + return -EBUSY; + if (ret) { + /* + * Synchronous rejection (e.g. -EAGAIN: CMQ full, no backlog). + * Free only the per-submit transients and keep the accumulated + * chunks intact so the caller can retry the identical final(). + * If no retry comes, exit_tfm reclaims the orphaned chunks; the + * per-tfm buffered-byte cap bounds how much stays pinned. + */ + if (rctx->total_len > 0) + cmh_dma_unmap_single(rctx->in_dma, rctx->total_len, + DMA_TO_DEVICE); + cmh_dma_unmap_single(rctx->tag_dma, POLY1305_DIGEST_SIZE, + DMA_FROM_DEVICE); + kfree(rctx->tag_buf); + rctx->tag_buf =3D NULL; + kfree_sensitive(rctx->buf); + rctx->buf =3D NULL; + /* Keep chunks + total_len so the retry rebuilds the buffer. */ + return ret; + } + + return -EINPROGRESS; + +out_unmap_tag: + cmh_dma_unmap_single(rctx->tag_dma, POLY1305_DIGEST_SIZE, + DMA_FROM_DEVICE); +out_free_tag: + kfree(rctx->tag_buf); +out_free_buf: + kfree_sensitive(rctx->buf); + rctx->buf =3D NULL; +out_free_chunks: + cmh_poly_free_chunks(rctx, tctx); + rctx->total_len =3D 0; + return ret; +} + +static int cmh_poly_export(struct ahash_request *req, void *out) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_poly_reqctx *rctx =3D ahash_request_ctx(req); + struct cmh_poly_export_state *state =3D out; + struct cmh_poly_chunk *chunk; + u32 offset =3D 0; + int ret; + + BUILD_BUG_ON(sizeof(struct poly1305_desc_ctx) > CMH_POLY_EXPORT_MAX); + + /* + * If more data is buffered than the flat window holds, switch to + * the Poly1305 library so a bounded, fixed-size state can be + * exported -- making export/import (clone) work at any length. + */ + if (!rctx->switched && rctx->total_len > CMH_POLY_EXPORT_MAX) { + ret =3D cmh_poly_switch_to_fb(req); + if (ret) + return ret; + } + + /* Zero the whole state buffer so no kernel memory leaks out. */ + memset(state, 0, crypto_ahash_statesize(tfm)); + + if (rctx->switched) { + state->format =3D CMH_POLY_FMT_FB; + memcpy(state->data, &rctx->fb_state, sizeof(rctx->fb_state)); + return 0; + } + + state->format =3D CMH_POLY_FMT_RAW; + state->total_len =3D rctx->total_len; + list_for_each_entry(chunk, &rctx->chunks, list) { + memcpy(state->data + offset, chunk->data, chunk->len); + offset +=3D chunk->len; + } + return 0; +} + +static int cmh_poly_import(struct ahash_request *req, const void *in) +{ + struct crypto_ahash *tfm =3D crypto_ahash_reqtfm(req); + struct cmh_poly_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cmh_poly_reqctx *rctx =3D ahash_request_ctx(req); + const struct cmh_poly_export_state *state =3D in; + struct cmh_poly_chunk *chunk; + + memset(rctx, 0, sizeof(*rctx)); + INIT_LIST_HEAD(&rctx->chunks); + + /* Fallback-format state: restore the Poly1305 library state. */ + if (state->format =3D=3D CMH_POLY_FMT_FB) { + memcpy(&rctx->fb_state, state->data, sizeof(rctx->fb_state)); + rctx->switched =3D true; + return 0; + } + + if (state->format !=3D CMH_POLY_FMT_RAW) + return -EINVAL; + + if (state->total_len > CMH_POLY_EXPORT_MAX) + return -EINVAL; + + if (state->total_len) { + chunk =3D kmalloc(sizeof(*chunk) + state->total_len, + req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC); + if (!chunk) + return -ENOMEM; + chunk->len =3D state->total_len; + memcpy(chunk->data, state->data, state->total_len); + list_add_tail(&chunk->list, &rctx->chunks); + spin_lock_bh(&tctx->chunk_lock); + list_add_tail(&chunk->tfm_node, &tctx->all_chunks); + spin_unlock_bh(&tctx->chunk_lock); + rctx->total_len =3D state->total_len; + } + return 0; +} + +static int cmh_poly_finup(struct ahash_request *req) +{ + int err; + + err =3D cmh_poly_update(req); + if (err) + return err; + return cmh_poly_final(req); +} + +static int cmh_poly_digest(struct ahash_request *req) +{ + int err; + + err =3D cmh_poly_init(req); + if (err) + return err; + return cmh_poly_finup(req); +} + +static int cmh_poly_init_tfm(struct crypto_ahash *tfm) +{ + struct cmh_poly_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + + memset(tctx, 0, sizeof(*tctx)); + spin_lock_init(&tctx->chunk_lock); + INIT_LIST_HEAD(&tctx->all_chunks); + crypto_ahash_set_reqsize(tfm, sizeof(struct cmh_poly_reqctx)); + return 0; +} + +static void cmh_poly_exit_tfm(struct crypto_ahash *tfm) +{ + struct cmh_poly_tfm_ctx *tctx =3D crypto_ahash_ctx(tfm); + struct cmh_poly_chunk *c, *tmp; + + /* Free any orphaned chunks (e.g. testmgr export/reimport poison) */ + spin_lock_bh(&tctx->chunk_lock); + list_for_each_entry_safe(c, tmp, &tctx->all_chunks, tfm_node) { + list_del(&c->tfm_node); + tctx->tfm_buffered -=3D c->len; + kfree_sensitive(c); + } + spin_unlock_bh(&tctx->chunk_lock); + + if (tctx->has_key) { + cmh_dma_unmap_single(tctx->rkey_dma, CCP_POLY_KEY_SIZE, + DMA_TO_DEVICE); + cmh_dma_unmap_single(tctx->skey_dma, CCP_POLY_KEY_SIZE, + DMA_TO_DEVICE); + } + kfree_sensitive(tctx->key); + tctx->key =3D NULL; +} + +static struct ahash_alg cmh_poly1305_alg =3D { + .init =3D cmh_poly_init, + .update =3D cmh_poly_update, + .final =3D cmh_poly_final, + .finup =3D cmh_poly_finup, + .digest =3D cmh_poly_digest, + .export =3D cmh_poly_export, + .import =3D cmh_poly_import, + .setkey =3D cmh_poly_setkey, + .init_tfm =3D cmh_poly_init_tfm, + .exit_tfm =3D cmh_poly_exit_tfm, + .halg =3D { + .digestsize =3D POLY1305_DIGEST_SIZE, + .statesize =3D CMH_POLY_STATE_SIZE, + .base =3D { + .cra_name =3D "poly1305", + .cra_driver_name =3D "rambus-cmh-poly1305", + .cra_priority =3D 300, + .cra_flags =3D CRYPTO_ALG_KERN_DRIVER_ONLY | + CRYPTO_ALG_NO_FALLBACK | + CRYPTO_ALG_ASYNC | + CRYPTO_ALG_REQ_VIRT, + .cra_blocksize =3D POLY1305_BLOCK_SIZE, + .cra_ctxsize =3D sizeof(struct cmh_poly_tfm_ctx), + .cra_module =3D THIS_MODULE, + }, + }, +}; + +/** + * cmh_ccp_poly_register() - Register Poly1305 hash algorithm with the cry= pto framework + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_ccp_poly_register(void) +{ + int ret; + + if (!cmh_core_present(CMH_CORE_CCP)) + return 0; + + ret =3D crypto_register_ahash(&cmh_poly1305_alg); + if (ret) + dev_err(cmh_dev(), "cmh_ccp_poly: failed to register poly1305 (rc=3D%d)\= n", + ret); + else + dev_dbg(cmh_dev(), "cmh_ccp_poly: registered poly1305\n"); + + return ret; +} + +/** + * cmh_ccp_poly_unregister() - Unregister Poly1305 hash algorithm from the= crypto framework + */ +void cmh_ccp_poly_unregister(void) +{ + if (!cmh_core_present(CMH_CORE_CCP)) + return; + + crypto_unregister_ahash(&cmh_poly1305_alg); + dev_dbg(cmh_dev(), "cmh_ccp_poly: unregistered poly1305\n"); +} diff --git a/drivers/crypto/cmh/cmh_main.c b/drivers/crypto/cmh/cmh_main.c index 410941cc04e4..0738b9815a11 100644 --- a/drivers/crypto/cmh/cmh_main.c +++ b/drivers/crypto/cmh/cmh_main.c @@ -38,6 +38,7 @@ #include "cmh_sm3.h" #include "cmh_aes.h" #include "cmh_sm4.h" +#include "cmh_ccp.h" #include "cmh_mgmt.h" #include "cmh_registers.h" #include "cmh_debugfs.h" @@ -264,6 +265,21 @@ static int cmh_probe(struct platform_device *pdev) if (ret) goto err_sm4_cmac_register; =20 + /* Register CCP ChaCha20 skcipher algorithm */ + ret =3D cmh_ccp_register(); + if (ret) + goto err_ccp_register; + + /* Register CCP ChaCha20-Poly1305 AEAD (RFC 7539) */ + ret =3D cmh_ccp_aead_register(); + if (ret) + goto err_ccp_aead_register; + + /* Register CCP Poly1305 shash algorithm */ + ret =3D cmh_ccp_poly_register(); + if (ret) + goto err_ccp_poly_register; + /* Register key management device (/dev/cmh_mgmt) */ ret =3D cmh_mgmt_register(); if (ret) @@ -274,6 +290,12 @@ static int cmh_probe(struct platform_device *pdev) return 0; =20 err_mgmt_register: + cmh_ccp_poly_unregister(); +err_ccp_poly_register: + cmh_ccp_aead_unregister(); +err_ccp_aead_register: + cmh_ccp_unregister(); +err_ccp_register: cmh_sm4_cmac_unregister(); err_sm4_cmac_register: cmh_sm4_aead_unregister(); @@ -322,6 +344,9 @@ static void cmh_remove(struct platform_device *pdev) cfg =3D &dev->config; =20 cmh_mgmt_unregister(); + cmh_ccp_poly_unregister(); + cmh_ccp_aead_unregister(); + cmh_ccp_unregister(); cmh_sm4_cmac_unregister(); cmh_sm4_aead_unregister(); cmh_sm4_unregister(); diff --git a/drivers/crypto/cmh/include/cmh_ccp.h b/drivers/crypto/cmh/incl= ude/cmh_ccp.h new file mode 100644 index 000000000000..363d208cbceb --- /dev/null +++ b/drivers/crypto/cmh/include/cmh_ccp.h @@ -0,0 +1,24 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- CCP Crypto API Drivers + * + * Registers CCP algorithms with the Linux crypto subsystem: + * skcipher: chacha20 + * shash: poly1305 + * aead: rfc7539(chacha20poly1305) + */ + +#ifndef CMH_CCP_H +#define CMH_CCP_H + +int cmh_ccp_register(void); +void cmh_ccp_unregister(void); + +int cmh_ccp_aead_register(void); +void cmh_ccp_aead_unregister(void); + +int cmh_ccp_poly_register(void); +void cmh_ccp_poly_unregister(void); + +#endif /* CMH_CCP_H */ --=20 2.43.7 From nobody Fri Sep 25 01:20:34 2026 Received: from BN1PR04CU002.outbound.protection.outlook.com (mail-eastus2azon11020140.outbound.protection.outlook.com [52.101.56.140]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 84F3049E14C; Thu, 17 Sep 2026 22:59:39 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.56.140 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789685982; cv=fail; b=M1mk1yzEUY4Z9rL/JDgA9oX0QMNUU6g3Z+pEeYGMHi+ORqML2rR20DV5kfx5nxk64FoAlQoRSd3heHpJbBHT+WU+OHt41pMGTlg/HCQs2cYaVJ2dhx956xkjq8jthZ6+5YhExBF7q0UkXBRTz91kj3k5l42PFY8/LMrJqQFA11Q= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789685982; c=relaxed/simple; bh=3sh2Rq1tx22uLYO9aKy0cbuniCHP37qy6UxGE3owcg0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=mrZJ29FtA2CTzBqawjrb0tLstwdS84sGcg4fgaamVQHODyr/EcNibhqQYYabAD688z0j4kJtrAWvV9ugNcYa5xwn5ZKTr3g4elT9A2mb3LCyQW/1Y8E6ybI9DoafEuCllF1we6wTb5shnYrTaN4qDx7mjeXkfPfEg0ZhcrhDpsY= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com; spf=fail smtp.mailfrom=rambus.com; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b=054H6JKJ; arc=fail smtp.client-ip=52.101.56.140 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=rambus.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b="054H6JKJ" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=DWnhb04H1sO1TpO833MNfR+vkakiqpb2cYq9uEskr/EQsPovsinhOSrTBVAylIgssUM64NaZlRGJQlMMo3B7jbs3gkOWDrDhmRmtI08fpk+NqJSNiF/tn/i7F73vKSBk21pl5k+DhUTLzRbM7Z80Z7jbGuUp0KW3UqEI6Oa5eO8ZbkJPBH5jL7d/l7OVgGuYDDRIXqkhpcMhBT+WBtzTpoDdA3lrCSBn6+ku6KbgXPjTU0vJqMMvtnH1ekIcSgqqQpfmIyjMwT099Fzft2m/DQFVnoU6SFT+DVQ1Twf3F0ykhpWscEbH78c5vOYl2wl0d7hkHzSCuKZ8lufeEtWJhg== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=CwwcBtRE1uljw0HPn8RwOtf8PSk7ACHE8ZA0Q76mNkc=; b=FELr7RwThLi+m4WDiQSi31nNNAz6OznC5khHaOT23eGB6hIm97Q/yttCpDRU6NJ79gmqiDF0ne+9u+P9gvTSayDGmOEjZgj/XUArhOeV2dU2yU6mc1QwU8mV67kbj7i5Bv+AI+AEYZcVdjhEHLBs4h97l8ktfkx4onj/6qy22x8t4jtaY2kBN9hvUGpgutxsTEPWYPTnbMbKAWQNGfn4TwWNidf8K/aR6+/SYlZM0CR2s1u5cxtVGxK+3JZJAtCn2qxd6aG/hYO7FrYnM+telDgj0rv2yVcjru3kHRifaGWeEjMbAleBBAwqcoePszNdy+Db2/Tzfbqh6tV8SjPaWw== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 192.86.86.210) smtp.rcpttodomain=cryptography.com smtp.mailfrom=rambus.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=rambus.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=rambus.com; s=selector1; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=CwwcBtRE1uljw0HPn8RwOtf8PSk7ACHE8ZA0Q76mNkc=; b=054H6JKJpPJf3qCB78v+162WGZ5cAgBB9svEdmp8te6ja7qm7WPDi0IyMx4DWXshsdfUzE7mDTpM/XCybgAPMUsU4qs1OoUM+iT+81+NHQtwyzDXIEbOIFzRn5p2nuBKzGcF4xzytV8sxdOJ1n0F9MI1pCZG/eFB+C+DSBPiTk0wWPhq7SbDOdzxF8IvP4IXcKwJWRQQbevgg6jnaJHmdg49Ps+r4tOILBhN8kpQPc0YoA6LImofWUyPFfcdrPsoBIu5fqS/xvASo89ISbzyrko7JP5NistfhgTjUyf35YhDcqalV99fGPh9glLx90JWwKShwQ9lv/GnxVpmnzBh2w== Received: from SJ0PR13CA0154.namprd13.prod.outlook.com (2603:10b6:a03:2c7::9) by BY1PR04MB9633.namprd04.prod.outlook.com (2603:10b6:a03:5b8::20) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.428.13; Thu, 17 Sep 2026 22:59:34 +0000 Received: from SJ1PEPF000026C8.namprd04.prod.outlook.com (2603:10b6:a03:2c7:cafe::78) by SJ0PR13CA0154.outlook.office365.com (2603:10b6:a03:2c7::9) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.451.8 via Frontend Transport; Thu, 17 Sep 2026 22:59:34 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 192.86.86.210) smtp.mailfrom=rambus.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=rambus.com; Received-SPF: Pass (protection.outlook.com: domain of rambus.com designates 192.86.86.210 as permitted sender) receiver=protection.outlook.com; client-ip=192.86.86.210; helo=hqxsv-psmtppxy02.rambus.com; pr=C Received: from hqxsv-psmtppxy02.rambus.com (192.86.86.210) by SJ1PEPF000026C8.mail.protection.outlook.com (10.167.244.105) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.451.8 via Frontend Transport; Thu, 17 Sep 2026 22:59:33 +0000 Received: from hqxsv-cmdev3-aousherovitch.rambus.com (hqn-lb-int-float.rambus.com [10.12.20.20]) by hqxsv-psmtppxy02.rambus.com (Postfix) with ESMTP id 4D2201801768; Thu, 17 Sep 2026 22:59:32 +0000 (UTC) From: Alex Ousherovitch To: Albert Ou , Alex Ousherovitch , Conor Dooley , "David S. Miller" , Herbert Xu , Jonathan Corbet , Krzysztof Kozlowski , Palmer Dabbelt , Paul Walmsley , Rob Herring , Saravanakrishnan Krishnamoorthy , Shuah Khan Cc: Alexandre Ghiti , devicetree@vger.kernel.org, Joel Wittenauer , linux-api@vger.kernel.org, linux-crypto@vger.kernel.org, linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-riscv@lists.infradead.org, Shuah Khan , Thi Nguyen Subject: [PATCH v5 11/19] crypto: cmh - add DRBG hwrng Date: Thu, 17 Sep 2026 15:59:20 -0700 Message-ID: <20260917225929.2494111-12-aousherovitch@rambus.com> X-Mailer: git-send-email 2.43.7 In-Reply-To: <20260917225929.2494111-1-aousherovitch@rambus.com> References: <20260917225929.2494111-1-aousherovitch@rambus.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: SJ1PEPF000026C8:EE_|BY1PR04MB9633:EE_ X-MS-Office365-Filtering-Correlation-Id: aa50b22f-9eb1-4656-7824-08df150f5713 X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|376014|7416014|36860700016|82310400026|1800799024|23010399003|10067099003|921020|6133799003|3023799007|18002099003|22082099003|56012099006|11063799006; X-Microsoft-Antispam-Message-Info: s2TKjjH/xcRXXwfbBJDfOD7RrQqCwUXVPbs/3LihN9X+brQDkmEL2NdwUPoNSM1bUcLSL9UIQEdw1M7DE8J4+Zza9QUpcHt6iTysd0DfpbONHkvTt7XQ5zo3RpJR7+K8DNA4IJANoM8lZHum9H33AtMGpr83p5U/5BcU4L3buKyp5nqYhSdb/a9eTOUh9ysx6BMkXRNkZz9lnaffW/6F2g9zZjo3g9iFH1R20PgqqGLWMHLR21wkYgMl2DXcTl/85ITmFp+sF7jnO0aAYfOeImfjzUUM6YU3a3KJ6TEwoOvs4lQB5U7b+cguG8CoZnEF0WohN+X4Dt3b+aFoVSmbEoPSp3AIa/OpshMqaMMoNSwA/mQI1ibQjBnvaz2kDpUJRMSJQoT2PxCVjHL/HlJ04ldyo6N99sSQeXwyaVe6FPTx8gyeCWtHt8TL4/dtXE/9PGibZvv5c0k2/5qhslD0vVWBr1x1uPukvG4pxKUV5JrefnB/hn+sL2IYz3RG4P0cVC7mjj05e4EDzqGJt2sLGdOVQgQFzSDZF/p4fKDKCkNGUfTAKU6px+ayG8GUXxfzbTW87WMsCDEi75IWBCV6/HXpMU5LYvmfcAIBBVhaSsKVnpn4k6ghhAmlB18YUr4xcMf/SxnDc9K2k56CsHsYNHU3BK0KKbibnfJKMRg48Kdd5FyULFLz2D8odQv40lfCpM3KxMwvVX5dprrpEFnNLf/fFuWvggd7+KvbsqxPFqQ2/vIE/ndfOjkGRWjDuMxW X-Forefront-Antispam-Report: CIP:192.86.86.210;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:hqxsv-psmtppxy02.rambus.com;PTR:InfoDomainNonexistent;CAT:NONE;SFS:(13230040)(376014)(7416014)(36860700016)(82310400026)(1800799024)(23010399003)(10067099003)(921020)(6133799003)(3023799007)(18002099003)(22082099003)(56012099006)(11063799006);DIR:OUT;SFP:1102; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: QiojtGVtn+6nCkLo7WZzZCfo061f6HIAKZ6bmGQWIYLybky+RBUabD0MP0pO6hjz6rGAO0x2AIti0BCK98liMdINN25cVkecSy2wLd5Wv6vcJRA5TXOCr1f1mNrTCgc0KOrIWtC1Rpcv2As35PeESMk7ORhLCc7vB2JTrUBTiHoCACUHF2/qBtgsXexB8r+eyxZEz8wv9Dl7BAW5yad3tJ/d8N1riCUnJoT8eEuWhNGX8NdlGhvW9T1SAA01Vwdg0fLhmTgr1ZfhRPTZmlB+0QMX3Yc52mJ5jFteK7r7jkB2xZaFgqptDmcf/C8FKWq4v0dbL9npLUs1xDihtTAKmb7LqQX15S5lnlm3uavGb2+zQ3MPLLYrXZhnF85PHDsWJT00m3sUIoyySn+sqov3PYXsXY7ar/uuCEp3FjzaCvIVHHsrctCtnJSq2/JIdUs8 X-OriginatorOrg: rambus.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 17 Sep 2026 22:59:33.8975 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: aa50b22f-9eb1-4656-7824-08df150f5713 X-MS-Exchange-CrossTenant-Id: bd0ba799-c2b9-413c-9c56-5d1731c4827c X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=bd0ba799-c2b9-413c-9c56-5d1731c4827c;Ip=[192.86.86.210];Helo=[hqxsv-psmtppxy02.rambus.com] X-MS-Exchange-CrossTenant-AuthSource: SJ1PEPF000026C8.namprd04.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: BY1PR04MB9633 Content-Type: text/plain; charset="utf-8" Register the CMH DRBG core (core ID 0x0f) as an hwrng provider. The hardware implements a NIST SP 800-90A compliant DRBG with automatic self-seeding. Signed-off-by: Alex Ousherovitch Co-developed-by: Saravanakrishnan Krishnamoorthy Signed-off-by: Saravanakrishnan Krishnamoorthy --- drivers/crypto/cmh/Makefile | 3 +- drivers/crypto/cmh/cmh_main.c | 9 ++ drivers/crypto/cmh/cmh_rng.c | 281 ++++++++++++++++++++++++++++++++++ 3 files changed, 292 insertions(+), 1 deletion(-) create mode 100644 drivers/crypto/cmh/cmh_rng.c diff --git a/drivers/crypto/cmh/Makefile b/drivers/crypto/cmh/Makefile index ef24879ba1f1..407fd2870f6d 100644 --- a/drivers/crypto/cmh/Makefile +++ b/drivers/crypto/cmh/Makefile @@ -28,7 +28,8 @@ cmh-y :=3D \ cmh_sm4_cmac.o \ cmh_ccp.o \ cmh_ccp_aead.o \ - cmh_ccp_poly.o + cmh_ccp_poly.o \ + cmh_rng.o =20 # Management ioctl device (/dev/cmh_mgmt): key lifecycle, PKE, PQC ioctls. cmh-$(CONFIG_CRYPTO_DEV_CMH_MGMT) +=3D \ diff --git a/drivers/crypto/cmh/cmh_main.c b/drivers/crypto/cmh/cmh_main.c index 0738b9815a11..b06a5ea0df4f 100644 --- a/drivers/crypto/cmh/cmh_main.c +++ b/drivers/crypto/cmh/cmh_main.c @@ -36,6 +36,7 @@ #include "cmh_cshake.h" #include "cmh_kmac.h" #include "cmh_sm3.h" +#include "cmh_rng.h" #include "cmh_aes.h" #include "cmh_sm4.h" #include "cmh_ccp.h" @@ -235,6 +236,11 @@ static int cmh_probe(struct platform_device *pdev) if (ret) goto err_sm3_register; =20 + /* Register hwrng backed by DRBG core */ + ret =3D cmh_rng_register(pdev); + if (ret) + goto err_rng_register; + /* Register AES skcipher algorithms */ ret =3D cmh_aes_register(); if (ret) @@ -308,6 +314,8 @@ static int cmh_probe(struct platform_device *pdev) err_aes_aead_register: cmh_aes_unregister(); err_aes_register: + cmh_rng_unregister(); +err_rng_register: cmh_sm3_unregister(); err_sm3_register: cmh_kmac_unregister(); @@ -353,6 +361,7 @@ static void cmh_remove(struct platform_device *pdev) cmh_aes_cmac_unregister(); cmh_aes_aead_unregister(); cmh_aes_unregister(); + cmh_rng_unregister(); cmh_sm3_unregister(); cmh_kmac_unregister(); cmh_cshake_unregister(); diff --git a/drivers/crypto/cmh/cmh_rng.c b/drivers/crypto/cmh/cmh_rng.c new file mode 100644 index 000000000000..a73634024e13 --- /dev/null +++ b/drivers/crypto/cmh/cmh_rng.c @@ -0,0 +1,281 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- Hardware RNG (DRBG) Driver + * + * Implements a Linux hwrng backed by the CMH DRBG core. Each .read() + * builds a 3-entry VCQ (header + GENERATE + FLUSH) and submits it + * synchronously through the Transaction Manager. + * + * DRBG configuration (CONFIG) is a management-host operation in the + * CMH security model. The driver attempts CONFIG at probe with the + * hardcoded ratio/strength defaults; this succeeds in stateless mode + * (any host may CONFIG) or when this host is the management host. On + * -EPERM the driver logs a notice and continues -- GENERATE works once + * the management host configures the DRBG. + * + * The management host (or any privileged user-space process) can also + * reconfigure the DRBG at runtime via CMH_IOCTL_DRBG_CONFIG. + */ + +#include +#include +#include +#include +#include +#include + +#include "cmh_rng.h" +#include "cmh_vcq.h" +#include "cmh_txn.h" +#include "cmh_dma.h" +#include "cmh_sys.h" +#include "cmh_config.h" + +/* VCQ layout for .read(): header + GENERATE + FLUSH =3D 3 entries. */ +#define DRBG_READ_VCQ_CMDS 3 + +/* VCQ layout for CONFIG: header + RESET + CONFIG + FLUSH =3D 4 entries. */ +#define DRBG_CONFIG_VCQ_CMDS 4 + +/* + * DRBG parameters -- hardcoded to production defaults. + * Entropy ratio 0 =3D 1:1 (full entropy), security strength 0x10 =3D 256-= bit. + */ +#define CMH_DRBG_ENTROPY_RATIO 0 +#define CMH_DRBG_SECURITY_STRENGTH 0x10 + +static unsigned int drbg_timeout_ms =3D 500; + +/* VCQ Builders */ + +static void vcq_add_drbg_generate(struct vcq_cmd *slot, u64 dst_phys, u32 = len) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(CORE_ID_DRBG, 0, 1, DRBG_CMD_GENERATE); + slot->hwc.drbg.cmd_generate.dst =3D dst_phys; + slot->hwc.drbg.cmd_generate.len =3D len; +} + +/* + * Maximum bytes per DRBG GENERATE request. The kernel calls .read() + * repeatedly to fill larger requests, so capping here is safe. + * 32 bytes matches the 256-bit security strength natural output size. + */ +#define CMH_DRBG_MAX_GENERATE 32U +/* Backoff before re-issuing a blocking read after a transient DRBG error.= */ +#define CMH_DRBG_RETRY_BACKOFF_MS 20U + +/* hwrng .read() callback */ + +static int cmh_rng_read(struct hwrng *rng, void *data, size_t max, bool wa= it) +{ + struct cmh_dma_orphan *orphan; + struct vcq_cmd vcq[DRBG_READ_VCQ_CMDS]; + dma_addr_t dma_addr; + void *dmabuf; + size_t nbytes; + int ret; + + if (max =3D=3D 0) + return 0; + + /* + * Our path uses GFP_KERNEL allocations and synchronous VCQ + * submission -- both may sleep. When the caller indicates + * non-blocking context (!wait), return 0 ("no data yet") so + * the hwrng core retries later. + */ + if (!wait) + return 0; + + nbytes =3D min_t(size_t, max, CMH_DRBG_MAX_GENERATE); + + orphan =3D kmalloc_obj(*orphan, GFP_KERNEL); + if (!orphan) + return -ENOMEM; + + dmabuf =3D kmalloc(nbytes, GFP_KERNEL); + if (!dmabuf) { + kfree(orphan); + return -ENOMEM; + } + + dma_addr =3D cmh_dma_map_single(dmabuf, nbytes, DMA_FROM_DEVICE); + if (cmh_dma_map_error(dma_addr)) { + kfree(dmabuf); + kfree(orphan); + return -ENOMEM; + } + + orphan->buf =3D dmabuf; + orphan->addr =3D dma_addr; + orphan->len =3D nbytes; + orphan->dir =3D DMA_FROM_DEVICE; + + vcq_set_header(&vcq[0], DRBG_READ_VCQ_CMDS); + vcq_add_drbg_generate(&vcq[1], dma_addr, nbytes); + vcq_add_flush(&vcq[2], CORE_ID_DRBG); + + /* + * Use the noabort variant: if the MBX is occupied by a slow + * operation (e.g. SLH-DSA sign at 120 s), we must not issue + * MBX_COMMAND_ABORT -- that would kill the unrelated in-flight + * VCQ. On timeout with an in-flight VCQ (-EINPROGRESS), the + * orphan callback defers DMA cleanup until the RH fires. + */ + ret =3D cmh_tm_submit_sync_noabort(vcq, DRBG_READ_VCQ_CMDS, 1, + msecs_to_jiffies(drbg_timeout_ms), + cmh_dma_orphan_free, orphan); + if (ret =3D=3D -EINPROGRESS) { + /* + * The orphan callback owns dmabuf and frees it on VCQ + * completion. Return 0 (not -EAGAIN): .read() only runs with + * wait=3Dtrue (see the !wait early return above), and the hwrng + * core forwards a negative errno straight to a blocking read + * whereas a 0 return makes it retry. + */ + return 0; + } + + /* Normal path or cancelled-from-queue: caller owns DMA */ + cmh_dma_unmap_single(dma_addr, nbytes, DMA_FROM_DEVICE); + kfree(orphan); + + if (ret) { + /* + * .read() only runs with wait=3Dtrue (see the !wait early + * return above). For known transient conditions return 0 so + * the hwrng core retries the blocking read; a negative errno + * here would be forwarded to userspace on a blocking fd + * (e.g. -EAGAIN violates POSIX). Propagate genuinely + * unexpected failures so real faults are not masked into an + * indefinite retry loop. + */ + switch (ret) { + case -EAGAIN: + case -EBUSY: + case -ETIMEDOUT: + case -EIO: + /* + * -ENODEV: the TM is not running -- occurs when the + * hwrng kthread (PF_NOFREEZE, not frozen during + * suspend) calls .read() while the device is suspended. + * Treat as transient: the TM restarts on resume. + */ + case -ENODEV: + dev_dbg_ratelimited(cmh_dev(), + "rng: transient DRBG failure (rc=3D%d)\n", + ret); + kfree_sensitive(dmabuf); + /* + * Back off before the hwrng core re-issues the + * blocking read: rng_dev_read() loops on a 0 return + * with only a conditional need_resched, so a persistent + * transient fault would otherwise spin the CPU. + */ + msleep(CMH_DRBG_RETRY_BACKOFF_MS); + return 0; + default: + dev_err_ratelimited(cmh_dev(), + "rng: DRBG generate failed (rc=3D%d)\n", + ret); + kfree_sensitive(dmabuf); + return ret; + } + } + + memcpy(data, dmabuf, nbytes); + kfree_sensitive(dmabuf); + + return nbytes; +} + +/* Registration */ + +static bool cmh_rng_registered; + +static struct hwrng cmh_hwrng =3D { + .name =3D "rambus-cmh-drbg", + .read =3D cmh_rng_read, +}; + +/** + * cmh_rng_register() - Register the CMH hardware RNG device + * @pdev: Platform device for the CMH accelerator + * + * Attempt a DRBG CONFIG VCQ (best effort; -EPERM when not the management + * host is non-fatal), then register the hwrng device with the kernel + * hwrng framework. + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_rng_register(struct platform_device *pdev) +{ + struct vcq_cmd cfg_vcq[DRBG_CONFIG_VCQ_CMDS]; + int ret; + + /* + * The hwrng core elevates a zero quality to full trust (1024) for + * a hardware RNG, so cmh_hwrng.quality is left at 0. Set it in the + * hwrng initializer to lower the entropy estimate if a platform + * requires it. + * + * DRBG CONFIG is a management-host operation. Attempt it: it + * succeeds in stateless mode (any host) or when we are the + * management host. On -EPERM (not the management host) continue + * without error -- GENERATE works once the management host has + * configured the DRBG. + */ + vcq_set_header(&cfg_vcq[0], DRBG_CONFIG_VCQ_CMDS); + vcq_add_drbg_reset(&cfg_vcq[1]); + vcq_add_drbg_config(&cfg_vcq[2], CMH_DRBG_ENTROPY_RATIO, + CMH_DRBG_SECURITY_STRENGTH); + vcq_add_flush(&cfg_vcq[3], CORE_ID_DRBG); + ret =3D cmh_tm_submit_sync(cfg_vcq, DRBG_CONFIG_VCQ_CMDS, 1); + if (ret =3D=3D -EPERM) + dev_notice(&pdev->dev, + "rng: DRBG config not permitted (not management host); assuming exte= rnal configuration\n"); + else if (ret) + dev_warn(&pdev->dev, + "rng: DRBG config failed (rc=3D%d)\n", ret); + + ret =3D hwrng_register(&cmh_hwrng); + if (ret) { + dev_err(&pdev->dev, "rng: hwrng_register failed (rc=3D%d)\n", + ret); + return ret; + } + + cmh_rng_registered =3D true; + return 0; +} + +/** + * cmh_rng_unregister() - Unregister the CMH hardware RNG device + * + * Unregisters the hwrng device from the kernel hwrng framework if it + * was previously registered. + */ +void cmh_rng_unregister(void) +{ + if (!cmh_rng_registered) + return; + hwrng_unregister(&cmh_hwrng); + cmh_rng_registered =3D false; +} + +/* -- debugfs timeout accessor ------------------------------------------ = */ + +#ifdef CONFIG_CRYPTO_DEV_CMH_DEBUG +/** + * cmh_rng_timeout_drbg_ptr() - Return pointer to drbg_timeout_ms for debu= gfs + * + * Exposes the DRBG operation timeout for runtime tuning via debugfs + * config/ directory. + * + * Return: pointer to the static drbg_timeout_ms variable. + */ +unsigned int *cmh_rng_timeout_drbg_ptr(void) { return &drbg_timeout_ms; } +#endif --=20 2.43.7 From nobody Fri Sep 25 01:20:34 2026 Received: from SA9PR02CU001.outbound.protection.outlook.com (mail-southcentralusazon11023117.outbound.protection.outlook.com [40.93.196.117]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C9CAC4B0E56; Thu, 17 Sep 2026 22:59:42 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=40.93.196.117 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789685989; cv=fail; b=RjxBdDqKbfkXprf3sbNgEYjAQYXaB1ry4LkkMvPXFQ6bL9BJ2TEjTgj6nC8te5GreRJ3dm31DgSQv7QiYtj3fk+81xVtqXFqsMwpAHi0sVg52tqS21mOnb07SiN7BK6lK9aXrCxvzIK3Ocu41QLg5Lff+ttotpUohDjtkH1hZco= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789685989; c=relaxed/simple; bh=aYmyIc7yrrJVuH9cAcHmUByL0mtthYRNnrmrvIX+jow=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=XZG18EG/dK2HEVAkC7FE+pWeuM7EMF/oKVWvLW+m1tJZT4Q0BBb4tQ46+l3MxMzjEJqCoan93j2/kz5bACRX8+LUBRrJsgzYC0oPmqDGIY4ipi7WZ+nottQoL8ydtrTjwP8meYZ7OHyd1wgFg+kdt/x/E78biAQG2Fz+3G5gry8= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com; spf=fail smtp.mailfrom=rambus.com; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b=1dNICHlm; arc=fail smtp.client-ip=40.93.196.117 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=rambus.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b="1dNICHlm" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=pwkwSbfkfaEK3K/hJz4kjEAnWxF37h5XAMp+48cOHbasToLmrfQ415ln2w6H8omkhzJCUar6drt8igdhiXXy8V1F7c9sTZyY6NOJeO3S3yWQ3aUwV3ydSACG/Jzi05HfAhGOE703bMfpxUSxoK1/jB99doraqSXrbZKMH+qepy1Dn3xyTOiCjR/zcthwdy+tcTyPByT1I4071Dhy44Rn838MN0ZyaKFuWHMIrsTGNBieLwRGTUgfz9KL4ynvhyWZOm9XzmSyae3wXEOTacgXhygNcsVBkLqDuAKgaFbJ3DR0eVv2gpl9DLWloIKcJJMaZgwFMqj1uZybRCYej0YtUA== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=z4g/088oTrSVBi4FAMzBlamE8qDbpanGfBDCx6o3TrQ=; b=Rhph00IE7i7UYZLo2TJpv+LSMgWYRibKw8Yb7kPGYP2pjZp+fCr74Ktd/ep+WfJKU7vMJ0mUtNFX5yXu3FMgPkkDYKzTPNuBIfDgt6gIjFfHOip85af519kfSyiWPIBcdy4I1ymOBxK8zqKynn/c001g5Ce7+zmLsP55ZgbO0r1hp/emXm8VgXCjJxu7+AOMQ549pzoGlV/N+ENQ6kdZl8uoAB0LiWKNIZbP4LxVnkL92XJmy9QHvEtzvIR1TMUgjihWmA4simZh3YEeAsR0TgACJXo38TJfO3ezraSsQPCOxeF2FeSIRUEjsQbZfJa42XkZptCyJSLxZTHSFKMOpA== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 192.86.86.210) smtp.rcpttodomain=cryptography.com smtp.mailfrom=rambus.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=rambus.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=rambus.com; s=selector1; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=z4g/088oTrSVBi4FAMzBlamE8qDbpanGfBDCx6o3TrQ=; b=1dNICHlm5kWLqo9JkLNhIXxTNMsQO3AONnbZDWFka706koJ3paQxJ5jRhDMyr19ADI0/UpLJ1uRnoj723/ooDbakmEmC8UHIMYzWpW66AAYi5LznphC7gUHgL9lin4Eap678my8tN+h10HuL7eDlSUtGw8riwOiu3aMAuUjj5TwbmdInLnsQgBmfC7BgFh8cxaG3UTnxfVHqydUq7pi7X9gnpbiEROIKx6FxQpBF3KTEGz8ml+KtaAH1TbTVhHn7OZPfd7J28is9VfCZjPaYBCeNVnlePRWXQGZOu0X8zBz/L8GVYEdQD46E13x7e57Fn6lbjZikqzTR6Fbv4iHZdA== Received: from SJ0PR13CA0064.namprd13.prod.outlook.com (2603:10b6:a03:2c4::9) by CH0PR04MB8179.namprd04.prod.outlook.com (2603:10b6:610:fc::17) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.428.12; Thu, 17 Sep 2026 22:59:34 +0000 Received: from SJ1PEPF000026C7.namprd04.prod.outlook.com (2603:10b6:a03:2c4:cafe::96) by SJ0PR13CA0064.outlook.office365.com (2603:10b6:a03:2c4::9) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.451.8 via Frontend Transport; Thu, 17 Sep 2026 22:59:34 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 192.86.86.210) smtp.mailfrom=rambus.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=rambus.com; Received-SPF: Pass (protection.outlook.com: domain of rambus.com designates 192.86.86.210 as permitted sender) receiver=protection.outlook.com; client-ip=192.86.86.210; helo=hqxsv-psmtppxy02.rambus.com; pr=C Received: from hqxsv-psmtppxy02.rambus.com (192.86.86.210) by SJ1PEPF000026C7.mail.protection.outlook.com (10.167.244.104) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.451.8 via Frontend Transport; Thu, 17 Sep 2026 22:59:33 +0000 Received: from hqxsv-cmdev3-aousherovitch.rambus.com (hqn-lb-int-float.rambus.com [10.12.20.20]) by hqxsv-psmtppxy02.rambus.com (Postfix) with ESMTP id 56D9D1801769; Thu, 17 Sep 2026 22:59:32 +0000 (UTC) From: Alex Ousherovitch To: Albert Ou , Alex Ousherovitch , Conor Dooley , "David S. Miller" , Herbert Xu , Jonathan Corbet , Krzysztof Kozlowski , Palmer Dabbelt , Paul Walmsley , Rob Herring , Saravanakrishnan Krishnamoorthy , Shuah Khan Cc: Alexandre Ghiti , devicetree@vger.kernel.org, Joel Wittenauer , linux-api@vger.kernel.org, linux-crypto@vger.kernel.org, linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-riscv@lists.infradead.org, Shuah Khan , Thi Nguyen Subject: [PATCH v5 12/19] crypto: cmh - add RSA akcipher Date: Thu, 17 Sep 2026 15:59:21 -0700 Message-ID: <20260917225929.2494111-13-aousherovitch@rambus.com> X-Mailer: git-send-email 2.43.7 In-Reply-To: <20260917225929.2494111-1-aousherovitch@rambus.com> References: <20260917225929.2494111-1-aousherovitch@rambus.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: SJ1PEPF000026C7:EE_|CH0PR04MB8179:EE_ X-MS-Office365-Filtering-Correlation-Id: 7cc00dcb-632c-4d50-683c-08df150f56fa X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|7416014|82310400026|376014|36860700016|23010399003|1800799024|921020|18002099003|22082099003|6133799003|56012099006|11063799006|10067099003|3023799007; X-Microsoft-Antispam-Message-Info: YQNggozaMcf3dFv8DnbDbeBp0kWEll5iIrhVvie91CR7ix12E3HTEWdTRjr+YUt2FA5L9RDh1dzPWwtJ7iB9XtpC0ndyO5HlDb/kTgg444MtxQ3kqB00F+ZSJUpYR8XitsHk6TUuifSwXDRiRoJRZIczTNVlaTxaGW8NGCD050LDzTEoMFkCsteXnGj6GzYcvd7qjJjRNEDzOVz9VVbvyWnR+Uq1tsefv8gHcEhh1KF0usgohSEutuFB/UoRRWwaVCUGGaZOfVFRwuxvs6yHyVV0P3YwWoCebsHbfIsctZQW7Of11vRfTVbcC/utA9LscnSUHKdLO4OyC9wyPsCiKgvb8OktR06fZLwOQ+YQaIL97qiV+4pezQ6xPYgIEOMImiFA80M/WBo4YW76FFkE1sV0Zr+/a29oN4OkHXTfUc5x6EjC9MQzT6mpwYlhZo5C3BSek7ACpsIYT2oG9Y5a44ffif8qV7mo0MLimEk5Vtqkc70pVtdHlJcIX0o0FVQ7Lgf2By/bcrATApeLCwHaEcgb74llAutpTyw1RmPFksRs3EgJ58RGllz882bf2M41xbrB2tGdQnhq3l5lqapMaLvswrvxXQuuw2+Sqnmubxf8X0xeu60bWWMIjm0Vuo/vZYM1yX8kjEeOj8vsNRZGU7aaeODpqG7RTSQmYLSw42DMsrK+BPrEd4840xLbUL2Gop5B8kfmRO3qQo+J0EjxeZaVDDrEf/8uBMuuXbSxtvGH2t7kcNGV0mL5utNkbJdM X-Forefront-Antispam-Report: CIP:192.86.86.210;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:hqxsv-psmtppxy02.rambus.com;PTR:InfoDomainNonexistent;CAT:NONE;SFS:(13230040)(7416014)(82310400026)(376014)(36860700016)(23010399003)(1800799024)(921020)(18002099003)(22082099003)(6133799003)(56012099006)(11063799006)(10067099003)(3023799007);DIR:OUT;SFP:1102; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: GtCbhnGgPG+wLQ/6b6BRpfOqXdtK4xViWKKhgSU6p7XhgI73BzsCrVmMPn9motj6TFjSVkPEV9kn/q8hqSwPa5sooEDZKiqS1sXz7gaZKwI3oWJKnJzxfoq9kv6pzJZqJuo555stusVnE0bgFEkHflEsx+yYSbI9i6OIvqCKypw57W6IaCZsvEB8KrCb76MJAJAa0hiu1j8leuXaK/hS1s3ndivFMng70gUEi0QlIqA2FpUL+2x0+tGQnGhKZsKrwMIDiWJx7uqNsQ9T1aJO1ZsnCIJ3rCSrzO7mrOwy7lQRAQB/24I6/HUYMCSr8cQNF9WbxYNKd4T9HFGBb7W36Q86SA1sz6D1z43/7oj1pDewMi98iNfn8VzXbwLgCYYP9TK/lf5NM1xYn4hR7Mh+fynHvnlJbVGWfrJp4fj59nyIc0YhG6nwOPUmvTZPR6yT X-OriginatorOrg: rambus.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 17 Sep 2026 22:59:33.7383 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: 7cc00dcb-632c-4d50-683c-08df150f56fa X-MS-Exchange-CrossTenant-Id: bd0ba799-c2b9-413c-9c56-5d1731c4827c X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=bd0ba799-c2b9-413c-9c56-5d1731c4827c;Ip=[192.86.86.210];Helo=[hqxsv-psmtppxy02.rambus.com] X-MS-Exchange-CrossTenant-AuthSource: SJ1PEPF000026C7.namprd04.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: CH0PR04MB8179 Content-Type: text/plain; charset="utf-8" Register the RSA akcipher algorithm using the CMH PKE core (core ID 0x0a). Supports encrypt, decrypt, sign, and verify operations with 2048, 3072, and 4096-bit keys. 512- and 1024-bit keys are also accepted for legacy/test interoperability. Includes common PKE helpers shared by subsequent ECDSA and ECDH patches. Signed-off-by: Alex Ousherovitch Co-developed-by: Saravanakrishnan Krishnamoorthy Signed-off-by: Saravanakrishnan Krishnamoorthy --- drivers/crypto/cmh/Makefile | 4 +- drivers/crypto/cmh/cmh_main.c | 9 + drivers/crypto/cmh/cmh_pke_common.c | 578 +++++++++++++++++++++++++ drivers/crypto/cmh/cmh_pke_rsa.c | 644 ++++++++++++++++++++++++++++ 4 files changed, 1234 insertions(+), 1 deletion(-) create mode 100644 drivers/crypto/cmh/cmh_pke_common.c create mode 100644 drivers/crypto/cmh/cmh_pke_rsa.c diff --git a/drivers/crypto/cmh/Makefile b/drivers/crypto/cmh/Makefile index 407fd2870f6d..cdbcc8cdac5f 100644 --- a/drivers/crypto/cmh/Makefile +++ b/drivers/crypto/cmh/Makefile @@ -29,7 +29,9 @@ cmh-y :=3D \ cmh_ccp.o \ cmh_ccp_aead.o \ cmh_ccp_poly.o \ - cmh_rng.o + cmh_rng.o \ + cmh_pke_common.o \ + cmh_pke_rsa.o =20 # Management ioctl device (/dev/cmh_mgmt): key lifecycle, PKE, PQC ioctls. cmh-$(CONFIG_CRYPTO_DEV_CMH_MGMT) +=3D \ diff --git a/drivers/crypto/cmh/cmh_main.c b/drivers/crypto/cmh/cmh_main.c index b06a5ea0df4f..23346bbf11b0 100644 --- a/drivers/crypto/cmh/cmh_main.c +++ b/drivers/crypto/cmh/cmh_main.c @@ -40,6 +40,7 @@ #include "cmh_aes.h" #include "cmh_sm4.h" #include "cmh_ccp.h" +#include "cmh_pke.h" #include "cmh_mgmt.h" #include "cmh_registers.h" #include "cmh_debugfs.h" @@ -286,6 +287,11 @@ static int cmh_probe(struct platform_device *pdev) if (ret) goto err_ccp_poly_register; =20 + /* Register PKE RSA akcipher */ + ret =3D cmh_pke_rsa_register(); + if (ret) + goto err_pke_rsa_register; + /* Register key management device (/dev/cmh_mgmt) */ ret =3D cmh_mgmt_register(); if (ret) @@ -296,6 +302,8 @@ static int cmh_probe(struct platform_device *pdev) return 0; =20 err_mgmt_register: + cmh_pke_rsa_unregister(); +err_pke_rsa_register: cmh_ccp_poly_unregister(); err_ccp_poly_register: cmh_ccp_aead_unregister(); @@ -352,6 +360,7 @@ static void cmh_remove(struct platform_device *pdev) cfg =3D &dev->config; =20 cmh_mgmt_unregister(); + cmh_pke_rsa_unregister(); cmh_ccp_poly_unregister(); cmh_ccp_aead_unregister(); cmh_ccp_unregister(); diff --git a/drivers/crypto/cmh/cmh_pke_common.c b/drivers/crypto/cmh/cmh_p= ke_common.c new file mode 100644 index 000000000000..ab3e2eb7d3f8 --- /dev/null +++ b/drivers/crypto/cmh/cmh_pke_common.c @@ -0,0 +1,578 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- PKE Common VCQ Builders + * + * VCQ builder functions for all PKE core commands. Each builder + * populates a single vcq_cmd slot with the appropriate magic, + * command ID, byte-swap flags, and command-specific payload. + * + * RSA commands always use PKE_SWAP_FLAGS (VCQ_FLAG_SWAP_BYTES | + * VCQ_FLAG_SWAP_WORDS). EC Weierstrass curves (NIST P-*, Brainpool, + * secp256k1, SM2) use PKE_SWAP_FLAGS; Edwards curves (25519, 448) + * use no swap flags. SM2 commands use per-command flags documented + * in the eSW ABI. + * + * Callers combine these with vcq_set_header() + vcq_add_flush() + * and submit via cmh_tm_submit_sync(). + */ + +#include + +#include "cmh_pke.h" + +/** + * vcq_add_pke_flush() - Add a PKE flush command to a VCQ slot + * @slot: VCQ command slot to populate + * @core_id: PKE hardware core ID + * + * Populates @slot with a flush command for the specified PKE core. + */ +void vcq_add_pke_flush(struct vcq_cmd *slot, u32 core_id) +{ + vcq_add_flush(slot, core_id); +} + +/* RSA */ + +/** + * vcq_add_pke_rsa_enc() - Build a VCQ command for RSA public-key encrypti= on + * @slot: VCQ command slot to populate + * @core_id: PKE hardware core ID + * @bits: RSA key size in bits + * @e_len: Length of the public exponent in bytes + * @e_dma: DMA address of public exponent buffer + * @n_dma: DMA address of modulus buffer + * @m_dma: DMA address of plaintext message buffer + * @c_dma: DMA address of ciphertext output buffer + * @flags: VCQ command flags + */ +void vcq_add_pke_rsa_enc(struct vcq_cmd *slot, u32 core_id, u32 bits, u32 = e_len, + u64 e_dma, u64 n_dma, u64 m_dma, u64 c_dma, + u32 flags) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, flags, 1, PKE_CMD_RSA_ENC); + slot->hwc.pke.cmd_rsa_enc.bits =3D bits; + slot->hwc.pke.cmd_rsa_enc.e_len =3D e_len; + slot->hwc.pke.cmd_rsa_enc.e =3D e_dma; + slot->hwc.pke.cmd_rsa_enc.n =3D n_dma; + slot->hwc.pke.cmd_rsa_enc.m =3D m_dma; + slot->hwc.pke.cmd_rsa_enc.c =3D c_dma; +} + +/** + * vcq_add_pke_rsa_dec() - Build a VCQ command for RSA private-key decrypt= ion + * @slot: VCQ command slot to populate + * @core_id: PKE hardware core ID + * @bits: RSA key size in bits + * @e_len: Length of the public exponent in bytes + * @e_dma: DMA address of public exponent buffer + * @n_dma: DMA address of modulus buffer + * @c_dma: DMA address of ciphertext input buffer + * @m_dma: DMA address of plaintext output buffer + * @d_ref: Datastore reference for the private exponent + * @flags: VCQ command flags + */ +void vcq_add_pke_rsa_dec(struct vcq_cmd *slot, u32 core_id, u32 bits, u32 = e_len, + u64 e_dma, u64 n_dma, u64 c_dma, u64 m_dma, + u64 d_ref, u32 flags) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, flags, 1, PKE_CMD_RSA_DEC); + slot->hwc.pke.cmd_rsa_dec.bits =3D bits; + slot->hwc.pke.cmd_rsa_dec.e_len =3D e_len; + slot->hwc.pke.cmd_rsa_dec.e =3D e_dma; + slot->hwc.pke.cmd_rsa_dec.n =3D n_dma; + slot->hwc.pke.cmd_rsa_dec.c =3D c_dma; + slot->hwc.pke.cmd_rsa_dec.m =3D m_dma; + slot->hwc.pke.cmd_rsa_dec.d =3D d_ref; +} + +/** + * vcq_add_pke_rsa_crt_dec() - Build a VCQ command for RSA-CRT decryption + * @slot: VCQ command slot to populate + * @core_id: PKE hardware core ID + * @bits: RSA key size in bits + * @e_len: Length of the public exponent in bytes + * @e_dma: DMA address of public exponent buffer + * @n_dma: DMA address of modulus buffer + * @c_dma: DMA address of ciphertext input buffer + * @m_dma: DMA address of plaintext output buffer + * @crt_ref: Datastore reference for CRT private key components + * @flags: VCQ command flags + */ +void vcq_add_pke_rsa_crt_dec(struct vcq_cmd *slot, u32 core_id, u32 bits, = u32 e_len, + u64 e_dma, u64 n_dma, u64 c_dma, u64 m_dma, + u64 crt_ref, u32 flags) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, flags, 1, PKE_CMD_RSA_CRT_DEC); + slot->hwc.pke.cmd_rsa_crt_dec.bits =3D bits; + slot->hwc.pke.cmd_rsa_crt_dec.e_len =3D e_len; + slot->hwc.pke.cmd_rsa_crt_dec.e =3D e_dma; + slot->hwc.pke.cmd_rsa_crt_dec.n =3D n_dma; + slot->hwc.pke.cmd_rsa_crt_dec.c =3D c_dma; + slot->hwc.pke.cmd_rsa_crt_dec.m =3D m_dma; + slot->hwc.pke.cmd_rsa_crt_dec.crt =3D crt_ref; +} + +/* ECDSA */ + +/** + * vcq_add_pke_ecdsa_verify() - Build a VCQ command for ECDSA signature ve= rification + * @slot: VCQ command slot to populate + * @core_id: PKE hardware core ID + * @curve: Curve identifier (e.g. NIST P-256, P-384, P-521) + * @dlen: Digest length in bytes + * @pk_dma: DMA address of public key buffer + * @dig_dma: DMA address of digest buffer + * @sig_dma: DMA address of signature buffer + * @rp_dma: DMA address of r-prime verification output buffer + * @flags: VCQ command flags + */ +void vcq_add_pke_ecdsa_verify(struct vcq_cmd *slot, u32 core_id, u32 curve= , u32 dlen, + u64 pk_dma, u64 dig_dma, u64 sig_dma, + u64 rp_dma, u32 flags) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, flags, 1, PKE_CMD_ECDSA_VERIFY); + slot->hwc.pke.cmd_ecdsa_verify.curve =3D curve; + slot->hwc.pke.cmd_ecdsa_verify.digest_len =3D dlen; + slot->hwc.pke.cmd_ecdsa_verify.public_key =3D pk_dma; + slot->hwc.pke.cmd_ecdsa_verify.digest =3D dig_dma; + slot->hwc.pke.cmd_ecdsa_verify.signature =3D sig_dma; + slot->hwc.pke.cmd_ecdsa_verify.rprime =3D rp_dma; +} + +/** + * vcq_add_pke_ecdsa_sign() - Build a VCQ command for ECDSA signing + * @slot: VCQ command slot to populate + * @core_id: PKE hardware core ID + * @curve: Curve identifier (e.g. NIST P-256, P-384, P-521) + * @sklen: Secret key length in bytes + * @dig_dma: DMA address of digest buffer + * @sig_dma: DMA address of signature output buffer + * @sk_ref: Datastore reference for the secret key + * @dlen: Digest length in bytes + * @flags: VCQ command flags + */ +void vcq_add_pke_ecdsa_sign(struct vcq_cmd *slot, u32 core_id, u32 curve, = u32 sklen, + u64 dig_dma, u64 sig_dma, u64 sk_ref, + u32 dlen, u32 flags) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, flags, 1, PKE_CMD_ECDSA_SIGN); + slot->hwc.pke.cmd_ecdsa_sign.curve =3D curve; + slot->hwc.pke.cmd_ecdsa_sign.secret_key_len =3D sklen; + slot->hwc.pke.cmd_ecdsa_sign.digest =3D dig_dma; + slot->hwc.pke.cmd_ecdsa_sign.signature =3D sig_dma; + slot->hwc.pke.cmd_ecdsa_sign.secret_key =3D sk_ref; + slot->hwc.pke.cmd_ecdsa_sign.digest_len =3D dlen; +} + +/** + * vcq_add_pke_ecdsa_pubgen() - Build a VCQ command for ECDSA public key g= eneration + * @slot: VCQ command slot to populate + * @core_id: PKE hardware core ID + * @curve: Curve identifier (e.g. NIST P-256, P-384, P-521) + * @sklen: Secret key length in bytes + * @pk_dma: DMA address of public key output buffer + * @sk_ref: Datastore reference for the secret key + * @flags: VCQ command flags + * + * Generates the public key from an existing private key stored in the + * datastore. + */ +void vcq_add_pke_ecdsa_pubgen(struct vcq_cmd *slot, u32 core_id, u32 curve= , u32 sklen, + u64 pk_dma, u64 sk_ref, u32 flags) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, flags, 1, PKE_CMD_ECDSA_PUBGEN); + slot->hwc.pke.cmd_ecdsa_pubgen.curve =3D curve; + slot->hwc.pke.cmd_ecdsa_pubgen.secret_key_len =3D sklen; + slot->hwc.pke.cmd_ecdsa_pubgen.public_key =3D pk_dma; + slot->hwc.pke.cmd_ecdsa_pubgen.secret_key =3D sk_ref; +} + +/** + * vcq_add_pke_ecdsa_keygen() - Build a VCQ command for ECDSA key pair gen= eration + * @slot: VCQ command slot to populate + * @core_id: PKE hardware core ID + * @curve: Curve identifier (e.g. NIST P-256, P-384, P-521) + * @sklen: Secret key length in bytes + * @sk_ref: Datastore reference for the generated secret key + * @sk_type: Datastore type for the secret key object + * @flags: VCQ command flags + */ +void vcq_add_pke_ecdsa_keygen(struct vcq_cmd *slot, u32 core_id, u32 curve= , u32 sklen, + u64 sk_ref, u32 sk_type, u32 flags) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, flags, 1, PKE_CMD_ECDSA_KEYGEN); + slot->hwc.pke.cmd_ecdsa_keygen.curve =3D curve; + slot->hwc.pke.cmd_ecdsa_keygen.secret_key_len =3D sklen; + slot->hwc.pke.cmd_ecdsa_keygen.secret_key =3D sk_ref; + slot->hwc.pke.cmd_ecdsa_keygen.secret_key_type =3D sk_type; +} + +/* ECDH */ + +/** + * vcq_add_pke_ecdh_keygen() - Build a VCQ command for ECDH key pair gener= ation + * @slot: VCQ command slot to populate + * @core_id: PKE hardware core ID + * @curve: Curve identifier (e.g. NIST P-256, P-384, P-521, X25519, X448) + * @sklen: Secret key length in bytes + * @pkx_dma: DMA address of public key X-coordinate output buffer + * @sk_ref: Datastore reference for the generated secret key + * @flags: VCQ command flags + */ +void vcq_add_pke_ecdh_keygen(struct vcq_cmd *slot, u32 core_id, u32 curve,= u32 sklen, + u64 pkx_dma, u64 sk_ref, u32 flags) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, flags, 1, PKE_CMD_ECDH_KEYGEN); + slot->hwc.pke.cmd_ecdh_keygen.curve =3D curve; + slot->hwc.pke.cmd_ecdh_keygen.secret_key_len =3D sklen; + slot->hwc.pke.cmd_ecdh_keygen.public_key_x =3D pkx_dma; + slot->hwc.pke.cmd_ecdh_keygen.secret_key =3D sk_ref; +} + +/** + * vcq_add_pke_ecdh() - Build a VCQ command for ECDH shared secret computa= tion + * @slot: VCQ command slot to populate + * @core_id: PKE hardware core ID + * @curve: Curve identifier (e.g. NIST P-256, P-384, P-521, X25519, X448) + * @sklen: Secret key length in bytes + * @sslen: Shared secret length in bytes + * @ss_type: Datastore type for the shared secret object + * @peer_dma: DMA address of peer public key buffer + * @sk_ref: Datastore reference for the local secret key + * @ss_ref: Datastore reference for the computed shared secret + * @flags: VCQ command flags + */ +void vcq_add_pke_ecdh(struct vcq_cmd *slot, u32 core_id, u32 curve, u32 sk= len, + u32 sslen, u32 ss_type, u64 peer_dma, u64 sk_ref, + u64 ss_ref, u32 flags) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, flags, 1, PKE_CMD_ECDH); + slot->hwc.pke.cmd_ecdh.curve =3D curve; + slot->hwc.pke.cmd_ecdh.secret_key_len =3D sklen; + slot->hwc.pke.cmd_ecdh.shared_secret_len =3D sslen; + slot->hwc.pke.cmd_ecdh.shared_secret_type =3D ss_type; + slot->hwc.pke.cmd_ecdh.peer_key_x =3D peer_dma; + slot->hwc.pke.cmd_ecdh.secret_key =3D sk_ref; + slot->hwc.pke.cmd_ecdh.shared_secret =3D ss_ref; +} + +/* EdDSA */ + +/** + * vcq_add_pke_eddsa_verify() - Build a VCQ command for EdDSA signature ve= rification + * @slot: VCQ command slot to populate + * @core_id: PKE hardware core ID + * @curve: Curve identifier (Ed25519 or Ed448) + * @dlen: Digest (message) length in bytes + * @pky_dma: DMA address of public key Y-coordinate buffer + * @dig_dma: DMA address of digest buffer + * @sig_dma: DMA address of signature buffer + * @rp_dma: DMA address of r-prime verification output buffer + * @flags: VCQ command flags + */ +void vcq_add_pke_eddsa_verify(struct vcq_cmd *slot, u32 core_id, u32 curve= , u32 dlen, + u64 pky_dma, u64 dig_dma, u64 sig_dma, + u64 rp_dma, u32 flags) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, flags, 1, PKE_CMD_EDDSA_VERIFY); + slot->hwc.pke.cmd_eddsa_verify.curve =3D curve; + slot->hwc.pke.cmd_eddsa_verify.digest_len =3D dlen; + slot->hwc.pke.cmd_eddsa_verify.public_key_y =3D pky_dma; + slot->hwc.pke.cmd_eddsa_verify.digest =3D dig_dma; + slot->hwc.pke.cmd_eddsa_verify.signature =3D sig_dma; + slot->hwc.pke.cmd_eddsa_verify.rprime =3D rp_dma; +} + +/** + * vcq_add_pke_eddsa_sign() - Build a VCQ command for EdDSA signing + * @slot: VCQ command slot to populate + * @core_id: PKE hardware core ID + * @curve: Curve identifier (Ed25519 or Ed448) + * @sklen: Secret key length in bytes + * @dig_dma: DMA address of digest (message) buffer + * @sig_dma: DMA address of signature output buffer + * @sk_ref: Datastore reference for the secret key + * @dlen: Digest (message) length in bytes + * @flags: VCQ command flags + */ +void vcq_add_pke_eddsa_sign(struct vcq_cmd *slot, u32 core_id, u32 curve, = u32 sklen, + u64 dig_dma, u64 sig_dma, u64 sk_ref, + u32 dlen, u32 flags) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, flags, 1, PKE_CMD_EDDSA_SIGN); + slot->hwc.pke.cmd_eddsa_sign.curve =3D curve; + slot->hwc.pke.cmd_eddsa_sign.secret_key_len =3D sklen; + slot->hwc.pke.cmd_eddsa_sign.digest =3D dig_dma; + slot->hwc.pke.cmd_eddsa_sign.signature =3D sig_dma; + slot->hwc.pke.cmd_eddsa_sign.secret_key =3D sk_ref; + slot->hwc.pke.cmd_eddsa_sign.digest_len =3D dlen; +} + +/** + * vcq_add_pke_eddsa_pubgen() - Build a VCQ command for EdDSA public key g= eneration + * @slot: VCQ command slot to populate + * @core_id: PKE hardware core ID + * @curve: Curve identifier (Ed25519 or Ed448) + * @sklen: Secret key length in bytes + * @pky_dma: DMA address of public key Y-coordinate output buffer + * @sk_ref: Datastore reference for the secret key + * @flags: VCQ command flags + * + * Generates the public key from an existing private key stored in the + * datastore. + */ +void vcq_add_pke_eddsa_pubgen(struct vcq_cmd *slot, u32 core_id, u32 curve= , u32 sklen, + u64 pky_dma, u64 sk_ref, u32 flags) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, flags, 1, PKE_CMD_EDDSA_PUBGEN); + slot->hwc.pke.cmd_eddsa_pubgen.curve =3D curve; + slot->hwc.pke.cmd_eddsa_pubgen.secret_key_len =3D sklen; + slot->hwc.pke.cmd_eddsa_pubgen.public_key_y =3D pky_dma; + slot->hwc.pke.cmd_eddsa_pubgen.secret_key =3D sk_ref; +} + +/** + * vcq_add_pke_eddsa_keygen_sca() - Build a VCQ command for EdDSA SCA key = generation + * @slot: VCQ command slot to populate + * @core_id: PKE hardware core ID + * @curve: Curve identifier (Ed448) + * @sk_ref: Datastore reference for the input secret key + * @sca_sk_ref: Datastore reference for the SCA-masked output key + * + * Blinds an Ed448 private key into a side-channel-protected masked + * form. No byte-swap flags are used (CRI reference uses flags=3D0). + */ +void vcq_add_pke_eddsa_keygen_sca(struct vcq_cmd *slot, u32 core_id, u32 c= urve, + u64 sk_ref, u64 sca_sk_ref) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, + PKE_CMD_EDDSA_PRIV_KEYGEN_SCA); + slot->hwc.pke.cmd_eddsa_keygen_sca.curve =3D curve; + slot->hwc.pke.cmd_eddsa_keygen_sca.secret_key =3D sk_ref; + slot->hwc.pke.cmd_eddsa_keygen_sca.sca_secret_key =3D sca_sk_ref; +} + +/* SM2 */ + +/** + * vcq_add_pke_sm2_ecdh_keygen() - Build a VCQ command for SM2 ECDH epheme= ral key generation + * @slot: VCQ command slot to populate + * @core_id: PKE hardware core ID + * @nonce_dma: DMA address of nonce input buffer + * @session_key_dma: DMA address of session key output buffer + * @nonce_len: Nonce length in bytes + * @flags: VCQ command flags + */ +void vcq_add_pke_sm2_ecdh_keygen(struct vcq_cmd *slot, u32 core_id, u64 no= nce_dma, + u64 session_key_dma, u32 nonce_len, u32 flags) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, flags, 1, + PKE_CMD_SM2_ECDH_KEYGEN); + slot->hwc.pke.cmd_sm2_ecdh_keygen.nonce =3D nonce_dma; + slot->hwc.pke.cmd_sm2_ecdh_keygen.session_key =3D session_key_dma; + slot->hwc.pke.cmd_sm2_ecdh_keygen.nonce_len =3D nonce_len; +} + +/** + * vcq_add_pke_sm2_ecdh() - Build a VCQ command for SM2 ECDH shared secret= computation + * @slot: VCQ command slot to populate + * @core_id: PKE hardware core ID + * @nonce_len: Nonce length in bytes + * @private_key_len: Private key length in bytes + * @nonce_dma: DMA address of nonce buffer + * @peer_pk_dma: DMA address of peer public key buffer + * @peer_sk_dma: DMA address of peer session key buffer + * @priv_ref: Datastore reference for the local private key + * @sp_ref: Datastore reference for the shared point output + * @sp_type: Datastore type for the shared point object + * @flags: VCQ command flags + */ +void vcq_add_pke_sm2_ecdh(struct vcq_cmd *slot, u32 core_id, u32 nonce_len, + u32 private_key_len, u64 nonce_dma, + u64 peer_pk_dma, u64 peer_sk_dma, + u64 priv_ref, u64 sp_ref, u32 sp_type, u32 flags) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, flags, 1, PKE_CMD_SM2_ECDH); + slot->hwc.pke.cmd_sm2_ecdh.nonce_len =3D nonce_len; + slot->hwc.pke.cmd_sm2_ecdh.private_key_len =3D private_key_len; + slot->hwc.pke.cmd_sm2_ecdh.nonce =3D nonce_dma; + slot->hwc.pke.cmd_sm2_ecdh.peer_public_key =3D peer_pk_dma; + slot->hwc.pke.cmd_sm2_ecdh.peer_session_key =3D peer_sk_dma; + slot->hwc.pke.cmd_sm2_ecdh.private_key =3D priv_ref; + slot->hwc.pke.cmd_sm2_ecdh.shared_point =3D sp_ref; + slot->hwc.pke.cmd_sm2_ecdh.shared_point_type =3D sp_type; +} + +/** + * vcq_add_pke_sm2_dec_point() - Build a VCQ command for SM2 decryption po= int multiplication + * @slot: VCQ command slot to populate + * @core_id: PKE hardware core ID + * @ct_len: Ciphertext length in bytes + * @pk_len: Private key length in bytes + * @ct_dma: DMA address of ciphertext input buffer + * @dp_dma: DMA address of decryption point output buffer + * @priv_ref: Datastore reference for the private key + * @flags: VCQ command flags + */ +void vcq_add_pke_sm2_dec_point(struct vcq_cmd *slot, u32 core_id, u32 ct_l= en, + u32 pk_len, u64 ct_dma, u64 dp_dma, + u64 priv_ref, u32 flags) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, flags, 1, PKE_CMD_SM2_DEC_POINT); + slot->hwc.pke.cmd_sm2_dec_point.ciphertext_len =3D ct_len; + slot->hwc.pke.cmd_sm2_dec_point.private_key_len =3D pk_len; + slot->hwc.pke.cmd_sm2_dec_point.ciphertext =3D ct_dma; + slot->hwc.pke.cmd_sm2_dec_point.dec_point =3D dp_dma; + slot->hwc.pke.cmd_sm2_dec_point.private_key =3D priv_ref; +} + +/** + * vcq_add_pke_sm2_enc_point() - Build a VCQ command for SM2 encryption po= int multiplication + * @slot: VCQ command slot to populate + * @core_id: PKE hardware core ID + * @nonce_dma: DMA address of nonce buffer + * @pk_dma: DMA address of public key buffer + * @ct_dma: DMA address of ciphertext header output buffer + * @ep_dma: DMA address of encryption point output buffer + * @nonce_len: Nonce length in bytes + * @flags: VCQ command flags + */ +void vcq_add_pke_sm2_enc_point(struct vcq_cmd *slot, u32 core_id, u64 nonc= e_dma, + u64 pk_dma, u64 ct_dma, u64 ep_dma, + u32 nonce_len, u32 flags) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, flags, 1, PKE_CMD_SM2_ENC_POINT); + slot->hwc.pke.cmd_sm2_enc_point.nonce =3D nonce_dma; + slot->hwc.pke.cmd_sm2_enc_point.public_key =3D pk_dma; + slot->hwc.pke.cmd_sm2_enc_point.ciphertext =3D ct_dma; + slot->hwc.pke.cmd_sm2_enc_point.enc_point =3D ep_dma; + slot->hwc.pke.cmd_sm2_enc_point.nonce_len =3D nonce_len; +} + +/** + * vcq_add_pke_sm2_id_digest() - Build a VCQ command for SM2 identity dige= st computation + * @slot: VCQ command slot to populate + * @core_id: PKE hardware core ID + * @id_dma: DMA address of identity string buffer + * @pk_dma: DMA address of public key buffer + * @dig_dma: DMA address of digest output buffer + * @id_len: Identity string length in bytes + * @flags: VCQ command flags + */ +void vcq_add_pke_sm2_id_digest(struct vcq_cmd *slot, u32 core_id, u64 id_d= ma, + u64 pk_dma, u64 dig_dma, u32 id_len, + u32 flags) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, flags, 1, PKE_CMD_SM2_ID_DIGEST); + slot->hwc.pke.cmd_sm2_id_digest.id =3D id_dma; + slot->hwc.pke.cmd_sm2_id_digest.public_key =3D pk_dma; + slot->hwc.pke.cmd_sm2_id_digest.digest =3D dig_dma; + slot->hwc.pke.cmd_sm2_id_digest.id_len =3D id_len; +} + +/** + * vcq_add_pke_sm2_ecdh_hash() - Build a VCQ command for SM2 ECDH key deri= vation hash + * @slot: VCQ command slot to populate + * @core_id: PKE hardware core ID + * @peer_dig_dma: DMA address of peer identity digest buffer + * @dig_dma: DMA address of local identity digest buffer + * @sp_ref: Datastore reference for the shared point + * @sk_ref: Datastore reference for the derived shared key output + * @sk_type: Datastore type for the shared key object + * @flags: VCQ command flags + */ +void vcq_add_pke_sm2_ecdh_hash(struct vcq_cmd *slot, u32 core_id, u64 peer= _dig_dma, + u64 dig_dma, u64 sp_ref, u64 sk_ref, + u32 sk_type, u32 flags) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, flags, 1, PKE_CMD_SM2_ECDH_HASH); + slot->hwc.pke.cmd_sm2_ecdh_hash.peer_id_digest =3D peer_dig_dma; + slot->hwc.pke.cmd_sm2_ecdh_hash.id_digest =3D dig_dma; + slot->hwc.pke.cmd_sm2_ecdh_hash.shared_point =3D sp_ref; + slot->hwc.pke.cmd_sm2_ecdh_hash.shared_key =3D sk_ref; + slot->hwc.pke.cmd_sm2_ecdh_hash.shared_key_type =3D sk_type; +} + +/** + * vcq_add_pke_sm2_dec_hash() - Build a VCQ command for SM2 decryption has= h verification + * @slot: VCQ command slot to populate + * @core_id: PKE hardware core ID + * @ct_dma: DMA address of ciphertext input buffer + * @dp_dma: DMA address of decryption point buffer + * @pt_dma: DMA address of plaintext output buffer + * @ct_len: Ciphertext length in bytes + * @flags: VCQ command flags + */ +void vcq_add_pke_sm2_dec_hash(struct vcq_cmd *slot, u32 core_id, u64 ct_dm= a, + u64 dp_dma, u64 pt_dma, u32 ct_len, u32 flags) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, flags, 1, PKE_CMD_SM2_DEC_HASH); + slot->hwc.pke.cmd_sm2_dec_hash.ciphertext =3D ct_dma; + slot->hwc.pke.cmd_sm2_dec_hash.dec_point =3D dp_dma; + slot->hwc.pke.cmd_sm2_dec_hash.plaintext =3D pt_dma; + slot->hwc.pke.cmd_sm2_dec_hash.ciphertext_len =3D ct_len; +} + +/** + * vcq_add_pke_sm2_enc_hash() - Build a VCQ command for SM2 encryption has= h computation + * @slot: VCQ command slot to populate + * @core_id: PKE hardware core ID + * @msg_dma: DMA address of plaintext message buffer + * @ep_dma: DMA address of encryption point buffer + * @ct_dma: DMA address of ciphertext output buffer + * @msg_len: Message length in bytes + * @flags: VCQ command flags + */ +void vcq_add_pke_sm2_enc_hash(struct vcq_cmd *slot, u32 core_id, u64 msg_d= ma, + u64 ep_dma, u64 ct_dma, u32 msg_len, u32 flags) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, flags, 1, PKE_CMD_SM2_ENC_HASH); + slot->hwc.pke.cmd_sm2_enc_hash.message =3D msg_dma; + slot->hwc.pke.cmd_sm2_enc_hash.enc_point =3D ep_dma; + slot->hwc.pke.cmd_sm2_enc_hash.ciphertext =3D ct_dma; + slot->hwc.pke.cmd_sm2_enc_hash.message_len =3D msg_len; +} diff --git a/drivers/crypto/cmh/cmh_pke_rsa.c b/drivers/crypto/cmh/cmh_pke_= rsa.c new file mode 100644 index 000000000000..251ae8275865 --- /dev/null +++ b/drivers/crypto/cmh/cmh_pke_rsa.c @@ -0,0 +1,644 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- RSA akcipher Driver + * + * Registers "rsa" akcipher algorithm with the Linux crypto subsystem + * (priority 300, overrides software rsa-generic at 100). + * + * Raw RSA operations only (m^e mod n / c^d mod n). The kernel's + * pkcs1pad() template wraps this for PKCS#1 v1.5 / PSS / OAEP. + * + * Key format: DER-encoded ASN.1, parsed by kernel rsa_parse_pub_key() + * / rsa_parse_priv_key() helpers. + * + * Private key via cmh_key_ctx: raw keys written via SYS_REF_TEMP. + * Datastore-referenced keys are only reachable through the ioctl + * path (cmh_mgmt.c). + */ + +#include +#include +#include +#include +#include +#include +#include + +#include "cmh_pke.h" +#include "cmh_sys.h" +#include "cmh_sys_abi.h" +#include "cmh_txn.h" +#include "cmh_dma.h" +#include "cmh_key.h" + +struct cmh_rsa_tfm_ctx { + struct cmh_key_ctx key; /* private key (raw d only) */ + u8 *n; /* modulus (big-endian) */ + u8 *e; /* public exponent (big-endian) */ + size_t n_sz; + size_t e_sz; + u32 bits; /* key size in bits */ +}; + +static inline struct cmh_rsa_tfm_ctx *cmh_rsa_ctx(struct crypto_akcipher *= tfm) +{ + return akcipher_tfm_ctx(tfm); +} + +struct cmh_rsa_reqctx { + u8 *e_buf; + u8 *n_buf; + u8 *m_buf; + u8 *c_buf; + dma_addr_t e_dma; + dma_addr_t n_dma; + dma_addr_t m_dma; + dma_addr_t c_dma; + u32 key_bytes; + u32 e_padded; + u32 n_sz; +}; + +static u32 cmh_rsa_key_bits(size_t n_sz) +{ + /* + * Only accept exact modulus sizes supported by the hardware. + * The programmed RSA width must match the actual modulus buffer + * length; rounding a shorter modulus up to the next size would + * let the device read past the end of the DMA buffer. + */ + switch (n_sz) { + case 64: + return 512; + case 128: + return 1024; + case 256: + return 2048; + case 384: + return 3072; + case 512: + return 4096; + default: + return 0; + } +} + +static void cmh_rsa_enc_complete(void *data, int error) +{ + struct akcipher_request *req =3D data; + struct cmh_rsa_reqctx *rctx =3D akcipher_request_ctx(req); + + if (error =3D=3D -EINPROGRESS) { + cmh_complete(&req->base, error); + return; + } + + if (!cmh_dma_map_error(rctx->c_dma)) + cmh_dma_unmap_single(rctx->c_dma, rctx->key_bytes, + DMA_FROM_DEVICE); + if (!cmh_dma_map_error(rctx->m_dma)) + cmh_dma_unmap_single(rctx->m_dma, rctx->key_bytes, + DMA_TO_DEVICE); + if (!cmh_dma_map_error(rctx->n_dma)) + cmh_dma_unmap_single(rctx->n_dma, rctx->n_sz, + DMA_TO_DEVICE); + if (!cmh_dma_map_error(rctx->e_dma)) + cmh_dma_unmap_single(rctx->e_dma, rctx->e_padded, + DMA_TO_DEVICE); + + if (!error) { + int nents; + + nents =3D sg_nents_for_len(req->dst, rctx->key_bytes); + if (nents < 0 || + sg_copy_from_buffer(req->dst, nents, + rctx->c_buf, + rctx->key_bytes) !=3D rctx->key_bytes) + error =3D -EINVAL; + else + req->dst_len =3D rctx->key_bytes; + } + + kfree(rctx->c_buf); + rctx->c_buf =3D NULL; + kfree_sensitive(rctx->m_buf); + rctx->m_buf =3D NULL; + kfree(rctx->n_buf); + rctx->n_buf =3D NULL; + kfree(rctx->e_buf); + rctx->e_buf =3D NULL; + cmh_complete(&req->base, error); +} + +/* + * RSA encrypt: c =3D m^e mod n (public key operation) + * Also used for signature verification (verify =3D encrypt for raw RSA). + */ +static int cmh_rsa_enc(struct akcipher_request *req) +{ + struct crypto_akcipher *tfm =3D crypto_akcipher_reqtfm(req); + struct cmh_rsa_tfm_ctx *ctx =3D cmh_rsa_ctx(tfm); + struct cmh_rsa_reqctx *rctx =3D akcipher_request_ctx(req); + u32 key_bytes =3D ctx->bits / 8; + u32 e_padded =3D ALIGN(ctx->e_sz, 4); + struct core_dispatch d =3D cmh_core_select_instance(CMH_CORE_PKE); + struct vcq_cmd vcq[PKE_VCQ_CMDS_MIN]; + int ret, nents; + gfp_t gfp; + + if (!ctx->n || !ctx->e) + return -EINVAL; + if (req->src_len > key_bytes) + return -EINVAL; + if (req->dst_len < key_bytes) { + req->dst_len =3D key_bytes; + return -EOVERFLOW; + } + + gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + + memset(rctx, 0, sizeof(*rctx)); + rctx->key_bytes =3D key_bytes; + rctx->e_padded =3D e_padded; + rctx->n_sz =3D ctx->n_sz; + rctx->e_dma =3D DMA_MAPPING_ERROR; + rctx->n_dma =3D DMA_MAPPING_ERROR; + rctx->m_dma =3D DMA_MAPPING_ERROR; + rctx->c_dma =3D DMA_MAPPING_ERROR; + + rctx->e_buf =3D kzalloc(e_padded, gfp); + rctx->n_buf =3D kmemdup(ctx->n, ctx->n_sz, gfp); + rctx->m_buf =3D kzalloc(key_bytes, gfp); + rctx->c_buf =3D kzalloc(key_bytes, gfp); + if (!rctx->e_buf || !rctx->n_buf || !rctx->m_buf || !rctx->c_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + memcpy(rctx->e_buf + e_padded - ctx->e_sz, ctx->e, ctx->e_sz); + + nents =3D sg_nents_for_len(req->src, req->src_len); + if (nents < 0 || + sg_pcopy_to_buffer(req->src, nents, + rctx->m_buf + key_bytes - req->src_len, + req->src_len, 0) !=3D req->src_len) { + ret =3D -EINVAL; + goto out_free; + } + + rctx->e_dma =3D cmh_dma_map_single(rctx->e_buf, e_padded, + DMA_TO_DEVICE); + rctx->n_dma =3D cmh_dma_map_single(rctx->n_buf, ctx->n_sz, + DMA_TO_DEVICE); + rctx->m_dma =3D cmh_dma_map_single(rctx->m_buf, key_bytes, + DMA_TO_DEVICE); + rctx->c_dma =3D cmh_dma_map_single(rctx->c_buf, key_bytes, + DMA_FROM_DEVICE); + + if (cmh_dma_map_error(rctx->e_dma) || + cmh_dma_map_error(rctx->n_dma) || + cmh_dma_map_error(rctx->m_dma) || + cmh_dma_map_error(rctx->c_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + vcq_set_header(&vcq[0], PKE_VCQ_CMDS_MIN); + vcq_add_pke_rsa_enc(&vcq[1], d.core_id, ctx->bits, e_padded, + rctx->e_dma, rctx->n_dma, rctx->m_dma, + rctx->c_dma, PKE_SWAP_FLAGS); + vcq_add_pke_flush(&vcq[2], d.core_id); + + ret =3D cmh_tm_submit_async(vcq, PKE_VCQ_CMDS_MIN, 1, d.mbx_idx, + cmh_rsa_enc_complete, req, + !!(req->base.flags & + CRYPTO_TFM_REQ_MAY_BACKLOG), 0); + if (ret =3D=3D -EBUSY) + return -EBUSY; + if (!ret) + return -EINPROGRESS; + +out_unmap: + if (!cmh_dma_map_error(rctx->c_dma)) + cmh_dma_unmap_single(rctx->c_dma, key_bytes, + DMA_FROM_DEVICE); + if (!cmh_dma_map_error(rctx->m_dma)) + cmh_dma_unmap_single(rctx->m_dma, key_bytes, + DMA_TO_DEVICE); + if (!cmh_dma_map_error(rctx->n_dma)) + cmh_dma_unmap_single(rctx->n_dma, ctx->n_sz, + DMA_TO_DEVICE); + if (!cmh_dma_map_error(rctx->e_dma)) + cmh_dma_unmap_single(rctx->e_dma, e_padded, + DMA_TO_DEVICE); + +out_free: + kfree(rctx->c_buf); + kfree_sensitive(rctx->m_buf); + kfree(rctx->n_buf); + kfree(rctx->e_buf); + return ret; +} + +static void cmh_rsa_dec_complete(void *data, int error) +{ + struct akcipher_request *req =3D data; + struct cmh_rsa_reqctx *rctx =3D akcipher_request_ctx(req); + + if (error =3D=3D -EINPROGRESS) { + cmh_complete(&req->base, error); + return; + } + + if (!cmh_dma_map_error(rctx->m_dma)) + cmh_dma_unmap_single(rctx->m_dma, rctx->key_bytes, + DMA_FROM_DEVICE); + if (!cmh_dma_map_error(rctx->c_dma)) + cmh_dma_unmap_single(rctx->c_dma, rctx->key_bytes, + DMA_TO_DEVICE); + if (!cmh_dma_map_error(rctx->n_dma)) + cmh_dma_unmap_single(rctx->n_dma, rctx->n_sz, + DMA_TO_DEVICE); + if (!cmh_dma_map_error(rctx->e_dma)) + cmh_dma_unmap_single(rctx->e_dma, rctx->e_padded, + DMA_TO_DEVICE); + + if (!error) { + int nents; + + nents =3D sg_nents_for_len(req->dst, rctx->key_bytes); + if (nents < 0 || + sg_copy_from_buffer(req->dst, nents, + rctx->m_buf, + rctx->key_bytes) !=3D rctx->key_bytes) + error =3D -EINVAL; + else + req->dst_len =3D rctx->key_bytes; + } + + kfree_sensitive(rctx->m_buf); + rctx->m_buf =3D NULL; + kfree(rctx->c_buf); + rctx->c_buf =3D NULL; + kfree(rctx->n_buf); + rctx->n_buf =3D NULL; + kfree(rctx->e_buf); + rctx->e_buf =3D NULL; + cmh_complete(&req->base, error); +} + +/* + * RSA decrypt: m =3D c^d mod n (private key operation) + * Also used for signing (sign =3D decrypt for raw RSA). + * + * Private key 'd' is written via SYS_REF_TEMP inline. + */ +static int cmh_rsa_dec(struct akcipher_request *req) +{ + struct crypto_akcipher *tfm =3D crypto_akcipher_reqtfm(req); + struct cmh_rsa_tfm_ctx *ctx =3D cmh_rsa_ctx(tfm); + struct cmh_rsa_reqctx *rctx =3D akcipher_request_ctx(req); + u32 key_bytes =3D ctx->bits / 8; + u32 e_padded =3D ALIGN(ctx->e_sz, 4); + struct vcq_cmd vcq[PKE_VCQ_CMDS_MAX]; + struct core_dispatch dd; + int ret, idx, nents; + gfp_t gfp; + + if (ctx->key.mode !=3D CMH_KEY_RAW) + return -EINVAL; + if (!ctx->n || !ctx->e) + return -EINVAL; + if (req->src_len > key_bytes) + return -EINVAL; + if (req->dst_len < key_bytes) { + req->dst_len =3D key_bytes; + return -EOVERFLOW; + } + + gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + + memset(rctx, 0, sizeof(*rctx)); + rctx->key_bytes =3D key_bytes; + rctx->e_padded =3D e_padded; + rctx->n_sz =3D ctx->n_sz; + rctx->e_dma =3D DMA_MAPPING_ERROR; + rctx->n_dma =3D DMA_MAPPING_ERROR; + rctx->m_dma =3D DMA_MAPPING_ERROR; + rctx->c_dma =3D DMA_MAPPING_ERROR; + + rctx->e_buf =3D kzalloc(e_padded, gfp); + rctx->n_buf =3D kmemdup(ctx->n, ctx->n_sz, gfp); + rctx->c_buf =3D kzalloc(key_bytes, gfp); + rctx->m_buf =3D kzalloc(key_bytes, gfp); + if (!rctx->e_buf || !rctx->n_buf || !rctx->c_buf || !rctx->m_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + memcpy(rctx->e_buf + e_padded - ctx->e_sz, ctx->e, ctx->e_sz); + + nents =3D sg_nents_for_len(req->src, req->src_len); + if (nents < 0 || + sg_pcopy_to_buffer(req->src, nents, + rctx->c_buf + key_bytes - req->src_len, + req->src_len, 0) !=3D req->src_len) { + ret =3D -EINVAL; + goto out_free; + } + + rctx->e_dma =3D cmh_dma_map_single(rctx->e_buf, e_padded, + DMA_TO_DEVICE); + rctx->n_dma =3D cmh_dma_map_single(rctx->n_buf, ctx->n_sz, + DMA_TO_DEVICE); + rctx->c_dma =3D cmh_dma_map_single(rctx->c_buf, key_bytes, + DMA_TO_DEVICE); + rctx->m_dma =3D cmh_dma_map_single(rctx->m_buf, key_bytes, + DMA_FROM_DEVICE); + + if (cmh_dma_map_error(rctx->e_dma) || + cmh_dma_map_error(rctx->n_dma) || + cmh_dma_map_error(rctx->c_dma) || + cmh_dma_map_error(rctx->m_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + dd =3D cmh_core_select_instance(CMH_CORE_PKE); + + idx =3D 1; + vcq_add_sys_write(&vcq[idx], SYS_REF_TEMP, ctx->key.raw.dma, + SYS_REF_NONE, ctx->key.raw.len, + ctx->key.raw.sys_type); + vcq[idx].id |=3D PKE_SWAP_FLAGS; + idx++; + vcq_add_pke_rsa_dec(&vcq[idx++], dd.core_id, ctx->bits, e_padded, + rctx->e_dma, rctx->n_dma, rctx->c_dma, + rctx->m_dma, SYS_REF_TEMP, PKE_SWAP_FLAGS); + vcq_add_pke_flush(&vcq[idx++], dd.core_id); + vcq_set_header(&vcq[0], idx); + + ret =3D cmh_tm_submit_async(vcq, idx, 1, dd.mbx_idx, + cmh_rsa_dec_complete, req, + !!(req->base.flags & + CRYPTO_TFM_REQ_MAY_BACKLOG), 0); + if (ret =3D=3D -EBUSY) + return -EBUSY; + if (!ret) + return -EINPROGRESS; + +out_unmap: + if (!cmh_dma_map_error(rctx->m_dma)) + cmh_dma_unmap_single(rctx->m_dma, key_bytes, + DMA_FROM_DEVICE); + if (!cmh_dma_map_error(rctx->c_dma)) + cmh_dma_unmap_single(rctx->c_dma, key_bytes, + DMA_TO_DEVICE); + if (!cmh_dma_map_error(rctx->n_dma)) + cmh_dma_unmap_single(rctx->n_dma, ctx->n_sz, + DMA_TO_DEVICE); + if (!cmh_dma_map_error(rctx->e_dma)) + cmh_dma_unmap_single(rctx->e_dma, e_padded, + DMA_TO_DEVICE); + +out_free: + kfree_sensitive(rctx->m_buf); + kfree(rctx->c_buf); + kfree(rctx->n_buf); + kfree(rctx->e_buf); + return ret; +} + +static int cmh_rsa_set_pub_key(struct crypto_akcipher *tfm, + const void *key, unsigned int keylen) +{ + struct cmh_rsa_tfm_ctx *ctx =3D cmh_rsa_ctx(tfm); + struct rsa_key rsa =3D {}; + u32 bits; + int ret; + + /* + * Re-keying: release any private-key state (and its persistent DMA + * mapping) left by a prior set_priv_key so it is not leaked until + * exit_tfm and cannot be mistaken for the new public key. + */ + cmh_key_destroy(&ctx->key); + + ret =3D rsa_parse_pub_key(&rsa, key, keylen); + if (ret) + return ret; + + /* Strip ASN.1 leading zero padding from modulus */ + while (rsa.n_sz > 0 && rsa.n[0] =3D=3D 0) { + rsa.n++; + rsa.n_sz--; + } + + bits =3D cmh_rsa_key_bits(rsa.n_sz); + if (!bits) + return -EINVAL; + + /* Reject an exponent wider than the modulus (HW buffer bound). */ + if (!rsa.e_sz || rsa.e_sz > bits / 8) + return -EINVAL; + + kfree(ctx->n); + kfree(ctx->e); + ctx->n =3D NULL; + ctx->e =3D NULL; + ctx->n_sz =3D 0; + ctx->e_sz =3D 0; + + ctx->n =3D kmemdup(rsa.n, rsa.n_sz, GFP_KERNEL); + ctx->e =3D kmemdup(rsa.e, rsa.e_sz, GFP_KERNEL); + if (!ctx->n || !ctx->e) { + kfree(ctx->n); + kfree(ctx->e); + ctx->n =3D NULL; + ctx->e =3D NULL; + return -ENOMEM; + } + + ctx->n_sz =3D rsa.n_sz; + ctx->e_sz =3D rsa.e_sz; + ctx->bits =3D bits; + + return 0; +} + +static int cmh_rsa_set_priv_key(struct crypto_akcipher *tfm, + const void *key, unsigned int keylen) +{ + struct cmh_rsa_tfm_ctx *ctx =3D cmh_rsa_ctx(tfm); + struct rsa_key rsa =3D {}; + u32 bits, key_bytes; + u8 *d_padded; + int ret; + + ret =3D rsa_parse_priv_key(&rsa, key, keylen); + if (ret) + return ret; + + /* Strip ASN.1 leading zero padding from modulus */ + while (rsa.n_sz > 0 && rsa.n[0] =3D=3D 0) { + rsa.n++; + rsa.n_sz--; + } + + bits =3D cmh_rsa_key_bits(rsa.n_sz); + if (!bits || !rsa.d_sz) + return -EINVAL; + + key_bytes =3D bits / 8; + + /* Strip ASN.1 leading zero padding from private exponent */ + while (rsa.d_sz > 0 && rsa.d[0] =3D=3D 0) { + rsa.d++; + rsa.d_sz--; + } + + if (!rsa.d_sz || rsa.d_sz > key_bytes) + return -EINVAL; + + /* Reject an exponent wider than the modulus (HW buffer bound). */ + if (!rsa.e_sz || rsa.e_sz > key_bytes) + return -EINVAL; + + kfree(ctx->n); + kfree(ctx->e); + ctx->n =3D NULL; + ctx->e =3D NULL; + ctx->n_sz =3D 0; + ctx->e_sz =3D 0; + + ctx->n =3D kmemdup(rsa.n, rsa.n_sz, GFP_KERNEL); + ctx->e =3D kmemdup(rsa.e, rsa.e_sz, GFP_KERNEL); + if (!ctx->n || !ctx->e) { + ret =3D -ENOMEM; + goto err; + } + + ctx->n_sz =3D rsa.n_sz; + ctx->e_sz =3D rsa.e_sz; + ctx->bits =3D bits; + + /* + * Left-pad d to key_bytes (big-endian alignment). + * The CMH eSW resolves SYS_REF_TEMP by checking + * hdr->len >=3D key_bytes, so the written buffer must + * be at least key_bytes wide. + */ + d_padded =3D kzalloc(key_bytes, GFP_KERNEL); + if (!d_padded) { + ret =3D -ENOMEM; + goto err; + } + memcpy(d_padded + key_bytes - rsa.d_sz, rsa.d, rsa.d_sz); + + ret =3D cmh_key_setkey_raw(&ctx->key, d_padded, key_bytes, + CORE_ID_PKE); + kfree_sensitive(d_padded); + if (ret) + goto err; + + return 0; +err: + kfree(ctx->n); + kfree(ctx->e); + ctx->n =3D NULL; + ctx->e =3D NULL; + ctx->n_sz =3D 0; + ctx->e_sz =3D 0; + ctx->bits =3D 0; + return ret; +} + +static unsigned int cmh_rsa_max_size(struct crypto_akcipher *tfm) +{ + struct cmh_rsa_tfm_ctx *ctx =3D cmh_rsa_ctx(tfm); + + return ctx->n_sz; +} + +static int cmh_rsa_init_tfm(struct crypto_akcipher *tfm) +{ + struct cmh_rsa_tfm_ctx *ctx =3D cmh_rsa_ctx(tfm); + + memset(ctx, 0, sizeof(*ctx)); + tfm->reqsize =3D sizeof(struct cmh_rsa_reqctx); + return 0; +} + +static void cmh_rsa_exit_tfm(struct crypto_akcipher *tfm) +{ + struct cmh_rsa_tfm_ctx *ctx =3D cmh_rsa_ctx(tfm); + + cmh_key_destroy(&ctx->key); + kfree(ctx->n); + kfree(ctx->e); + ctx->n =3D NULL; + ctx->e =3D NULL; +} + +/* + * Raw RSA stays as akcipher (encrypt/decrypt only). The kernel's + * rsassa-pkcs1 sig template wraps our akcipher for sign/verify, + * matching the upstream split (rsa.c =3D akcipher, + * rsassa-pkcs1.c =3D sig template). + */ +static struct akcipher_alg cmh_rsa_alg =3D { + .encrypt =3D cmh_rsa_enc, + .decrypt =3D cmh_rsa_dec, + .set_pub_key =3D cmh_rsa_set_pub_key, + .set_priv_key =3D cmh_rsa_set_priv_key, + .max_size =3D cmh_rsa_max_size, + .init =3D cmh_rsa_init_tfm, + .exit =3D cmh_rsa_exit_tfm, + .base =3D { + .cra_name =3D "rsa", + .cra_driver_name =3D "rambus-cmh-rsa", + .cra_priority =3D 300, + .cra_flags =3D CRYPTO_ALG_ASYNC, + .cra_module =3D THIS_MODULE, + .cra_ctxsize =3D sizeof(struct cmh_rsa_tfm_ctx), + }, +}; + +static bool cmh_rsa_registered; + +/** + * cmh_pke_rsa_register() - Register RSA akcipher algorithm with the crypt= o framework + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_pke_rsa_register(void) +{ + int ret; + + if (!cmh_core_present(CMH_CORE_PKE)) + return 0; + + ret =3D crypto_register_akcipher(&cmh_rsa_alg); + if (ret) { + dev_err(cmh_dev(), + "cmh: failed to register rsa akcipher (%d)\n", + ret); + return ret; + } + + cmh_rsa_registered =3D true; + return 0; +} + +/** + * cmh_pke_rsa_unregister() - Unregister RSA akcipher algorithm from the c= rypto framework + */ +void cmh_pke_rsa_unregister(void) +{ + if (cmh_rsa_registered) + crypto_unregister_akcipher(&cmh_rsa_alg); + cmh_rsa_registered =3D false; +} --=20 2.43.7 From nobody Fri Sep 25 01:20:34 2026 Received: from DM1PR04CU001.outbound.protection.outlook.com (mail-centralusazon11020087.outbound.protection.outlook.com [52.101.61.87]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 666E24B4035; Thu, 17 Sep 2026 22:59:51 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.61.87 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789685998; cv=fail; b=f5Rr/0K18FCE/qjX0/JE0qL6z70P73cNj3EfwwjG2KxcUCVh7XRzRmyi26ePOCTja4ATAo32ia8lfkCedFfImim6vCXjHIdzppSiShG6Hs9Q/BOA5kUgPxc8f4iRNh6OdkDQuQCqkJSXjDSC2yjT4pc86fhkf98kkPPJms2myFE= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789685998; c=relaxed/simple; bh=1B2gta500JZeA+yoUjbExcL7NVm2RNYOZohnml9gM60=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=U11LSqXZmHmhF1PGq+xecwtu1by4Mxm0431XzhUOXOwpiho+UpdAqbwCd2E5ZJDF8ySLVIpT5bwYTbBwcWZHP0+uRYwSK9dI8wi6rLbJ2Vul8s4KDgy4hh3cD1w93sNQuXS67cGXe1ptduWN1KPcKxFoRP4QjAC2STbtY79O9w8= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com; spf=fail smtp.mailfrom=rambus.com; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b=hJ3KWdXH; arc=fail smtp.client-ip=52.101.61.87 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=rambus.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b="hJ3KWdXH" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=dwFkGr1eNctC1FGS7hGQH9g06sI6T3HwjdTREh3IxRruPTrsn75Tv67QLFNQWktz/ikmVp6qOqLfmO5vrpG0SfWlKqsgeDrqOLAoZnRsBcK7rNwNizS6zTM7ZRa4nqZXALps8PwClKRyUOFjIR+B4HHYe3fm/Xyzls2ennBZrI2WdSNge8G9SspSrmzQzf09cD1V5y8tvVczKzJYzuHSsX4Ow1Te7KnggIZvw097SE7Y+k9Ee2QPb1Gi2hpzqZmygM8w7EbqPmU1PdMY1hAQezkjDAvSl+XfhzMpaytm+D9UvgnJsj/HoF/+lFSz27+ETQhGzQU9Lt4HHaWI3T8cHQ== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=X2VX9lKohkPpnh4eRIAnB5oaBaH3obl+RR6aSz4YAPk=; b=lut7XeF4t18jbozBH7sZ9EZ0T86u8D2q+AF64Cwl3enQFCxf0DH3Xv/fetMHRMxIn+lKVGbo1DcHIjICxHwa6xnfRLfrfhOzB1AFkr+w410jttVAbatD5tGlJrtZBLcwemrmFaqmjcdShjKGD1B1JlGCh84sPE65sqK3mV2/EJNlzLg67AhKQa/MPiGd26QQQuPw1FJSaNdnUkmX6wHU3gUEuYk5cbSt2VWvdf17PioBxa/weUqdBsUB4VKU98FeKncp//uReg0uUbtK8PkaB56mxbwUYnuGZ3JULs9DhXSccNCSHhf6lcRW9XIyu6j9AP8ruPqSRYBfJizAr7IFLA== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 192.86.86.210) smtp.rcpttodomain=cryptography.com smtp.mailfrom=rambus.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=rambus.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=rambus.com; s=selector1; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=X2VX9lKohkPpnh4eRIAnB5oaBaH3obl+RR6aSz4YAPk=; b=hJ3KWdXHiFmgoQX2CYuPyHsRozFs36dgEhSgpoAP2E+UkKMUra8dy/MMSPFx4R1Tee7wHVHZHsxxf4e36LQtmug7k5u6P8nMq/OLKhaXlP8SlZwsooa34po5f6XKCZJ4kb2muUtoaN/OZzaU/2kfHQuUEQXnwBuDQGIDpll787BBvKbXtDBkW+av6YpBwtnIZUtsj23TNbl5cSVPQTkCaenHAuq6psuDWulvhip6GRTY3mdoX9uJ86/dmfcVtdM78bLnW+R0uZzBoLnOq3bWGWjPbX6q5oEgyxRa9/VkFs4ZDiv2x3zB3+vcT6GCu1o6qH30D61QMIy6G6Ihdjy1tQ== Received: from BN9PR03CA0661.namprd03.prod.outlook.com (2603:10b6:408:10e::6) by MN2PR04MB6639.namprd04.prod.outlook.com (2603:10b6:208:1e8::9) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.428.12; Thu, 17 Sep 2026 22:59:36 +0000 Received: from BN2PEPF0000A804.namprd02.prod.outlook.com (2603:10b6:408:10e:cafe::78) by BN9PR03CA0661.outlook.office365.com (2603:10b6:408:10e::6) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.428.12 via Frontend Transport; Thu, 17 Sep 2026 22:59:36 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 192.86.86.210) smtp.mailfrom=rambus.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=rambus.com; Received-SPF: Pass (protection.outlook.com: domain of rambus.com designates 192.86.86.210 as permitted sender) receiver=protection.outlook.com; client-ip=192.86.86.210; helo=hqxsv-psmtppxy02.rambus.com; pr=C Received: from hqxsv-psmtppxy02.rambus.com (192.86.86.210) by BN2PEPF0000A804.mail.protection.outlook.com (10.167.245.168) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.451.8 via Frontend Transport; Thu, 17 Sep 2026 22:59:34 +0000 Received: from hqxsv-cmdev3-aousherovitch.rambus.com (hqn-lb-int-float.rambus.com [10.12.20.20]) by hqxsv-psmtppxy02.rambus.com (Postfix) with ESMTP id 632E4180176B; Thu, 17 Sep 2026 22:59:32 +0000 (UTC) From: Alex Ousherovitch To: Albert Ou , Alex Ousherovitch , Conor Dooley , "David S. Miller" , Herbert Xu , Jonathan Corbet , Krzysztof Kozlowski , Palmer Dabbelt , Paul Walmsley , Rob Herring , Saravanakrishnan Krishnamoorthy , Shuah Khan Cc: Alexandre Ghiti , devicetree@vger.kernel.org, Joel Wittenauer , linux-api@vger.kernel.org, linux-crypto@vger.kernel.org, linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-riscv@lists.infradead.org, Shuah Khan , Thi Nguyen Subject: [PATCH v5 13/19] crypto: cmh - add ECDSA/SM2 sig Date: Thu, 17 Sep 2026 15:59:22 -0700 Message-ID: <20260917225929.2494111-14-aousherovitch@rambus.com> X-Mailer: git-send-email 2.43.7 In-Reply-To: <20260917225929.2494111-1-aousherovitch@rambus.com> References: <20260917225929.2494111-1-aousherovitch@rambus.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: BN2PEPF0000A804:EE_|MN2PR04MB6639:EE_ X-MS-Office365-Filtering-Correlation-Id: a044d8d0-f228-468b-88b6-08df150f57b7 X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|36860700016|376014|23010399003|7416014|1800799024|82310400026|22082099003|18002099003|3023799007|921020|11063799006|10067099003|56012099006; X-Microsoft-Antispam-Message-Info: ZxI9qd+m72RLev2sKimTG0wrfjJkJ0fnutnQJp9m3wt5DhMsdV1nqRjS1WgHl1WSzooPG0BKs4GSBBu5svUcjBfW1K2PGNA3Qh6dW9QCEAMpB1CDCPfbY9xg6sqNnxsjviR09/7u7c00xOK1MCOz62Eo7r3XB4p5n7+ZGS5cQYKDjNXOS6p+rsQeZFZ9g/fdm9T1IHfuCHiaDNRntIQTPq0xuePh1mKDfM/OWhu34p+GnDU/RSK7wE3jl9W7vnaMJe4rrXWZicWnXgNJsFzNFQL1bLkFipfgjJUm/+0GRS7XMkVZZTp+KCTWxunpd+VIfrZi0LFUi1/MWEISc+nHNIaQVA/84+zfhQKwsI2SjtBJCq6kvOOSZMU9S6leOLYtZYEcja+PjL1JiMgvghViKb9xlzvmEp+h5O68Knft/sBbawSe7U4fbsA4ssIbMwumAFYTC2gP4ArTAb9Wk/1ZIcd7qHTgW4A11toZ7z6KdCzYaADKWoKqpMoSCanDcEO4Ea5ExMu0QeFesjMB5j9X93NrNyB8fdNWfvYIX1981GcqCwkJ47JE9vAHcsvvKHHaokQKGJ5tesyofsJhUEAP36OetQ/XdCJl+h2Ia5efyhGUEgfhL9cLGMpixrKXkYqekyp+pHkI6KtmU0pV1Rim9RejzHyT77IZrA12yrWuu3JLW+W/LM29zw2BKDbSpyis5wbJ7bldCi0NY5ZYn5PeMQQKvSJ1VkrB3uKstnYV3CVBaSyKJD8V8ji9oCnDTWUK X-Forefront-Antispam-Report: CIP:192.86.86.210;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:hqxsv-psmtppxy02.rambus.com;PTR:InfoDomainNonexistent;CAT:NONE;SFS:(13230040)(36860700016)(376014)(23010399003)(7416014)(1800799024)(82310400026)(22082099003)(18002099003)(3023799007)(921020)(11063799006)(10067099003)(56012099006);DIR:OUT;SFP:1102; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: aqMBHgVEpiduGQnRMeoePrCVQ6PktNdswVX+yzcHMB9DCaRM0sPGVwqjcQ5rXmvoif541fr7csa/Y7C4OhlRUh2ffw1OTtLehDnHuaw58r5fyA4lh9vxlu4qZpNnWpXUPGjeKM43E0lOKsxBBuIoKe3ps3wICVucEqFtJ5NHk5XZ+b4+8uxNwWObAbgugrmGac6R4ImntJM4sFsQyKDQ5VzJ84ne3j58wfUBxUpDmnJAsz/U5C3tA0x/8DwH6h3Moya/+h3hstsbut9KYCnOUkmFp5fnUwIO5+avjOPjAcTJ/WgvaRK0wIlkGunwVI3ShnkSb693UwfmEHmIbxD1Wqn+0+Zhzus3kK+o6Ebl3U4WEkWMgQf9NJWjCpkCQ7iUuNMmXcyvTQhJlgpryPC2l6FP56MN0YlFjh6KwHTiY788IxgVWB0Jn4/pRGBuVdA6 X-OriginatorOrg: rambus.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 17 Sep 2026 22:59:34.7922 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: a044d8d0-f228-468b-88b6-08df150f57b7 X-MS-Exchange-CrossTenant-Id: bd0ba799-c2b9-413c-9c56-5d1731c4827c X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=bd0ba799-c2b9-413c-9c56-5d1731c4827c;Ip=[192.86.86.210];Helo=[hqxsv-psmtppxy02.rambus.com] X-MS-Exchange-CrossTenant-AuthSource: BN2PEPF0000A804.namprd02.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: MN2PR04MB6639 Content-Type: text/plain; charset="utf-8" Register ECDSA and SM2 sig algorithms using the CMH PKE core. Supports P-256, P-384, P-521, and SM2 curves for sign and verify operations. SM2 is registered as verify-only via the crypto API; full SM2 operations (encrypt, decrypt, key exchange) are available through the /dev/cmh_mgmt ioctl interface. Signed-off-by: Alex Ousherovitch Co-developed-by: Saravanakrishnan Krishnamoorthy Signed-off-by: Saravanakrishnan Krishnamoorthy --- drivers/crypto/cmh/Makefile | 3 +- drivers/crypto/cmh/cmh_main.c | 8 + drivers/crypto/cmh/cmh_pke_ecdsa.c | 610 +++++++++++++++++++++++++++++ 3 files changed, 620 insertions(+), 1 deletion(-) create mode 100644 drivers/crypto/cmh/cmh_pke_ecdsa.c diff --git a/drivers/crypto/cmh/Makefile b/drivers/crypto/cmh/Makefile index cdbcc8cdac5f..ae1f74a93e99 100644 --- a/drivers/crypto/cmh/Makefile +++ b/drivers/crypto/cmh/Makefile @@ -31,7 +31,8 @@ cmh-y :=3D \ cmh_ccp_poly.o \ cmh_rng.o \ cmh_pke_common.o \ - cmh_pke_rsa.o + cmh_pke_rsa.o \ + cmh_pke_ecdsa.o =20 # Management ioctl device (/dev/cmh_mgmt): key lifecycle, PKE, PQC ioctls. cmh-$(CONFIG_CRYPTO_DEV_CMH_MGMT) +=3D \ diff --git a/drivers/crypto/cmh/cmh_main.c b/drivers/crypto/cmh/cmh_main.c index 23346bbf11b0..b5b12f56a68e 100644 --- a/drivers/crypto/cmh/cmh_main.c +++ b/drivers/crypto/cmh/cmh_main.c @@ -292,6 +292,11 @@ static int cmh_probe(struct platform_device *pdev) if (ret) goto err_pke_rsa_register; =20 + /* Register PKE ECDSA/SM2 sig */ + ret =3D cmh_pke_ecdsa_register(); + if (ret) + goto err_pke_ecdsa_register; + /* Register key management device (/dev/cmh_mgmt) */ ret =3D cmh_mgmt_register(); if (ret) @@ -302,6 +307,8 @@ static int cmh_probe(struct platform_device *pdev) return 0; =20 err_mgmt_register: + cmh_pke_ecdsa_unregister(); +err_pke_ecdsa_register: cmh_pke_rsa_unregister(); err_pke_rsa_register: cmh_ccp_poly_unregister(); @@ -360,6 +367,7 @@ static void cmh_remove(struct platform_device *pdev) cfg =3D &dev->config; =20 cmh_mgmt_unregister(); + cmh_pke_ecdsa_unregister(); cmh_pke_rsa_unregister(); cmh_ccp_poly_unregister(); cmh_ccp_aead_unregister(); diff --git a/drivers/crypto/cmh/cmh_pke_ecdsa.c b/drivers/crypto/cmh/cmh_pk= e_ecdsa.c new file mode 100644 index 000000000000..a23828750368 --- /dev/null +++ b/drivers/crypto/cmh/cmh_pke_ecdsa.c @@ -0,0 +1,610 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- ECDSA / SM2 Signature Driver (sig_alg, synchronous) + * + * Registers "ecdsa-nist-p256", "ecdsa-nist-p384", and "ecdsa-nist-p521" + * sig algorithms with sign, verify, set_pub_key, and set_priv_key callbac= ks. + * Registers "sm2" as verify-only (set_pub_key + verify); SM2 sign is + * provided via the cmh_mgmt ioctl path in cmh_pke_sm2.c. + * + * In-kernel consumers typically use verify-only (module signatures, IMA), + * but we provide sign as well for completeness -- matching the CMH eSW + * capability. + * + * Key format: Public key =3D raw 04 || X || Y (uncompressed). + * Signature format: struct ecdsa_raw_sig (two u64[ECC_MAX_DIGITS] arrays + * in VLI format -- native byte order, LE digit order) for both sign + * output and verify input. This matches the kernel crypto sig API. + * + * Private key via cmh_key_ctx: raw keys written via SYS_REF_TEMP. + * Datastore-referenced keys are only reachable through the ioctl + * path (cmh_mgmt.c). + * + * SM2 note: The SM2 sig entry is verify-only (no sign/set_priv_key). + * SM2 signature verification requires the digest to be SM3(ZA || M) + * where ZA =3D SM3(ENTLA || IDA || a || b || xG || yG || xA || yA). + * The ZA identity pre-hash is the caller's responsibility; the driver + * passes the digest directly to the CMH eSW SM2 verify engine. + */ + +#include +#include +#include +#include +#include +#include +#include + +#include "cmh_pke.h" +#include "cmh_sys.h" +#include "cmh_sys_abi.h" +#include "cmh_txn.h" +#include "cmh_dma.h" +#include "cmh_key.h" + +/* + * Number of ECC digits needed for a given coordinate byte length. + * P-256: 4, P-384: 6, P-521/SM2(clen=3D68): 9. + */ +static inline unsigned int clen_to_ndigits(u32 clen) +{ + return DIV_ROUND_UP(clen, sizeof(u64)); +} + +struct cmh_ecdsa_tfm_ctx { + struct cmh_key_ctx key; /* private key (raw only) */ + u8 *pub_key; /* uncompressed (x, y) without 04 prefix */ + u32 pub_key_len; + u32 curve; /* PKE_CURVE_* */ + u32 clen; /* coordinate length in bytes */ +}; + +static inline struct cmh_ecdsa_tfm_ctx *cmh_ecdsa_ctx(struct crypto_sig *t= fm) +{ + return crypto_sig_ctx(tfm); +} + +/* + * Convert one VLI component (u64 array, LE digit order, native byte order) + * to big-endian byte array of @out_len bytes. The VLI value is right-ali= gned + * in the output (leading zero padding added if ndigits*8 < out_len). If = the + * VLI is wider than @out_len, the skipped leading bytes must be zero; a + * non-zero leading byte is malformed caller input and returns -EBADMSG. + * + * Return: 0 on success, -EBADMSG if the value does not fit in @out_len. + */ +static int ecdsa_vli_to_be(const u64 *vli, unsigned int ndigits, + u8 *out, unsigned int out_len) +{ + unsigned int full_len =3D ndigits * sizeof(u64); + unsigned int i, skip; + + memset(out, 0, out_len); + + if (full_len <=3D out_len) { + /* VLI fits entirely -- write at right end of out */ + u8 *dst =3D out + (out_len - full_len); + + for (i =3D 0; i < ndigits; i++) + put_unaligned_be64(vli[ndigits - 1 - i], + &dst[i * sizeof(u64)]); + } else { + /* + * VLI wider than out -- the skipped leading bytes MUST be + * zero. This runs on caller-supplied signature data, so a + * non-zero high byte is malformed input: reject it with + * -EBADMSG rather than WARN_ON_ONCE (which would let an + * unprivileged caller trip panic_on_warn). + */ + u8 tmp[ECC_MAX_BYTES]; + + for (i =3D 0; i < ndigits; i++) + put_unaligned_be64(vli[ndigits - 1 - i], + &tmp[i * sizeof(u64)]); + skip =3D full_len - out_len; + if (memchr_inv(tmp, 0, skip)) + return -EBADMSG; + memcpy(out, tmp + skip, out_len); + } + + return 0; +} + +/* + * Convert big-endian byte array to VLI (u64 array, LE digit order). + * Output is zero-filled to @max_digits entries. + */ +static void ecdsa_be_to_vli(const u8 *in, unsigned int in_len, + u64 *vli, unsigned int max_digits) +{ + u8 tmp[ECC_MAX_BYTES]; + unsigned int full_len; + unsigned int i; + + if (WARN_ON_ONCE(max_digits > ECC_MAX_DIGITS)) + max_digits =3D ECC_MAX_DIGITS; + + full_len =3D max_digits * sizeof(u64); + + memset(tmp, 0, full_len); + if (in_len <=3D full_len) + memcpy(tmp + (full_len - in_len), in, in_len); + else + memcpy(tmp, in + (in_len - full_len), full_len); + + for (i =3D 0; i < max_digits; i++) { + unsigned int off =3D (max_digits - 1 - i) * sizeof(u64); + + vli[i] =3D get_unaligned_be64(&tmp[off]); + } +} + +/* + * Extract raw (r || s) big-endian byte arrays from struct ecdsa_raw_sig. + * Each component is written as @clen bytes into @raw_rs. + */ +static int ecdsa_sig_to_raw(const void *src, unsigned int slen, + u8 *raw_rs, u32 clen) +{ + const struct ecdsa_raw_sig *sig =3D src; + unsigned int ndigits =3D clen_to_ndigits(clen); + int ret; + + if (slen !=3D sizeof(struct ecdsa_raw_sig)) + return -EINVAL; + + ret =3D ecdsa_vli_to_be(sig->r, ndigits, raw_rs, clen); + if (ret) + return ret; + return ecdsa_vli_to_be(sig->s, ndigits, raw_rs + clen, clen); +} + +/* + * Encode raw (r || s) big-endian byte arrays into struct ecdsa_raw_sig. + * Returns sizeof(struct ecdsa_raw_sig) on success. + */ +static int ecdsa_raw_to_sig(const u8 *raw_rs, u32 clen, + void *dst, unsigned int dlen) +{ + struct ecdsa_raw_sig *sig =3D dst; + + if (dlen < sizeof(struct ecdsa_raw_sig)) + return -ENOSPC; + + memset(sig, 0, sizeof(*sig)); + ecdsa_be_to_vli(raw_rs, clen, sig->r, ECC_MAX_DIGITS); + ecdsa_be_to_vli(raw_rs + clen, clen, sig->s, ECC_MAX_DIGITS); + return sizeof(struct ecdsa_raw_sig); +} + +/* + * ECDSA verify (synchronous sig_alg) + * + * @src: struct ecdsa_raw_sig (VLI format) + * @slen: signature length (must be sizeof(struct ecdsa_raw_sig)) + * @digest: hash digest + * @dlen: digest length + * + * Returns 0 on successful verification, negative errno on failure. + */ +static int cmh_ecdsa_verify(struct crypto_sig *tfm, + const void *src, unsigned int slen, + const void *digest, unsigned int dlen) +{ + struct cmh_ecdsa_tfm_ctx *ctx =3D cmh_ecdsa_ctx(tfm); + u32 clen =3D ctx->clen; + u32 sig_raw_len =3D 2 * clen; + u32 copy_len =3D min_t(u32, dlen, clen); + struct core_dispatch d =3D cmh_core_select_instance(CMH_CORE_PKE); + struct vcq_cmd vcq[PKE_VCQ_CMDS_MIN]; + u8 *sig_raw =3D NULL, *dig_buf =3D NULL, *pk_buf =3D NULL, *rp_buf =3D NU= LL; + dma_addr_t pk_dma, dig_dma, sig_dma, rp_dma; + int ret; + + if (!ctx->pub_key) + return -EINVAL; + + sig_raw =3D kzalloc(sig_raw_len, GFP_KERNEL); + dig_buf =3D kzalloc(clen, GFP_KERNEL); + pk_buf =3D kmemdup(ctx->pub_key, ctx->pub_key_len, GFP_KERNEL); + rp_buf =3D kzalloc(clen, GFP_KERNEL); + if (!sig_raw || !dig_buf || !pk_buf || !rp_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + /* Extract raw (r, s) big-endian from VLI signature */ + ret =3D ecdsa_sig_to_raw(src, slen, sig_raw, clen); + if (ret) + goto out_free; + + /* + * Truncate or zero-pad digest to clen bytes, right-aligned. + * Matches ECDSA bits2int: use leftmost min(dlen, clen) bytes, + * zero-pad on the left when dlen < clen. + */ + memcpy(dig_buf + (clen - copy_len), digest, copy_len); + + pk_dma =3D cmh_dma_map_single(pk_buf, ctx->pub_key_len, DMA_TO_DEVICE); + dig_dma =3D cmh_dma_map_single(dig_buf, clen, DMA_TO_DEVICE); + sig_dma =3D cmh_dma_map_single(sig_raw, sig_raw_len, DMA_TO_DEVICE); + rp_dma =3D cmh_dma_map_single(rp_buf, clen, DMA_FROM_DEVICE); + + if (cmh_dma_map_error(pk_dma) || cmh_dma_map_error(dig_dma) || + cmh_dma_map_error(sig_dma) || cmh_dma_map_error(rp_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + vcq_set_header(&vcq[0], PKE_VCQ_CMDS_MIN); + vcq_add_pke_ecdsa_verify(&vcq[1], d.core_id, ctx->curve, clen, + pk_dma, dig_dma, sig_dma, rp_dma, + pke_swap_flags(ctx->curve)); + vcq_add_pke_flush(&vcq[2], d.core_id); + + ret =3D cmh_tm_submit_sync_mbx(vcq, PKE_VCQ_CMDS_MIN, 1, d.mbx_idx); + +out_unmap: + if (!cmh_dma_map_error(rp_dma)) + cmh_dma_unmap_single(rp_dma, clen, DMA_FROM_DEVICE); + if (!cmh_dma_map_error(sig_dma)) + cmh_dma_unmap_single(sig_dma, sig_raw_len, DMA_TO_DEVICE); + if (!cmh_dma_map_error(dig_dma)) + cmh_dma_unmap_single(dig_dma, clen, DMA_TO_DEVICE); + if (!cmh_dma_map_error(pk_dma)) + cmh_dma_unmap_single(pk_dma, ctx->pub_key_len, DMA_TO_DEVICE); + +out_free: + kfree(rp_buf); + kfree(pk_buf); + kfree(sig_raw); + kfree(dig_buf); + return ret; +} + +/* + * ECDSA sign (synchronous sig_alg) + * + * @src: hash digest + * @slen: digest length + * @dst: output buffer for struct ecdsa_raw_sig (VLI format) + * @dlen: output buffer length + * + * Returns sizeof(struct ecdsa_raw_sig) on success, negative errno on fail= ure. + */ +static int cmh_ecdsa_sign(struct crypto_sig *tfm, + const void *src, unsigned int slen, + void *dst, unsigned int dlen) +{ + struct cmh_ecdsa_tfm_ctx *ctx =3D cmh_ecdsa_ctx(tfm); + u32 clen =3D ctx->clen; + u32 sig_raw_len =3D 2 * clen; + u32 copy_len =3D min_t(u32, slen, clen); + struct core_dispatch dd; + struct vcq_cmd vcq[PKE_VCQ_CMDS_MAX]; + u8 *dig_buf =3D NULL, *sig_buf =3D NULL; + dma_addr_t dig_dma, sig_dma; + int ret, idx; + + if (ctx->key.mode !=3D CMH_KEY_RAW) + return -EINVAL; + if (dlen < sizeof(struct ecdsa_raw_sig)) + return -EINVAL; + + dig_buf =3D kzalloc(clen, GFP_KERNEL); + sig_buf =3D kzalloc(sig_raw_len, GFP_KERNEL); + if (!dig_buf || !sig_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + /* + * Truncate or zero-pad digest to clen bytes, right-aligned. + * Matches ECDSA bits2int: use leftmost min(slen, clen) bytes, + * zero-pad on the left when slen < clen. + */ + memcpy(dig_buf + (clen - copy_len), src, copy_len); + + dig_dma =3D cmh_dma_map_single(dig_buf, clen, DMA_TO_DEVICE); + sig_dma =3D cmh_dma_map_single(sig_buf, sig_raw_len, DMA_FROM_DEVICE); + + if (cmh_dma_map_error(dig_dma) || cmh_dma_map_error(sig_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + dd =3D cmh_core_select_instance(CMH_CORE_PKE); + + idx =3D 1; + vcq_add_sys_write(&vcq[idx], SYS_REF_TEMP, ctx->key.raw.dma, + SYS_REF_NONE, ctx->key.raw.len, + ctx->key.raw.sys_type); + vcq[idx].id |=3D pke_swap_flags(ctx->curve); + idx++; + vcq_add_pke_ecdsa_sign(&vcq[idx++], dd.core_id, ctx->curve, clen, + dig_dma, sig_dma, SYS_REF_TEMP, + clen, pke_swap_flags(ctx->curve)); + vcq_add_pke_flush(&vcq[idx++], dd.core_id); + vcq_set_header(&vcq[0], idx); + + ret =3D cmh_tm_submit_sync_mbx(vcq, idx, 1, dd.mbx_idx); + if (!ret) { + /* Sync bounce buffer so CPU sees the DMA-written signature */ + cmh_dma_sync_for_cpu(sig_dma, sig_raw_len, DMA_FROM_DEVICE); + + /* Encode raw (r||s) into VLI ecdsa_raw_sig for kernel API */ + ret =3D ecdsa_raw_to_sig(sig_buf, clen, dst, dlen); + } + +out_unmap: + if (!cmh_dma_map_error(sig_dma)) + cmh_dma_unmap_single(sig_dma, sig_raw_len, DMA_FROM_DEVICE); + if (!cmh_dma_map_error(dig_dma)) + cmh_dma_unmap_single(dig_dma, clen, DMA_TO_DEVICE); + +out_free: + kfree(sig_buf); + kfree(dig_buf); + return ret; +} + +static int cmh_ecdsa_set_pub_key(struct crypto_sig *tfm, + const void *key, unsigned int keylen) +{ + struct cmh_ecdsa_tfm_ctx *ctx =3D cmh_ecdsa_ctx(tfm); + const u8 *d =3D key; + u32 clen =3D ctx->clen; + u32 raw_clen; + + /* Accept 04 || X || Y (uncompressed point) */ + if (keylen < 1 || d[0] !=3D 0x04) + return -EINVAL; + d++; + keylen--; + + if (keylen & 1) + return -EINVAL; + raw_clen =3D keylen / 2; + + /* + * Kernel passes ceil(bits/8) per coordinate (e.g. 66 for P-521), + * but our HW ABI uses clen (ALIGN(66,4)=3D68 for P-521). + * Accept raw_clen <=3D clen and zero-pad on the left. + */ + if (raw_clen > clen || raw_clen =3D=3D 0) + return -EINVAL; + + kfree(ctx->pub_key); + ctx->pub_key =3D NULL; + ctx->pub_key_len =3D 0; + + ctx->pub_key =3D kzalloc(2 * clen, GFP_KERNEL); + if (!ctx->pub_key) + return -ENOMEM; + + /* Right-align each coordinate to clen bytes */ + memcpy(ctx->pub_key + (clen - raw_clen), d, raw_clen); + memcpy(ctx->pub_key + clen + (clen - raw_clen), d + raw_clen, + raw_clen); + ctx->pub_key_len =3D 2 * clen; + return 0; +} + +static int cmh_ecdsa_set_priv_key(struct crypto_sig *tfm, + const void *key, unsigned int keylen) +{ + struct cmh_ecdsa_tfm_ctx *ctx =3D cmh_ecdsa_ctx(tfm); + u32 clen =3D ctx->clen; + u32 api_len =3D DIV_ROUND_UP(pke_curve_bits(ctx->curve), 8); + u8 padded[ECC_MAX_BYTES]; + int ret; + + /* + * The crypto sig API passes the scalar as ceil(field_bits/8) bytes + * (e.g. 66 for P-521), which may be shorter than the HW width clen + * (ALIGN(.,4) =3D 68 for P-521). Accept exactly that width and + * left-pad to clen. A strict "keylen =3D=3D clen" wrongly rejected the + * valid 66-byte P-521 key; accepting "keylen <=3D clen" wrongly + * admitted malformed short keys. + */ + if (keylen !=3D api_len) + return -EINVAL; + + if (keylen =3D=3D clen) + return cmh_key_setkey_raw(&ctx->key, key, keylen, CORE_ID_PKE); + + memset(padded, 0, clen); + memcpy(padded + (clen - keylen), key, keylen); + ret =3D cmh_key_setkey_raw(&ctx->key, padded, clen, CORE_ID_PKE); + memzero_explicit(padded, clen); + return ret; +} + +static unsigned int cmh_ecdsa_key_size(struct crypto_sig *tfm) +{ + struct cmh_ecdsa_tfm_ctx *ctx =3D cmh_ecdsa_ctx(tfm); + + /* crypto_sig_keysize() returns bits, not bytes */ + return pke_curve_bits(ctx->curve); +} + +static unsigned int cmh_ecdsa_max_size(struct crypto_sig *tfm) +{ + return sizeof(struct ecdsa_raw_sig); +} + +static unsigned int cmh_ecdsa_digest_size(struct crypto_sig *tfm) +{ + /* + * Accept digests up to SHA-512 (64 bytes). Digests longer + * than the curve order are truncated per ECDSA bits2int. + * Matches kernel ecdsa_digest_size(). + */ + return SHA512_DIGEST_SIZE; +} + +static int cmh_ecdsa_p256_init(struct crypto_sig *tfm) +{ + struct cmh_ecdsa_tfm_ctx *ctx =3D cmh_ecdsa_ctx(tfm); + + memset(ctx, 0, sizeof(*ctx)); + ctx->curve =3D PKE_CURVE_P256; + ctx->clen =3D pke_curve_clen(PKE_CURVE_P256); + return 0; +} + +static int cmh_ecdsa_p384_init(struct crypto_sig *tfm) +{ + struct cmh_ecdsa_tfm_ctx *ctx =3D cmh_ecdsa_ctx(tfm); + + memset(ctx, 0, sizeof(*ctx)); + ctx->curve =3D PKE_CURVE_P384; + ctx->clen =3D pke_curve_clen(PKE_CURVE_P384); + return 0; +} + +static int cmh_ecdsa_p521_init(struct crypto_sig *tfm) +{ + struct cmh_ecdsa_tfm_ctx *ctx =3D cmh_ecdsa_ctx(tfm); + + memset(ctx, 0, sizeof(*ctx)); + ctx->curve =3D PKE_CURVE_P521; + ctx->clen =3D pke_curve_clen(PKE_CURVE_P521); + return 0; +} + +static int cmh_sm2_init(struct crypto_sig *tfm) +{ + struct cmh_ecdsa_tfm_ctx *ctx =3D cmh_ecdsa_ctx(tfm); + + memset(ctx, 0, sizeof(*ctx)); + ctx->curve =3D PKE_CURVE_SM2; + ctx->clen =3D pke_curve_clen(PKE_CURVE_SM2); + return 0; +} + +static void cmh_ecdsa_exit(struct crypto_sig *tfm) +{ + struct cmh_ecdsa_tfm_ctx *ctx =3D cmh_ecdsa_ctx(tfm); + + cmh_key_destroy(&ctx->key); + kfree(ctx->pub_key); + ctx->pub_key =3D NULL; +} + +static struct sig_alg cmh_ecdsa_algs[] =3D { + { + .sign =3D cmh_ecdsa_sign, + .verify =3D cmh_ecdsa_verify, + .set_pub_key =3D cmh_ecdsa_set_pub_key, + .set_priv_key =3D cmh_ecdsa_set_priv_key, + .key_size =3D cmh_ecdsa_key_size, + .max_size =3D cmh_ecdsa_max_size, + .digest_size =3D cmh_ecdsa_digest_size, + .init =3D cmh_ecdsa_p256_init, + .exit =3D cmh_ecdsa_exit, + .base =3D { + .cra_name =3D "ecdsa-nist-p256", + .cra_driver_name =3D "rambus-cmh-ecdsa-nist-p256", + .cra_priority =3D 300, + .cra_module =3D THIS_MODULE, + .cra_ctxsize =3D sizeof(struct cmh_ecdsa_tfm_ctx), + }, + }, + { + .sign =3D cmh_ecdsa_sign, + .verify =3D cmh_ecdsa_verify, + .set_pub_key =3D cmh_ecdsa_set_pub_key, + .set_priv_key =3D cmh_ecdsa_set_priv_key, + .key_size =3D cmh_ecdsa_key_size, + .max_size =3D cmh_ecdsa_max_size, + .digest_size =3D cmh_ecdsa_digest_size, + .init =3D cmh_ecdsa_p384_init, + .exit =3D cmh_ecdsa_exit, + .base =3D { + .cra_name =3D "ecdsa-nist-p384", + .cra_driver_name =3D "rambus-cmh-ecdsa-nist-p384", + .cra_priority =3D 300, + .cra_module =3D THIS_MODULE, + .cra_ctxsize =3D sizeof(struct cmh_ecdsa_tfm_ctx), + }, + }, + { + .sign =3D cmh_ecdsa_sign, + .verify =3D cmh_ecdsa_verify, + .set_pub_key =3D cmh_ecdsa_set_pub_key, + .set_priv_key =3D cmh_ecdsa_set_priv_key, + .key_size =3D cmh_ecdsa_key_size, + .max_size =3D cmh_ecdsa_max_size, + .digest_size =3D cmh_ecdsa_digest_size, + .init =3D cmh_ecdsa_p521_init, + .exit =3D cmh_ecdsa_exit, + .base =3D { + .cra_name =3D "ecdsa-nist-p521", + .cra_driver_name =3D "rambus-cmh-ecdsa-nist-p521", + .cra_priority =3D 300, + .cra_module =3D THIS_MODULE, + .cra_ctxsize =3D sizeof(struct cmh_ecdsa_tfm_ctx), + }, + }, + { + .verify =3D cmh_ecdsa_verify, + .set_pub_key =3D cmh_ecdsa_set_pub_key, + .key_size =3D cmh_ecdsa_key_size, + .max_size =3D cmh_ecdsa_max_size, + .digest_size =3D cmh_ecdsa_digest_size, + .init =3D cmh_sm2_init, + .exit =3D cmh_ecdsa_exit, + .base =3D { + .cra_name =3D "sm2", + .cra_driver_name =3D "rambus-cmh-sm2", + .cra_priority =3D 300, + .cra_module =3D THIS_MODULE, + .cra_ctxsize =3D sizeof(struct cmh_ecdsa_tfm_ctx), + }, + }, +}; + +/** + * cmh_pke_ecdsa_register() - Register ECDSA/SM2 sig algorithms with the c= rypto framework + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_pke_ecdsa_register(void) +{ + int ret, i; + + if (!cmh_core_present(CMH_CORE_PKE)) + return 0; + + for (i =3D 0; i < ARRAY_SIZE(cmh_ecdsa_algs); i++) { + ret =3D crypto_register_sig(&cmh_ecdsa_algs[i]); + if (ret) { + dev_err(cmh_dev(), "cmh: failed to register %s (%d)\n", + cmh_ecdsa_algs[i].base.cra_name, ret); + goto err_unregister; + } + } + + return 0; + +err_unregister: + while (i--) + crypto_unregister_sig(&cmh_ecdsa_algs[i]); + return ret; +} + +/** + * cmh_pke_ecdsa_unregister() - Unregister ECDSA/SM2 sig algorithms from t= he crypto framework + */ +void cmh_pke_ecdsa_unregister(void) +{ + int i =3D ARRAY_SIZE(cmh_ecdsa_algs); + + if (!cmh_core_present(CMH_CORE_PKE)) + return; + + while (i--) + crypto_unregister_sig(&cmh_ecdsa_algs[i]); +} --=20 2.43.7 From nobody Fri Sep 25 01:20:34 2026 Received: from SJ2PR03CU001.outbound.protection.outlook.com (mail-westusazon11022107.outbound.protection.outlook.com [52.101.43.107]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C36304B0E49; Thu, 17 Sep 2026 22:59:42 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.43.107 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789685987; cv=fail; b=SndzG10/yxUtl6YWIIAmhnwoKYfd9X7ZUlID+sNWsobacol0Q1WSOSztlMXP15Pl6H/B4LIKPEXs6c+EK3yLF3b/3LlInbSbFkbAK+2k9g0YqBtDbDu1nRu40rPFj3QIR0OTCULgsTJr0y79zqkS086duNVjUqtb76Db0iFGKLQ= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789685987; c=relaxed/simple; bh=pq9qDZb1y7iUwkrb8udABt2TRDiarngJsZtOyamc3T8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=KaMZHrghDPosqG0t2Q07gEehUFypJjNZA6kKiP5x+eTYbFS+sOIDwvRLAcZsg/U55qxKdOkOtz2YCA2bz5ft5O9IrqzEa2a7UHbr9BGeoMH4VjT907FrbI0aJRxPdjC9ak8fcAAXQcv9E3cdZTYhamOQTIMNYZi1ueIop1LI7BQ= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com; spf=fail smtp.mailfrom=rambus.com; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b=yjKVZfvj; arc=fail smtp.client-ip=52.101.43.107 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=rambus.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b="yjKVZfvj" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=S6okVh1VBv/L9SWHLVM7TYh3JFE0LrRCQPkF5TDiwCigjd0OXwTDj9eqehuWc9qtH8lzeTdRy8Th3qwiy47BY6MSulWeFfDjZx35inNUZrwEBENBMomcvNv30RCD8fwzmxaOTAMIDseNecdQnFFbuYOeImqIt7UInTHQLwBZgtIYUEkGIxQ5L1Qz4yZAdIDPP/cFx1LD8X1l4ZW4Bc2pMOWemnBmI2ADTkBaXOius/MGLqZhnR2RKeJkiHHuaL6RZTMEmKBK+0HWSGCNi6hc22DsCb9ceXwD90RqpK9Z9ekwzal41nnjs7cy2Mpw5UDzKlA4l0bdhqWlQbZCAbW5OA== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=x67nJblAhtktgljYbUc3Jl9/SnurDy8yuIGdUCdZl5M=; b=u4VGeqZaiHK8Dm3PPNnzl5Zgdv5jJGYcb1qF4jyGij68FnuOAF4xqnX6d5v8A3z68TKGvwH/Sgkl/ivwh3N7BntWHzYGWImYeBWw+MmgE7ZnmYXzZlKngfKi3wVTV9phWdfzE7CsL3j+LdH4sr5xe0xF5lQt958gR/CFfux0zfcJ8IKREc/jayMiiBpAe/Ocov/tMJ9wGzi1BBsoFdy3KudKV2+2ZUD0UQrReRMkYCMlqV8ycB0VWBafWABpVYWoG7b2lqmEbCEZYEv5oXaWzG18dXEdkSnSdBx+TI7o/7NB1s6Vadf0u8b+Tp0IlqF4Yj6OZNtBSYNm9W0GzUYEyQ== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 192.86.86.210) smtp.rcpttodomain=cryptography.com smtp.mailfrom=rambus.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=rambus.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=rambus.com; s=selector1; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=x67nJblAhtktgljYbUc3Jl9/SnurDy8yuIGdUCdZl5M=; b=yjKVZfvje4iQf5En3EMbHMnojXNbPV0sYYgSvM5ZdQSNu4IktcV4FRkPD+rDpRzSDiecdt++4qKrgm0K3x5oiuCM+weds/zZkao1CRNElLUu7M7x3WUddc2KNJ5K/gooG0qnpIWmJvVAphizkPAbGaPD2/5HXVUKBfDKZxnWxgX/d30zIRbobumuQvpk90aZ/knDX9WfUt0JqzXupP52sWMwlpwHTpw/mD1czrrB9a9Rz50oglrTJhEfhVEmaqrS7dw5hyI8vpyKsoiOAFXoFw3A3P8vGQ5Do/9gsQVXKHi2KQhFp8W9NXrQNsHghD2U39tdqwx7cM+BX9ujjJJb1Q== Received: from BN0PR10CA0029.namprd10.prod.outlook.com (2603:10b6:408:143::7) by PH0PR04MB8293.namprd04.prod.outlook.com (2603:10b6:510:107::17) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.428.9; Thu, 17 Sep 2026 22:59:35 +0000 Received: from BN2PEPF0000A800.namprd02.prod.outlook.com (2603:10b6:408:143:cafe::81) by BN0PR10CA0029.outlook.office365.com (2603:10b6:408:143::7) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.428.12 via Frontend Transport; Thu, 17 Sep 2026 22:59:35 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 192.86.86.210) smtp.mailfrom=rambus.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=rambus.com; Received-SPF: Pass (protection.outlook.com: domain of rambus.com designates 192.86.86.210 as permitted sender) receiver=protection.outlook.com; client-ip=192.86.86.210; helo=hqxsv-psmtppxy02.rambus.com; pr=C Received: from hqxsv-psmtppxy02.rambus.com (192.86.86.210) by BN2PEPF0000A800.mail.protection.outlook.com (10.167.245.167) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.451.8 via Frontend Transport; Thu, 17 Sep 2026 22:59:34 +0000 Received: from hqxsv-cmdev3-aousherovitch.rambus.com (hqn-lb-int-float.rambus.com [10.12.20.20]) by hqxsv-psmtppxy02.rambus.com (Postfix) with ESMTP id 6E658180176C; Thu, 17 Sep 2026 22:59:32 +0000 (UTC) From: Alex Ousherovitch To: Albert Ou , Alex Ousherovitch , Conor Dooley , "David S. Miller" , Herbert Xu , Jonathan Corbet , Krzysztof Kozlowski , Palmer Dabbelt , Paul Walmsley , Rob Herring , Saravanakrishnan Krishnamoorthy , Shuah Khan Cc: Alexandre Ghiti , devicetree@vger.kernel.org, Joel Wittenauer , linux-api@vger.kernel.org, linux-crypto@vger.kernel.org, linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-riscv@lists.infradead.org, Shuah Khan , Thi Nguyen Subject: [PATCH v5 14/19] crypto: cmh - add ECDH/X25519 kpp Date: Thu, 17 Sep 2026 15:59:23 -0700 Message-ID: <20260917225929.2494111-15-aousherovitch@rambus.com> X-Mailer: git-send-email 2.43.7 In-Reply-To: <20260917225929.2494111-1-aousherovitch@rambus.com> References: <20260917225929.2494111-1-aousherovitch@rambus.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: BN2PEPF0000A800:EE_|PH0PR04MB8293:EE_ X-MS-Office365-Filtering-Correlation-Id: 132b42bc-18de-4378-2902-08df150f57b4 X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|1800799024|23010399003|82310400026|7416014|376014|36860700016|10067099003|6133799003|3023799007|22082099003|18002099003|5023799004|56012099006|11063799006|921020; X-Microsoft-Antispam-Message-Info: HHAZ55Z4ZUi2UWJGbRuj5ctfGihZ239LFvmyarPcBsxW3QohBbLsuycHJ2GWZgJDTFpwnAeqiLLyybiI+al6+2CrfLu5k6c29dB/JfMCNuFKxe/F+lF5g88POmGMh0akwkZC0X8itgreJlfKrGy38rWWoRMyl9hgfKHARsr/iImiuvm+Ql8Y0hnmWukfDyWNoB1fPxGoZaWCTPpJTMwgEeOg4DqXITJm8jdKxUH8sxEvLTD5lAFMCnq3846QPyigmtvXrb0qxHB9QDU7XJ1o72yBJ4rdjiw0K0Ho9JVpVUcSYgU65d4ExKV+14kf9dxwRLV/L1cD/cuGhE9gpYypqvanullXDFIoBwPe8NTTU3WlGAPlgzZYn+RRfby73Ww609BdNsp38KERC4r/86tm3CZkK+7anSzi7FpqXLc1NwfP3XvmM0hwvFfeROsKhvAujkMcYFHAqNCg9++MBukxHiggpOwMseqt2rH28eauL2MJKP6TaVLP8d1iFjF8SZWPzuTzD2q1KJj8qmj86QgKx2d7X7U5QgmByO5kHBir7LRKzGfqyIuiu6me5/CwPjpfo8bBK550osy9C+u6DhdPuA5hFkEEyfUUSDlw9CCZ9EUg8cHPnZzVK3uC+ih0X+gBUbY6hQ0imh8TuWwnNrqJrCA3XV/qc4zc2PIODGvfF2QDxFmA+5bFQDUWb0jhPoKACzT60BjoQkeXLX/AlesVN2Xuako2J9/0klX/MqTg9wEhaHiQI7VMWa/P/OB0Fk2G X-Forefront-Antispam-Report: CIP:192.86.86.210;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:hqxsv-psmtppxy02.rambus.com;PTR:InfoDomainNonexistent;CAT:NONE;SFS:(13230040)(1800799024)(23010399003)(82310400026)(7416014)(376014)(36860700016)(10067099003)(6133799003)(3023799007)(22082099003)(18002099003)(5023799004)(56012099006)(11063799006)(921020);DIR:OUT;SFP:1102; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: 9IM5FyIlSKVVRqpjsctPDAGgIJsgDgW4VoBqM6LWDZTyuL9C1yvoXPsub00evOSulz3exjPp+T7NohFXh2Yyq5VznOd5Rfk0IBotJSVvrR8A3UcMIwFFCjDE2J7S27aqXbk7axvXkm4cy16McQKXpsah/etfC2DJ9AYAUp/SCUdfzRpkDBRhu5I8cz2ktjA7nt9i0BKpDUIQUQVrbRSNxul6s2JNxeyytcVY1sKDkVr2/bgf+/iLE0UwOOt49gSL8K/O4GZb1WK1PNh5W9aZqVm8n5u/fLGu2WVtfaIY5fV5mqu7gc66IjQbiiLIAUdh+Xts2pEatze1x3FXmKbVeBnipK9FEt1Oyif2xTIfXijBer5JAWS+FgDm0nOfDnHm3UfBb/0iqAURXsYbvZEkRX3nzA4NfTSWICyTr55oSJtDy+niRyuzfJtaqmj/umY0 X-OriginatorOrg: rambus.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 17 Sep 2026 22:59:34.7700 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: 132b42bc-18de-4378-2902-08df150f57b4 X-MS-Exchange-CrossTenant-Id: bd0ba799-c2b9-413c-9c56-5d1731c4827c X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=bd0ba799-c2b9-413c-9c56-5d1731c4827c;Ip=[192.86.86.210];Helo=[hqxsv-psmtppxy02.rambus.com] X-MS-Exchange-CrossTenant-AuthSource: BN2PEPF0000A800.namprd02.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: PH0PR04MB8293 Content-Type: text/plain; charset="utf-8" Register ECDH and X25519 kpp algorithms using the CMH PKE core. Supports P-256, P-384, and Curve25519 for key agreement. Signed-off-by: Alex Ousherovitch Co-developed-by: Saravanakrishnan Krishnamoorthy Signed-off-by: Saravanakrishnan Krishnamoorthy --- drivers/crypto/cmh/Makefile | 3 +- drivers/crypto/cmh/cmh_main.c | 8 + drivers/crypto/cmh/cmh_pke_ecdh.c | 817 ++++++++++++++++++++++++++++++ 3 files changed, 827 insertions(+), 1 deletion(-) create mode 100644 drivers/crypto/cmh/cmh_pke_ecdh.c diff --git a/drivers/crypto/cmh/Makefile b/drivers/crypto/cmh/Makefile index ae1f74a93e99..da420d0ec1fe 100644 --- a/drivers/crypto/cmh/Makefile +++ b/drivers/crypto/cmh/Makefile @@ -32,7 +32,8 @@ cmh-y :=3D \ cmh_rng.o \ cmh_pke_common.o \ cmh_pke_rsa.o \ - cmh_pke_ecdsa.o + cmh_pke_ecdsa.o \ + cmh_pke_ecdh.o =20 # Management ioctl device (/dev/cmh_mgmt): key lifecycle, PKE, PQC ioctls. cmh-$(CONFIG_CRYPTO_DEV_CMH_MGMT) +=3D \ diff --git a/drivers/crypto/cmh/cmh_main.c b/drivers/crypto/cmh/cmh_main.c index b5b12f56a68e..93d939e8fc4e 100644 --- a/drivers/crypto/cmh/cmh_main.c +++ b/drivers/crypto/cmh/cmh_main.c @@ -297,6 +297,11 @@ static int cmh_probe(struct platform_device *pdev) if (ret) goto err_pke_ecdsa_register; =20 + /* Register PKE ECDH/X25519 kpp */ + ret =3D cmh_pke_ecdh_register(); + if (ret) + goto err_pke_ecdh_register; + /* Register key management device (/dev/cmh_mgmt) */ ret =3D cmh_mgmt_register(); if (ret) @@ -307,6 +312,8 @@ static int cmh_probe(struct platform_device *pdev) return 0; =20 err_mgmt_register: + cmh_pke_ecdh_unregister(); +err_pke_ecdh_register: cmh_pke_ecdsa_unregister(); err_pke_ecdsa_register: cmh_pke_rsa_unregister(); @@ -367,6 +374,7 @@ static void cmh_remove(struct platform_device *pdev) cfg =3D &dev->config; =20 cmh_mgmt_unregister(); + cmh_pke_ecdh_unregister(); cmh_pke_ecdsa_unregister(); cmh_pke_rsa_unregister(); cmh_ccp_poly_unregister(); diff --git a/drivers/crypto/cmh/cmh_pke_ecdh.c b/drivers/crypto/cmh/cmh_pke= _ecdh.c new file mode 100644 index 000000000000..56c0f2629c6b --- /dev/null +++ b/drivers/crypto/cmh/cmh_pke_ecdh.c @@ -0,0 +1,817 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- ECDH / X25519 kpp Driver + * + * Registers "ecdh-nist-p256", "ecdh-nist-p384", and "curve25519" + * kpp algorithms with priority 300. + * + * - set_secret: decodes private key from kpp_secret + ecdh struct + * (NIST curves) or raw 32-byte scalar (Curve25519). + * Stores in cmh_key_ctx: raw keys written via SYS_REF_TEMP. + * Datastore-referenced keys are only reachable through the ioctl + * path (cmh_mgmt.c). + * + * - generate_public_key: PKE_CMD_ECDH_KEYGEN -> outputs X coordinate + * (NIST Weierstrass) or full public key (Edwards/Montgomery). + * For NIST curves, we generate X||Y by calling ECDSA_PUBGEN instead, + * matching the kernel ecdh.c pattern that outputs uncompressed X||Y. + * + * - compute_shared_secret: PKE_CMD_ECDH -> shared secret X coordinate. + */ + +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "cmh_pke.h" +#include "cmh_sys.h" +#include "cmh_sys_abi.h" +#include "cmh_txn.h" +#include "cmh_dma.h" +#include "cmh_key.h" + +/* + * ECDH key format: kpp_secret header + key_size(u16) + key data. + * We decode this inline to avoid depending on CONFIG_CRYPTO_ECDH. + */ +#define ECDH_KPP_SECRET_MIN_SIZE (sizeof(struct kpp_secret) + sizeof(unsig= ned short)) + +struct cmh_ecdh_tfm_ctx { + struct cmh_key_ctx key; + u32 curve; /* PKE_CURVE_* */ + u32 clen; /* coordinate length in bytes */ +}; + +static inline struct cmh_ecdh_tfm_ctx *cmh_ecdh_ctx(struct crypto_kpp *tfm) +{ + return kpp_tfm_ctx(tfm); +} + +/* + * Per-request context for ECDH/X25519 operations. + * + * generate_public_key and compute_shared_secret are both single-phase + * async VCQs. compute_shared_secret writes the private key into a + * per-mailbox scratch datastore object and targets SYS_REF_TEMP for the + * shared secret (see cmh_ecdh_key_scratch below). + * + * Both operations reference the transform's persistent private-key DMA + * buffer (ctx->key.raw.dma) directly in their VCQ, and that buffer is + * mapped for the lifetime of the transform (set_secret) and freed only by + * a re-key or cmh_key_destroy(). This is safe because the crypto KPP API + * serialises operations on a single transform: a caller must not re-key + * (set_secret) while a request is in flight, and async callers wait for + * the request completion before issuing the next operation. There is + * therefore no window in which an in-flight VCQ can reference a key buffer + * that a concurrent set_secret has freed. + */ +struct cmh_ecdh_reqctx { + /* Buffers */ + u8 *pk_buf; /* keygen: output public key */ + u8 *peer_buf; /* compute: peer public key */ + u8 *ss_buf; /* compute: shared secret output */ + /* DMA handles */ + dma_addr_t pk_dma; + dma_addr_t peer_dma; + dma_addr_t ss_dma; + /* Sizes and params */ + u32 out_len; /* keygen: public key size */ + u32 clen; + u32 peer_len; +}; + +/* + * Per-mailbox scratch datastore object holding the private key for + * compute_shared_secret. + * + * The CMH ECDH command writes its shared secret into a datastore object, + * not directly to host DMA. Allocating that object per request via + * SYS_CMD_NEW permanently consumes a slot every call, because the eSW + * datastore is a bump allocator with no per-object free. + * + * Instead we sys_write the private key into a persistent per-mailbox + * object (owned by that mailbox) and target SYS_REF_TEMP for the shared + * secret, which the eSW reclaims when sys_data reads it back. One object + * per mailbox preserves the round-robin PKE dispatch; the eSW runs a VCQ + * and its children to completion per mailbox, so concurrent requests on + * the same mailbox never observe each other's key in the shared object. + * + * Created on the first key set (deferred from register) and scrubbed at + * unregister, so a bring-up that never keys ECDH leaves the datastore + * empty for a DS import. Indexed by mbx_idx; entries for mailboxes + * without a PKE instance stay zero. + */ +static u64 *cmh_ecdh_key_scratch; +static u32 cmh_ecdh_scratch_count; +static bool cmh_ecdh_scratch_ready; /* DS objects created (one-shot) */ +static DEFINE_MUTEX(cmh_ecdh_scratch_lock); + +static int cmh_ecdh_scratch_ensure(void); + +/* + * cmh_ecdh_commit_key() - Ensure the key scratch exists, then stash the r= aw + * private key. Deferring scratch creation to the first key set keeps the + * datastore empty until ECDH is actually used. + */ +static int cmh_ecdh_commit_key(struct cmh_ecdh_tfm_ctx *ctx, + const u8 *key, u32 len) +{ + int ret =3D cmh_ecdh_scratch_ensure(); + + if (ret) + return ret; + return cmh_key_setkey_raw(&ctx->key, key, len, CORE_ID_PKE); +} + +/* + * set_secret: NIST curves decode kpp_secret + u16 key_size + raw scalar. + * Curve25519 uses raw 32-byte scalar directly. + */ +static int cmh_ecdh_set_secret_nist(struct crypto_kpp *tfm, + const void *buf, unsigned int len) +{ + struct cmh_ecdh_tfm_ctx *ctx =3D cmh_ecdh_ctx(tfm); + const u8 *ptr =3D buf; + struct kpp_secret secret; + unsigned short key_size; + int ret; + + if (!buf || len < ECDH_KPP_SECRET_MIN_SIZE) + return -EINVAL; + + memcpy(&secret, ptr, sizeof(secret)); + ptr +=3D sizeof(secret); + + if (secret.type !=3D CRYPTO_KPP_SECRET_TYPE_ECDH) + return -EINVAL; + if (len < secret.len) + return -EINVAL; + + memcpy(&key_size, ptr, sizeof(key_size)); + ptr +=3D sizeof(key_size); + + if (key_size =3D=3D 0) { + /* + * key_size =3D=3D 0: generate a validated random private key. + * Uses the kernel ECC library (FIPS 186-5 A.2.2) to ensure + * the scalar is in the valid range [2, n-3] for the curve. + */ + u64 priv[ECC_MAX_DIGITS]; + unsigned int ndigits =3D ctx->clen / sizeof(u64); + unsigned int curve_id; + u8 *rnd; + + if (secret.len !=3D ECDH_KPP_SECRET_MIN_SIZE) + return -EINVAL; + if (ndigits > ECC_MAX_DIGITS) + return -EINVAL; + /* Reject non-limb-aligned clen to prevent ndigits truncation */ + if (ctx->clen % sizeof(u64)) + return -EINVAL; + + if (ctx->curve =3D=3D PKE_CURVE_P256) + curve_id =3D ECC_CURVE_NIST_P256; + else if (ctx->curve =3D=3D PKE_CURVE_P384) + curve_id =3D ECC_CURVE_NIST_P384; + else + return -EINVAL; + + ret =3D ecc_gen_privkey(curve_id, ndigits, priv); + if (ret) { + memzero_explicit(priv, sizeof(priv)); + return ret; + } + + rnd =3D kmalloc(ctx->clen, GFP_KERNEL); + if (!rnd) { + memzero_explicit(priv, sizeof(priv)); + return -ENOMEM; + } + + /* Convert VLI (native LE-digit-order) to big-endian bytes */ + ecc_swap_digits(priv, (u64 *)rnd, ndigits); + memzero_explicit(priv, sizeof(priv)); + + ret =3D cmh_ecdh_commit_key(ctx, rnd, ctx->clen); + kfree_sensitive(rnd); + return ret; + } + + if (key_size !=3D ctx->clen) + return -EINVAL; + + if (secret.len !=3D ECDH_KPP_SECRET_MIN_SIZE + key_size) + return -EINVAL; + + /* + * Validate the raw NIST scalar is in range for the curve, matching the + * key_size=3D=3D0 path (which uses ecc_gen_privkey). Rejects degenerate= / + * out-of-range keys the HW might otherwise accept. + */ + { + u64 priv[ECC_MAX_DIGITS]; + unsigned int ndigits =3D ctx->clen / sizeof(u64); + unsigned int curve_id; + + if (ctx->clen % sizeof(u64) || ndigits > ECC_MAX_DIGITS) + return -EINVAL; + if (ctx->curve =3D=3D PKE_CURVE_P256) + curve_id =3D ECC_CURVE_NIST_P256; + else if (ctx->curve =3D=3D PKE_CURVE_P384) + curve_id =3D ECC_CURVE_NIST_P384; + else + return -EINVAL; + + ecc_digits_from_bytes(ptr, key_size, priv, ndigits); + ret =3D ecc_is_key_valid(curve_id, ndigits, priv, key_size); + memzero_explicit(priv, sizeof(priv)); + if (ret) + return ret; + } + + return cmh_ecdh_commit_key(ctx, ptr, key_size); +} + +static int cmh_ecdh_set_secret_x25519(struct crypto_kpp *tfm, + const void *buf, unsigned int len) +{ + struct cmh_ecdh_tfm_ctx *ctx =3D cmh_ecdh_ctx(tfm); + + if (len !=3D pke_curve_clen(PKE_CURVE_25519)) + return -EINVAL; + + return cmh_ecdh_commit_key(ctx, buf, len); +} + +static void cmh_ecdh_keygen_complete(void *data, int error) +{ + struct kpp_request *req =3D data; + struct cmh_ecdh_reqctx *rctx =3D kpp_request_ctx(req); + + if (error =3D=3D -EINPROGRESS) { + cmh_complete(&req->base, error); + return; + } + + if (!cmh_dma_map_error(rctx->pk_dma)) + cmh_dma_unmap_single(rctx->pk_dma, rctx->out_len, + DMA_FROM_DEVICE); + + if (!error) { + int nents; + + nents =3D sg_nents_for_len(req->dst, rctx->out_len); + if (nents < 0 || + sg_copy_from_buffer(req->dst, nents, + rctx->pk_buf, + rctx->out_len) !=3D rctx->out_len) + error =3D -EINVAL; + else + req->dst_len =3D rctx->out_len; + } + + kfree(rctx->pk_buf); + rctx->pk_buf =3D NULL; + cmh_complete(&req->base, error); +} + +/* + * generate_public_key: For NIST ECDH, use ECDH_KEYGEN which outputs + * the public key X-coordinate. But the kernel kpp interface expects + * uncompressed X||Y, so we use ECDSA_PUBGEN which gives us (X,Y). + * For Curve25519, ECDH_KEYGEN gives us the Montgomery u-coordinate + * which is the full public key. + */ +static int cmh_ecdh_generate_public_key(struct kpp_request *req) +{ + struct crypto_kpp *tfm =3D crypto_kpp_reqtfm(req); + struct cmh_ecdh_tfm_ctx *ctx =3D cmh_ecdh_ctx(tfm); + struct cmh_ecdh_reqctx *rctx =3D kpp_request_ctx(req); + u32 clen =3D ctx->clen; + bool is_25519 =3D (ctx->curve =3D=3D PKE_CURVE_25519); + u32 out_len =3D is_25519 ? clen : 2 * clen; + struct vcq_cmd vcq[PKE_VCQ_CMDS_MAX]; + struct core_dispatch dd; + u32 swap, dma_swap; + int ret, idx; + gfp_t gfp; + + if (ctx->key.mode !=3D CMH_KEY_RAW) + return -EINVAL; + if (req->dst_len < out_len) + return -EINVAL; + + gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + + memset(rctx, 0, sizeof(*rctx)); + rctx->out_len =3D out_len; + rctx->pk_dma =3D DMA_MAPPING_ERROR; + + rctx->pk_buf =3D kzalloc(out_len, gfp); + if (!rctx->pk_buf) + return -ENOMEM; + + rctx->pk_dma =3D cmh_dma_map_single(rctx->pk_buf, out_len, + DMA_FROM_DEVICE); + if (cmh_dma_map_error(rctx->pk_dma)) { + ret =3D -ENOMEM; + goto out_free; + } + + /* + * The SYS_REF_TEMP key write uses per-curve dma_swap (0 for the + * little-endian Curve25519 secret), but the PKE compute command + * itself takes PKE_SWAP_FLAGS for every curve: the HW emits the + * public-key / shared-secret coordinate in the order PKE_SWAP_FLAGS + * maps to the expected encoding (the mgmt ECDH keygen does the same + * and matches RFC 7748). Do NOT switch the command to dma_swap. + */ + swap =3D PKE_SWAP_FLAGS; + dma_swap =3D pke_swap_flags(ctx->curve); + + dd =3D cmh_core_select_instance(CMH_CORE_PKE); + + vcq_set_header(&vcq[0], PKE_VCQ_CMDS_MAX); + idx =3D 1; + vcq_add_sys_write(&vcq[idx], SYS_REF_TEMP, ctx->key.raw.dma, + SYS_REF_NONE, ctx->key.raw.len, + ctx->key.raw.sys_type); + vcq[idx].id |=3D dma_swap; + idx++; + if (is_25519) + vcq_add_pke_ecdh_keygen(&vcq[idx++], dd.core_id, ctx->curve, + clen, rctx->pk_dma, SYS_REF_TEMP, + swap); + else + vcq_add_pke_ecdsa_pubgen(&vcq[idx++], dd.core_id, + ctx->curve, clen, rctx->pk_dma, + SYS_REF_TEMP, swap); + vcq_add_pke_flush(&vcq[idx++], dd.core_id); + + ret =3D cmh_tm_submit_async(vcq, PKE_VCQ_CMDS_MAX, 1, dd.mbx_idx, + cmh_ecdh_keygen_complete, req, + !!(req->base.flags & + CRYPTO_TFM_REQ_MAY_BACKLOG), 0); + if (ret =3D=3D -EBUSY) + return -EBUSY; + if (!ret) + return -EINPROGRESS; + + if (!cmh_dma_map_error(rctx->pk_dma)) + cmh_dma_unmap_single(rctx->pk_dma, out_len, + DMA_FROM_DEVICE); + +out_free: + kfree(rctx->pk_buf); + return ret; +} + +static void cmh_ecdh_ss_complete(void *data, int error) +{ + struct kpp_request *req =3D data; + struct cmh_ecdh_reqctx *rctx =3D kpp_request_ctx(req); + + if (error =3D=3D -EINPROGRESS) { + cmh_complete(&req->base, error); + return; + } + + if (!cmh_dma_map_error(rctx->peer_dma)) + cmh_dma_unmap_single(rctx->peer_dma, rctx->peer_len, + DMA_TO_DEVICE); + if (!cmh_dma_map_error(rctx->ss_dma)) + cmh_dma_unmap_single(rctx->ss_dma, rctx->clen, + DMA_FROM_DEVICE); + + if (!error) { + int nents; + + nents =3D sg_nents_for_len(req->dst, rctx->clen); + if (nents < 0 || + sg_copy_from_buffer(req->dst, nents, + rctx->ss_buf, + rctx->clen) !=3D rctx->clen) + error =3D -EINVAL; + else + req->dst_len =3D rctx->clen; + } + + kfree(rctx->peer_buf); + rctx->peer_buf =3D NULL; + kfree_sensitive(rctx->ss_buf); + rctx->ss_buf =3D NULL; + cmh_complete(&req->base, error); +} + +/* + * compute_shared_secret: PKE_CMD_ECDH. + * + * req->src =3D peer public key (X||Y for NIST, raw 32B for Curve25519). + * Output =3D shared secret X coordinate (clen bytes). + * + * The CMH ECDH command stores the shared secret in a datastore object. + * We write the private key into the per-mailbox scratch object, run the + * ECDH with the shared secret targeting SYS_REF_TEMP, and read it back + * with SYS_CMD_DATA in the same VCQ -- the eSW reclaims the temp store on + * read, so no datastore slot leaks per request. + */ +static int cmh_ecdh_compute_shared_secret(struct kpp_request *req) +{ + struct crypto_kpp *tfm =3D crypto_kpp_reqtfm(req); + struct cmh_ecdh_tfm_ctx *ctx =3D cmh_ecdh_ctx(tfm); + struct cmh_ecdh_reqctx *rctx =3D kpp_request_ctx(req); + u32 clen =3D ctx->clen; + bool is_25519 =3D (ctx->curve =3D=3D PKE_CURVE_25519); + u32 peer_len =3D is_25519 ? clen : 2 * clen; + u32 ss_type =3D SYS_TYPE_SET(SYS_TYPE_FLAG_PT, CORE_ID_PKE); + struct vcq_cmd vcq[6]; + struct core_dispatch dd; + u64 key_ref; + u32 swap, dma_swap; + int ret, idx, nents; + gfp_t gfp; + + if (ctx->key.mode !=3D CMH_KEY_RAW) + return -EINVAL; + if (req->src_len < peer_len || req->dst_len < clen) + return -EINVAL; + + gfp =3D req->base.flags & CRYPTO_TFM_REQ_MAY_SLEEP ? + GFP_KERNEL : GFP_ATOMIC; + + memset(rctx, 0, sizeof(*rctx)); + rctx->clen =3D clen; + rctx->peer_len =3D peer_len; + rctx->pk_dma =3D DMA_MAPPING_ERROR; + rctx->peer_dma =3D DMA_MAPPING_ERROR; + rctx->ss_dma =3D DMA_MAPPING_ERROR; + + rctx->peer_buf =3D kmalloc(peer_len, gfp); + rctx->ss_buf =3D kzalloc(clen, gfp); + if (!rctx->peer_buf || !rctx->ss_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + nents =3D sg_nents_for_len(req->src, peer_len); + if (nents < 0 || + sg_pcopy_to_buffer(req->src, nents, rctx->peer_buf, + peer_len, 0) !=3D peer_len) { + ret =3D -EINVAL; + goto out_free; + } + + rctx->peer_dma =3D cmh_dma_map_single(rctx->peer_buf, peer_len, + DMA_TO_DEVICE); + rctx->ss_dma =3D cmh_dma_map_single(rctx->ss_buf, clen, + DMA_FROM_DEVICE); + + if (cmh_dma_map_error(rctx->peer_dma) || + cmh_dma_map_error(rctx->ss_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + dd =3D cmh_core_select_instance(CMH_CORE_PKE); + if (dd.mbx_idx < 0 || (u32)dd.mbx_idx >=3D cmh_ecdh_scratch_count || + !cmh_ecdh_key_scratch[dd.mbx_idx]) { + ret =3D -EIO; + goto out_unmap; + } + key_ref =3D cmh_ecdh_key_scratch[dd.mbx_idx]; + + /* + * The private-key write uses per-curve dma_swap (0 for the + * little-endian Curve25519 secret), but the PKE compute command + * itself takes PKE_SWAP_FLAGS for every curve: the HW emits the + * shared-secret coordinate in the order PKE_SWAP_FLAGS maps to the + * expected encoding (the mgmt ECDH path does the same and matches + * RFC 7748). Do NOT switch the command to dma_swap. + */ + swap =3D PKE_SWAP_FLAGS; + dma_swap =3D pke_swap_flags(ctx->curve); + + vcq_set_header(&vcq[0], 6); + idx =3D 1; + vcq_add_sys_write(&vcq[idx], key_ref, ctx->key.raw.dma, + SYS_REF_NONE, ctx->key.raw.len, + ctx->key.raw.sys_type); + vcq[idx].id |=3D dma_swap; + idx++; + vcq_add_pke_ecdh(&vcq[idx++], dd.core_id, ctx->curve, clen, + clen, ss_type, rctx->peer_dma, + key_ref, SYS_REF_TEMP, swap); + vcq_add_pke_flush(&vcq[idx++], dd.core_id); + vcq_add_sys_data(&vcq[idx], SYS_REF_TEMP, rctx->ss_dma, clen); + vcq[idx].id |=3D dma_swap; + idx++; + vcq_add_sys_flush(&vcq[idx++]); + + ret =3D cmh_tm_submit_async(vcq, 6, 1, dd.mbx_idx, + cmh_ecdh_ss_complete, req, + !!(req->base.flags & + CRYPTO_TFM_REQ_MAY_BACKLOG), 0); + if (ret =3D=3D -EBUSY) + return -EBUSY; + if (!ret) + return -EINPROGRESS; + +out_unmap: + if (!cmh_dma_map_error(rctx->ss_dma)) + cmh_dma_unmap_single(rctx->ss_dma, clen, + DMA_FROM_DEVICE); + if (!cmh_dma_map_error(rctx->peer_dma)) + cmh_dma_unmap_single(rctx->peer_dma, peer_len, + DMA_TO_DEVICE); + +out_free: + kfree_sensitive(rctx->ss_buf); + kfree(rctx->peer_buf); + return ret; +} + +static unsigned int cmh_ecdh_max_size(struct crypto_kpp *tfm) +{ + struct cmh_ecdh_tfm_ctx *ctx =3D cmh_ecdh_ctx(tfm); + + /* Max output =3D X||Y for generate_public_key (NIST) */ + return 2 * ctx->clen; +} + +static unsigned int cmh_x25519_max_size(struct crypto_kpp *tfm) +{ + return pke_curve_clen(PKE_CURVE_25519); /* single coordinate */ +} + +static int cmh_ecdh_p256_init(struct crypto_kpp *tfm) +{ + struct cmh_ecdh_tfm_ctx *ctx =3D cmh_ecdh_ctx(tfm); + + memset(ctx, 0, sizeof(*ctx)); + ctx->curve =3D PKE_CURVE_P256; + ctx->clen =3D pke_curve_clen(PKE_CURVE_P256); + tfm->reqsize =3D sizeof(struct cmh_ecdh_reqctx); + return 0; +} + +static int cmh_ecdh_p384_init(struct crypto_kpp *tfm) +{ + struct cmh_ecdh_tfm_ctx *ctx =3D cmh_ecdh_ctx(tfm); + + memset(ctx, 0, sizeof(*ctx)); + ctx->curve =3D PKE_CURVE_P384; + ctx->clen =3D pke_curve_clen(PKE_CURVE_P384); + tfm->reqsize =3D sizeof(struct cmh_ecdh_reqctx); + return 0; +} + +static int cmh_x25519_init(struct crypto_kpp *tfm) +{ + struct cmh_ecdh_tfm_ctx *ctx =3D cmh_ecdh_ctx(tfm); + + memset(ctx, 0, sizeof(*ctx)); + ctx->curve =3D PKE_CURVE_25519; + ctx->clen =3D pke_curve_clen(PKE_CURVE_25519); + tfm->reqsize =3D sizeof(struct cmh_ecdh_reqctx); + return 0; +} + +static void cmh_ecdh_exit(struct crypto_kpp *tfm) +{ + struct cmh_ecdh_tfm_ctx *ctx =3D cmh_ecdh_ctx(tfm); + + cmh_key_destroy(&ctx->key); +} + +static struct kpp_alg cmh_ecdh_algs[] =3D { + { + .set_secret =3D cmh_ecdh_set_secret_nist, + .generate_public_key =3D cmh_ecdh_generate_public_key, + .compute_shared_secret =3D cmh_ecdh_compute_shared_secret, + .max_size =3D cmh_ecdh_max_size, + .init =3D cmh_ecdh_p256_init, + .exit =3D cmh_ecdh_exit, + .base =3D { + .cra_name =3D "ecdh-nist-p256", + .cra_driver_name =3D "rambus-cmh-ecdh-nist-p256", + .cra_priority =3D 300, + .cra_flags =3D CRYPTO_ALG_ASYNC, + .cra_module =3D THIS_MODULE, + .cra_ctxsize =3D sizeof(struct cmh_ecdh_tfm_ctx), + }, + }, + { + .set_secret =3D cmh_ecdh_set_secret_nist, + .generate_public_key =3D cmh_ecdh_generate_public_key, + .compute_shared_secret =3D cmh_ecdh_compute_shared_secret, + .max_size =3D cmh_ecdh_max_size, + .init =3D cmh_ecdh_p384_init, + .exit =3D cmh_ecdh_exit, + .base =3D { + .cra_name =3D "ecdh-nist-p384", + .cra_driver_name =3D "rambus-cmh-ecdh-nist-p384", + .cra_priority =3D 300, + .cra_flags =3D CRYPTO_ALG_ASYNC, + .cra_module =3D THIS_MODULE, + .cra_ctxsize =3D sizeof(struct cmh_ecdh_tfm_ctx), + }, + }, + { + .set_secret =3D cmh_ecdh_set_secret_x25519, + .generate_public_key =3D cmh_ecdh_generate_public_key, + .compute_shared_secret =3D cmh_ecdh_compute_shared_secret, + .max_size =3D cmh_x25519_max_size, + .init =3D cmh_x25519_init, + .exit =3D cmh_ecdh_exit, + .base =3D { + .cra_name =3D "curve25519", + .cra_driver_name =3D "rambus-cmh-curve25519", + .cra_priority =3D 300, + .cra_flags =3D CRYPTO_ALG_ASYNC, + .cra_module =3D THIS_MODULE, + .cra_ctxsize =3D sizeof(struct cmh_ecdh_tfm_ctx), + }, + }, +}; + +/* + * cmh_ecdh_scratch_free() - Scrub and release the per-mailbox key scratch + * + * Best-effort: sys_grant with no access wipes the key bytes and clears the + * CID on each object (the datastore stack space is only reclaimed by a fu= ll + * reset). The grant runs on the owning mailbox so the eSW access check + * passes. + */ +static void cmh_ecdh_scratch_free(void) +{ + u32 i; + + if (!cmh_ecdh_key_scratch) + return; + + for (i =3D 0; i < cmh_ecdh_scratch_count; i++) { + struct vcq_cmd vcq[3]; + + if (!cmh_ecdh_key_scratch[i]) + continue; + + vcq_set_header(&vcq[0], 3); + vcq_add_sys_grant(&vcq[1], cmh_ecdh_key_scratch[i], 0, 0, 0); + vcq_add_sys_flush(&vcq[2]); + cmh_tm_submit_sync_mbx(vcq, 3, 1, (s32)i); + cmh_ecdh_key_scratch[i] =3D 0; + } + + kfree(cmh_ecdh_key_scratch); + cmh_ecdh_key_scratch =3D NULL; + cmh_ecdh_scratch_count =3D 0; + cmh_ecdh_scratch_ready =3D false; +} + +/* + * cmh_ecdh_scratch_ensure() - Create the per-PKE-mailbox key scratch on t= he + * first keyed use. + * + * Deferred from register time so that a bring-up which never sets an ECDH + * key leaves the datastore empty (a DS import requires an empty store). + * Idempotent and safe against concurrent set_secret callers; on failure it + * leaves cmh_ecdh_scratch_ready clear and any objects already created in + * place, so the next key set retries and skips them. Runs in process + * context (set_secret), so the synchronous submit may sleep. + */ +static int cmh_ecdh_scratch_ensure(void) +{ + u32 key_len =3D pke_curve_clen(PKE_CURVE_P384); + u32 n_pke =3D cmh_core_num_instances(CMH_CORE_PKE); + u64 *ref_buf; + u32 i; + int ret =3D 0; + + mutex_lock(&cmh_ecdh_scratch_lock); + if (cmh_ecdh_scratch_ready) + goto out; + + ref_buf =3D kzalloc_obj(u64, GFP_KERNEL); + if (!ref_buf) { + ret =3D -ENOMEM; + goto out; + } + + for (i =3D 0; i < n_pke; i++) { + struct core_dispatch d =3D cmh_core_get_instance(CMH_CORE_PKE, i); + struct vcq_cmd vcq[3]; + dma_addr_t ref_dma; + + if (d.mbx_idx < 0 || (u32)d.mbx_idx >=3D cmh_ecdh_scratch_count) { + ret =3D -EINVAL; + goto out_free; + } + if (cmh_ecdh_key_scratch[d.mbx_idx]) + continue; + + *ref_buf =3D 0; + ref_dma =3D cmh_dma_map_single(ref_buf, sizeof(*ref_buf), + DMA_FROM_DEVICE); + if (cmh_dma_map_error(ref_dma)) { + ret =3D -ENOMEM; + goto out_free; + } + + vcq_set_header(&vcq[0], 3); + vcq_add_sys_new(&vcq[1], 0, ref_dma, key_len); + vcq_add_sys_flush(&vcq[2]); + ret =3D cmh_tm_submit_sync_mbx(vcq, 3, 1, d.mbx_idx); + if (!ret) { + cmh_dma_sync_for_cpu(ref_dma, sizeof(*ref_buf), + DMA_FROM_DEVICE); + cmh_ecdh_key_scratch[d.mbx_idx] =3D *ref_buf; + } + cmh_dma_unmap_single(ref_dma, sizeof(*ref_buf), + DMA_FROM_DEVICE); + if (ret) + goto out_free; + if (!cmh_ecdh_key_scratch[d.mbx_idx]) { + ret =3D -EIO; + goto out_free; + } + } + + cmh_ecdh_scratch_ready =3D true; +out_free: + kfree(ref_buf); +out: + mutex_unlock(&cmh_ecdh_scratch_lock); + return ret; +} + +/** + * cmh_pke_ecdh_register() - Register ECDH kpp algorithms with the crypto = framework + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_pke_ecdh_register(void) +{ + u32 n_mbx =3D cmh_tm_mbx_count(); + int ret, i; + + if (!cmh_core_present(CMH_CORE_PKE)) + return 0; + + /* + * Host-side index array only; the per-mailbox datastore objects are + * created lazily on the first key set (cmh_ecdh_scratch_ensure) so a + * bring-up that never keys ECDH leaves the datastore empty for a DS + * import. + */ + cmh_ecdh_key_scratch =3D kcalloc(n_mbx, sizeof(*cmh_ecdh_key_scratch), + GFP_KERNEL); + if (!cmh_ecdh_key_scratch) + return -ENOMEM; + cmh_ecdh_scratch_count =3D n_mbx; + + for (i =3D 0; i < ARRAY_SIZE(cmh_ecdh_algs); i++) { + ret =3D crypto_register_kpp(&cmh_ecdh_algs[i]); + if (ret) { + dev_err(cmh_dev(), "cmh: failed to register %s (%d)\n", + cmh_ecdh_algs[i].base.cra_name, ret); + goto err_unregister; + } + } + + return 0; + +err_unregister: + while (i--) + crypto_unregister_kpp(&cmh_ecdh_algs[i]); + cmh_ecdh_scratch_free(); + return ret; +} + +/** + * cmh_pke_ecdh_unregister() - Unregister ECDH kpp algorithms from the cry= pto framework + */ +void cmh_pke_ecdh_unregister(void) +{ + int i =3D ARRAY_SIZE(cmh_ecdh_algs); + + if (!cmh_core_present(CMH_CORE_PKE)) + return; + + while (i--) + crypto_unregister_kpp(&cmh_ecdh_algs[i]); + + /* + * Safe to free the shared scratch now: crypto_unregister_kpp() above + * blocks until every TFM is released and no request is in flight, so + * no cmh_ecdh_compute_shared_secret() can still be reading it. + */ + cmh_ecdh_scratch_free(); +} --=20 2.43.7 From nobody Fri Sep 25 01:20:34 2026 Received: from DM1PR04CU001.outbound.protection.outlook.com (mail-centralusazon11020136.outbound.protection.outlook.com [52.101.61.136]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BBDAE47D927; Thu, 17 Sep 2026 23:00:32 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.61.136 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789686036; cv=fail; b=j9owCvXmKmQoSh2QpWpPiQHrwyOwMdu6AGoSwE4vQvVuWWdkQDRLX0NoKBkTE9Rqo1g+n0j1uY/y3guglHWDs1xRS5U1HdCQCiOuCXADP2aJWi5Z929RvLixyPOkL8bawXcmuEATBEFE5Lt5Gbd5mw/ok9yDa91ICVlLZL659Lo= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789686036; c=relaxed/simple; bh=0Cnc1DkAptjYYIsYGfbGXqUQxiPms2HEG2hq+5FiZ68=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=lFunwSKx+CsZ6zmqNk3tMAwnro+Dyijo6GFLY2RtJgY5T1hvNVZKeRBSlkZS6C4hv3PlMrEjvbfeE7n+gYdKHXQJq/Dp3G1N5O90QsonJv8bW4lxCc5yCpc296EIatNhVo+Gt7LRaHQTZwO9lu2pOnW1u0Wj7lPbL+5nKbJ3TGI= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com; spf=fail smtp.mailfrom=rambus.com; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b=xHiSU6fq; arc=fail smtp.client-ip=52.101.61.136 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=rambus.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b="xHiSU6fq" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=kZD3BPsS7Q1LBrwLzcgYaPlrAJLzwMHfSc6MWCZgAhCbC/dGeXQNCJpxjA10XZBrAQOaBX7jodbM95TGg1OaP6Oq71H97FaRvdRfJkbcW7sNxlnDRs38hADgMBXxCTdhflqgiYapjG/ITKRZFRuGWvQz71zaVlIcG66FMbNCoflJfD9k9O/WIoUKi1g90w0IbfwxiNeJw5X1IaVVooZvMstx0gqCFywjamwrt0C+FaahYG1bE0sRoDnmgZQF0hSUc9IkZ1MiMH/ypUVVIcLxlWrWiGhYWVD90lv+jhkJYNMOrgBZkFR/R8Q+X4p9QccxU4hnsAIgqboY8U/LiN5SPw== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=23jVenue7GBFJK9hg5/EylL7K6WSTFd308Je/a2McuU=; b=HOfx9IQsY/M1XqXKscKs/kTUKxBCIckmGncqoC2aMZpfnUY4RafHpaAH45YC7QKa0RBnB5FcUT59huI9oivv/hFlkInlIwLrlv1Ykj75spahQvwMlSjNYx189ZJPb2YKFEhvzBL6K4ntPIAcDgI16KCeQZoTXw3wra2HhjffLTVn8jGMXMcCXPyxbrGon4RjIAQo6D7oz4P4Imtyh6uvpWGcLslqMq1MJ/H1jWjdAbvgye7LRWBR2iSIxocLKu8ryj9LsUPAl5WME4CeVkFoFci3KCEcz7qkDwMKc+n7gr7jQVJP2oaazpIbn98vxDwUwpXWge84DcNAqpLL3MAf5A== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 192.86.86.210) smtp.rcpttodomain=cryptography.com smtp.mailfrom=rambus.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=rambus.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=rambus.com; s=selector1; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=23jVenue7GBFJK9hg5/EylL7K6WSTFd308Je/a2McuU=; b=xHiSU6fqo+pdZ8h1i8XwrSE+osUBWmEQkgDADfQqOKt00lynpRV5pKflHNY0uEwH934vd8qv3j+R6tVwfz/Ry4EGzCcUZDQgcC8UxVSA91h9s6nSmqD0UU3dATKSRNloASXcPH1p8EnXr7OOSvd6fjDgVsHeD+6M0pk32n1DDC/9zQ4d1vkjhJso8KC+E54YzFHbtazWZ1liqK0wXQxk2uMbfThq9ztjE79SBmy2MKkAlVBRXA/nrABxTpMfO/FrerU0c4SAzjr0TYKAAKEMZYZUFF0zcFbB2p15ilePH6Qy6vRUroFg278bMRoIU8cHj12WVRUR8LhshZaKkjiz5Q== Received: from SJ0PR13CA0226.namprd13.prod.outlook.com (2603:10b6:a03:2c1::21) by SN4PR04MB10229.namprd04.prod.outlook.com (2603:10b6:806:216::8) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.406.12; Thu, 17 Sep 2026 22:59:35 +0000 Received: from BY1PEPF0001AE1A.namprd04.prod.outlook.com (2603:10b6:a03:2c1:cafe::71) by SJ0PR13CA0226.outlook.office365.com (2603:10b6:a03:2c1::21) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.451.8 via Frontend Transport; Thu, 17 Sep 2026 22:59:34 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 192.86.86.210) smtp.mailfrom=rambus.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=rambus.com; Received-SPF: Pass (protection.outlook.com: domain of rambus.com designates 192.86.86.210 as permitted sender) receiver=protection.outlook.com; client-ip=192.86.86.210; helo=hqxsv-psmtppxy02.rambus.com; pr=C Received: from hqxsv-psmtppxy02.rambus.com (192.86.86.210) by BY1PEPF0001AE1A.mail.protection.outlook.com (10.167.242.102) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.451.8 via Frontend Transport; Thu, 17 Sep 2026 22:59:34 +0000 Received: from hqxsv-cmdev3-aousherovitch.rambus.com (hqn-lb-int-float.rambus.com [10.12.20.20]) by hqxsv-psmtppxy02.rambus.com (Postfix) with ESMTP id 7AA45180176D; Thu, 17 Sep 2026 22:59:32 +0000 (UTC) From: Alex Ousherovitch To: Albert Ou , Alex Ousherovitch , Conor Dooley , "David S. Miller" , Herbert Xu , Jonathan Corbet , Krzysztof Kozlowski , Palmer Dabbelt , Paul Walmsley , Rob Herring , Saravanakrishnan Krishnamoorthy , Shuah Khan Cc: Alexandre Ghiti , devicetree@vger.kernel.org, Joel Wittenauer , linux-api@vger.kernel.org, linux-crypto@vger.kernel.org, linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-riscv@lists.infradead.org, Shuah Khan , Thi Nguyen Subject: [PATCH v5 15/19] crypto: cmh - add ML-KEM/ML-DSA (QSE) Date: Thu, 17 Sep 2026 15:59:24 -0700 Message-ID: <20260917225929.2494111-16-aousherovitch@rambus.com> X-Mailer: git-send-email 2.43.7 In-Reply-To: <20260917225929.2494111-1-aousherovitch@rambus.com> References: <20260917225929.2494111-1-aousherovitch@rambus.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: BY1PEPF0001AE1A:EE_|SN4PR04MB10229:EE_ X-MS-Office365-Filtering-Correlation-Id: 07577089-6300-49a0-46ca-08df150f576e X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|7416014|1800799024|376014|23010399003|82310400026|36860700016|921020|6133799003|56012099006|11063799006|18002099003|22082099003|3023799007|10067099003; X-Microsoft-Antispam-Message-Info: f85clLlK7xRuhX2+cBV3b/u3t6Lfuuzk7nQVpB8dCevU863Fx2l3y284BbHOs/6TPIkWuM0BBCHfgXp/GHjKpp9Wq+RBOkG/H+pJvYmRrstjtgYVW8IZEqSpQSq/Coc8Qukv2YyST26dqXU3zNgAhy6gV8Fssvoy7nrExU7bO44mwB8KMNDJUdEyWPS5Pi50Awy2s+6eJ6Mw4n9+ntoGKrfUR5p4PSBiUwAOgeSQ/EXnyyojScfESaQ7hJly/O6fpGK7C7neuRoBvSJND1L1mjl9HpA3fVLp8wpAv4I1YF3fWtVu2nwSsOkGkDUSKtcAkwjzASacPRX4jEcmLbnrgLOqFmkVzV8UfQ2Ka6hkaz7eEFUasvCUgujXWUf/gJwvNnOJjKfhnyiFfqpEqM3VdX5sB6vOoWX70mEUPTMcp5/8JhZ96hLpRWO7csyhk4+340PpEhdV0wzORPUvWoYUsXiznV/G55X/0Yzas0MmPsXhRh+jKTjqv3GU7x2/qk3lF2fe1e0H8I7lqiYNiasMAMfD3DP4/XXI0VVtunk5RK0nP6qo4UY8edFxPbeEwvIN4k22wGo2De7d1xafhSv/mv6RleXijeC+Oxuho4HmRjjLuGOro6B76UuYkmP2/lZEWmgQB1FopH6+xgCDYjFHeGek8sV0OxUmZ1OjgOH4aa4ajHOfyIH4e0Nyj4d3lwyR4Cb6jVocc8YlYWoXFPfSgoz35wuFmNURjB29SB6MmBj6q1hWf3D5zIU3ynY+6tGs X-Forefront-Antispam-Report: CIP:192.86.86.210;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:hqxsv-psmtppxy02.rambus.com;PTR:InfoDomainNonexistent;CAT:NONE;SFS:(13230040)(7416014)(1800799024)(376014)(23010399003)(82310400026)(36860700016)(921020)(6133799003)(56012099006)(11063799006)(18002099003)(22082099003)(3023799007)(10067099003);DIR:OUT;SFP:1102; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: S7UoPM129yUKGNoI9gUFCJuKNk5YiymE8UDjpy1ec1ZhqSkSfZSGQ4UrHxWU5+kY+XFgOCtUC774UXI9io9/22BH2f28nzm+THWsxT9TFkLKVL1NZFqpf0cLiJvOx9ApP2LSAptdqR//3UZe7VRQ2B6+CahEI/zv8VevcHhqgrT+b62AUpoxpsYlFzSzOrDgBsC2W6N6j6/kRRnaTE1wVAEcvO6nOg/7v2XdRKfpFlMw3DX3uAX4Q12nGez8YMl8bSd2eNL8X2qdxe8Acid+VfNlLP5ZTkOEKArgNaD0olM/8+r6bKIMNjPd7dWoFMjAk7WlDDA0uiO64Blj/Y0d+ixQkDqf33FUtu3YwwbP94J3M6TsGq0uSZgd8N22+YPoXEZLLoVrXWmox2S5ICO2lOlPEwjgqGdUJCN8pnfgZuj8Dd0xZo6FE1K2OHRkaNEO X-OriginatorOrg: rambus.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 17 Sep 2026 22:59:34.4984 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: 07577089-6300-49a0-46ca-08df150f576e X-MS-Exchange-CrossTenant-Id: bd0ba799-c2b9-413c-9c56-5d1731c4827c X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=bd0ba799-c2b9-413c-9c56-5d1731c4827c;Ip=[192.86.86.210];Helo=[hqxsv-psmtppxy02.rambus.com] X-MS-Exchange-CrossTenant-AuthSource: BY1PEPF0001AE1A.namprd04.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: SN4PR04MB10229 Content-Type: text/plain; charset="utf-8" Register ML-KEM (Kyber) and ML-DSA (Dilithium) algorithms using the CMH QSE core (core ID 0x09). ML-KEM is ioctl-only (keygen, encaps, decaps). ML-DSA is registered as a sig algorithm with priority 5001 to override the kernel's verify-only mldsa implementation at priority 5000. This follows the established pattern where hardware drivers override software-only fallbacks (e.g. ccp at 300 over generic AES at 100, qat similarly). The CMH driver provides full HW-accelerated sign + verify vs the kernel's verify-only software implementation. Includes cmh_pqc_sizes.c with compile-time tables of PQC key and signature sizes for all supported parameter sets. Signed-off-by: Alex Ousherovitch Co-developed-by: Saravanakrishnan Krishnamoorthy Signed-off-by: Saravanakrishnan Krishnamoorthy --- drivers/crypto/cmh/Makefile | 5 +- drivers/crypto/cmh/cmh_main.c | 9 + drivers/crypto/cmh/cmh_pqc_mldsa.c | 415 +++++++++++++++++++++++++++++ drivers/crypto/cmh/cmh_pqc_sizes.c | 39 +++ drivers/crypto/cmh/cmh_qse.c | 211 +++++++++++++++ 5 files changed, 678 insertions(+), 1 deletion(-) create mode 100644 drivers/crypto/cmh/cmh_pqc_mldsa.c create mode 100644 drivers/crypto/cmh/cmh_pqc_sizes.c create mode 100644 drivers/crypto/cmh/cmh_qse.c diff --git a/drivers/crypto/cmh/Makefile b/drivers/crypto/cmh/Makefile index da420d0ec1fe..7196d992487e 100644 --- a/drivers/crypto/cmh/Makefile +++ b/drivers/crypto/cmh/Makefile @@ -33,7 +33,10 @@ cmh-y :=3D \ cmh_pke_common.o \ cmh_pke_rsa.o \ cmh_pke_ecdsa.o \ - cmh_pke_ecdh.o + cmh_pke_ecdh.o \ + cmh_qse.o \ + cmh_pqc_mldsa.o \ + cmh_pqc_sizes.o =20 # Management ioctl device (/dev/cmh_mgmt): key lifecycle, PKE, PQC ioctls. cmh-$(CONFIG_CRYPTO_DEV_CMH_MGMT) +=3D \ diff --git a/drivers/crypto/cmh/cmh_main.c b/drivers/crypto/cmh/cmh_main.c index 93d939e8fc4e..7dfcdf395175 100644 --- a/drivers/crypto/cmh/cmh_main.c +++ b/drivers/crypto/cmh/cmh_main.c @@ -41,6 +41,7 @@ #include "cmh_sm4.h" #include "cmh_ccp.h" #include "cmh_pke.h" +#include "cmh_pqc.h" #include "cmh_mgmt.h" #include "cmh_registers.h" #include "cmh_debugfs.h" @@ -302,6 +303,11 @@ static int cmh_probe(struct platform_device *pdev) if (ret) goto err_pke_ecdh_register; =20 + /* Register PQC ML-KEM/ML-DSA */ + ret =3D cmh_pqc_mldsa_register(); + if (ret) + goto err_pqc_mldsa_register; + /* Register key management device (/dev/cmh_mgmt) */ ret =3D cmh_mgmt_register(); if (ret) @@ -312,6 +318,8 @@ static int cmh_probe(struct platform_device *pdev) return 0; =20 err_mgmt_register: + cmh_pqc_mldsa_unregister(); +err_pqc_mldsa_register: cmh_pke_ecdh_unregister(); err_pke_ecdh_register: cmh_pke_ecdsa_unregister(); @@ -374,6 +382,7 @@ static void cmh_remove(struct platform_device *pdev) cfg =3D &dev->config; =20 cmh_mgmt_unregister(); + cmh_pqc_mldsa_unregister(); cmh_pke_ecdh_unregister(); cmh_pke_ecdsa_unregister(); cmh_pke_rsa_unregister(); diff --git a/drivers/crypto/cmh/cmh_pqc_mldsa.c b/drivers/crypto/cmh/cmh_pq= c_mldsa.c new file mode 100644 index 000000000000..c37ba4bae89c --- /dev/null +++ b/drivers/crypto/cmh/cmh_pqc_mldsa.c @@ -0,0 +1,415 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- ML-DSA Signature Driver (sig_alg, synchronous) + * + * Registers "mldsa44", "mldsa65", "mldsa87" sig algorithms + * with sign, verify, set_pub_key, and set_priv_key callbacks. + * + * Key format: + * Public key =3D raw pk bytes (1312 / 1952 / 2592 bytes) + * Private key =3D raw sk bytes (2560 / 4032 / 4896 bytes) + * + * Sign: src =3D message bytes (up to 10240 bytes), dst =3D raw signature + * Verify: src =3D raw signature, digest =3D message bytes + * + * Non-masked mode only for sig_alg API. + * Masked mode available through /dev/cmh_mgmt ioctl. + */ + +#include +#include +#include +#include +#include + +#include "cmh_sys.h" +#include "cmh_qse_abi.h" +#include "cmh_txn.h" +#include "cmh_dma.h" +#include "cmh_key.h" +#include "cmh_pqc.h" + +struct cmh_mldsa_tfm_ctx { + struct cmh_key_ctx key; /* private key (raw only) */ + u8 *pub_key; + u32 pub_key_len; + u32 mode; /* ML_DSA_MODE_44/65/87 */ + int mode_idx; /* index into size tables */ +}; + +static inline struct cmh_mldsa_tfm_ctx *cmh_mldsa_ctx(struct crypto_sig *t= fm) +{ + return crypto_sig_ctx(tfm); +} + +/* + * ML-DSA sign (synchronous sig_alg) + * + * @src: message bytes + * @slen: message length + * @dst: signature output buffer + * @dlen: output buffer length + * + * Returns signature length on success, negative errno on failure. + */ +static int cmh_mldsa_sign(struct crypto_sig *tfm, + const void *src, unsigned int slen, + void *dst, unsigned int dlen) +{ + struct cmh_mldsa_tfm_ctx *ctx =3D cmh_mldsa_ctx(tfm); + int mi =3D ctx->mode_idx; + u32 sig_size =3D ml_dsa_sig_size[mi]; + u32 sk_size =3D ml_dsa_sk_size[mi]; + struct vcq_cmd vcq[QSE_VCQ_CMDS_MIN]; + struct core_dispatch dd; + u8 *m_buf =3D NULL, *sig_buf =3D NULL; + dma_addr_t m_dma =3D DMA_MAPPING_ERROR; + dma_addr_t sig_dma =3D DMA_MAPPING_ERROR; + int ret, idx; + + if (ctx->key.mode !=3D CMH_KEY_RAW) + return -EINVAL; + if (dlen < sig_size) + return -EINVAL; + if (slen > ML_DSA_MAX_MLEN) + return -EINVAL; + + sig_buf =3D kzalloc(sig_size, GFP_KERNEL); + if (!sig_buf) { + ret =3D -ENOMEM; + goto out_free; + } + /* ML-DSA permits an empty message (FIPS 204); skip 0-length DMA. */ + if (slen) { + m_buf =3D kmemdup(src, slen, GFP_KERNEL); + if (!m_buf) { + ret =3D -ENOMEM; + goto out_free; + } + } + + if (ctx->key.raw.len !=3D sk_size) { + ret =3D -EINVAL; + goto out_free; + } + + if (slen) { + m_dma =3D cmh_dma_map_single(m_buf, slen, DMA_TO_DEVICE); + if (cmh_dma_map_error(m_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + } + sig_dma =3D cmh_dma_map_single(sig_buf, sig_size, DMA_FROM_DEVICE); + + if (cmh_dma_map_error(sig_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + dd =3D cmh_core_select_instance(CMH_CORE_QSE); + + vcq_set_header(&vcq[0], QSE_VCQ_CMDS_MIN); + idx =3D 1; + vcq_add_qse_ml_dsa_sign(&vcq[idx++], dd.core_id, ctx->mode, + QSE_FLAG_USE_RNG, + 0, slen ? m_dma : 0, ctx->key.raw.dma, + sig_dma, slen, false); + vcq_add_qse_flush(&vcq[idx++], dd.core_id); + + ret =3D cmh_tm_submit_sync_mbx(vcq, QSE_VCQ_CMDS_MIN, 1, + dd.mbx_idx); + if (!ret) { + /* Sync bounce buffer so CPU sees the DMA-written signature */ + cmh_dma_sync_for_cpu(sig_dma, sig_size, DMA_FROM_DEVICE); + memcpy(dst, sig_buf, sig_size); + ret =3D sig_size; + } + +out_unmap: + if (!cmh_dma_map_error(sig_dma)) + cmh_dma_unmap_single(sig_dma, sig_size, DMA_FROM_DEVICE); + if (!cmh_dma_map_error(m_dma)) + cmh_dma_unmap_single(m_dma, slen, DMA_TO_DEVICE); + +out_free: + kfree(sig_buf); + kfree(m_buf); + return ret; +} + +/* + * ML-DSA verify (synchronous sig_alg) + * + * @src: raw signature + * @slen: signature length + * @digest: message bytes + * @dlen: message length + * + * Returns 0 on successful verification, negative errno on failure. + */ +static int cmh_mldsa_verify(struct crypto_sig *tfm, + const void *src, unsigned int slen, + const void *digest, unsigned int dlen) +{ + struct cmh_mldsa_tfm_ctx *ctx =3D cmh_mldsa_ctx(tfm); + int mi =3D ctx->mode_idx; + u32 sig_size =3D ml_dsa_sig_size[mi]; + u32 pk_size =3D ml_dsa_pk_size[mi]; + struct core_dispatch d =3D cmh_core_select_instance(CMH_CORE_QSE); + struct vcq_cmd vcq[QSE_VCQ_CMDS_MIN]; + u8 *sig_buf =3D NULL, *m_buf =3D NULL, *pk_buf =3D NULL; + dma_addr_t sig_dma =3D DMA_MAPPING_ERROR; + dma_addr_t m_dma =3D DMA_MAPPING_ERROR; + dma_addr_t pk_dma =3D DMA_MAPPING_ERROR; + int ret; + + if (!ctx->pub_key) + return -EINVAL; + if (slen !=3D sig_size) + return -EINVAL; + if (dlen > ML_DSA_MAX_MLEN) + return -EINVAL; + + sig_buf =3D kmemdup(src, slen, GFP_KERNEL); + pk_buf =3D kmemdup(ctx->pub_key, pk_size, GFP_KERNEL); + if (!sig_buf || !pk_buf) { + ret =3D -ENOMEM; + goto out_free; + } + /* ML-DSA permits an empty message (FIPS 204); skip 0-length DMA. */ + if (dlen) { + m_buf =3D kmemdup(digest, dlen, GFP_KERNEL); + if (!m_buf) { + ret =3D -ENOMEM; + goto out_free; + } + } + + sig_dma =3D cmh_dma_map_single(sig_buf, sig_size, DMA_TO_DEVICE); + pk_dma =3D cmh_dma_map_single(pk_buf, pk_size, DMA_TO_DEVICE); + if (dlen) { + m_dma =3D cmh_dma_map_single(m_buf, dlen, DMA_TO_DEVICE); + if (cmh_dma_map_error(m_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + } + + if (cmh_dma_map_error(sig_dma) || + cmh_dma_map_error(pk_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + vcq_set_header(&vcq[0], QSE_VCQ_CMDS_MIN); + vcq_add_qse_ml_dsa_verify(&vcq[1], d.core_id, ctx->mode, 0, + dlen ? m_dma : 0, pk_dma, sig_dma, dlen); + vcq_add_qse_flush(&vcq[2], d.core_id); + + ret =3D cmh_tm_submit_sync_mbx(vcq, QSE_VCQ_CMDS_MIN, 1, d.mbx_idx); + +out_unmap: + if (!cmh_dma_map_error(pk_dma)) + cmh_dma_unmap_single(pk_dma, pk_size, DMA_TO_DEVICE); + if (!cmh_dma_map_error(m_dma)) + cmh_dma_unmap_single(m_dma, dlen, DMA_TO_DEVICE); + if (!cmh_dma_map_error(sig_dma)) + cmh_dma_unmap_single(sig_dma, sig_size, DMA_TO_DEVICE); + +out_free: + kfree(pk_buf); + kfree(m_buf); + kfree(sig_buf); + return ret; +} + +static int cmh_mldsa_set_pub_key(struct crypto_sig *tfm, + const void *key, unsigned int keylen) +{ + struct cmh_mldsa_tfm_ctx *ctx =3D cmh_mldsa_ctx(tfm); + u32 expected =3D ml_dsa_pk_size[ctx->mode_idx]; + + if (keylen !=3D expected) + return -EINVAL; + + kfree(ctx->pub_key); + ctx->pub_key =3D NULL; + ctx->pub_key_len =3D 0; + + ctx->pub_key =3D kmemdup(key, keylen, GFP_KERNEL); + if (!ctx->pub_key) + return -ENOMEM; + + ctx->pub_key_len =3D keylen; + return 0; +} + +static int cmh_mldsa_set_priv_key(struct crypto_sig *tfm, + const void *key, unsigned int keylen) +{ + struct cmh_mldsa_tfm_ctx *ctx =3D cmh_mldsa_ctx(tfm); + u32 expected =3D ml_dsa_sk_size[ctx->mode_idx]; + + if (keylen !=3D expected) + return -EINVAL; + + return cmh_key_setkey_raw(&ctx->key, key, keylen, CORE_ID_QSE); +} + +static unsigned int cmh_mldsa_key_size(struct crypto_sig *tfm) +{ + struct cmh_mldsa_tfm_ctx *ctx =3D cmh_mldsa_ctx(tfm); + + /* crypto_sig_keysize() returns bits, not bytes */ + return ml_dsa_pk_size[ctx->mode_idx] * 8; +} + +static unsigned int cmh_mldsa_max_size(struct crypto_sig *tfm) +{ + struct cmh_mldsa_tfm_ctx *ctx =3D cmh_mldsa_ctx(tfm); + + return ml_dsa_sig_size[ctx->mode_idx]; +} + +static int cmh_mldsa_44_init(struct crypto_sig *tfm) +{ + struct cmh_mldsa_tfm_ctx *ctx =3D cmh_mldsa_ctx(tfm); + + memset(ctx, 0, sizeof(*ctx)); + ctx->mode =3D ML_DSA_MODE_44; + ctx->mode_idx =3D 0; + return 0; +} + +static int cmh_mldsa_65_init(struct crypto_sig *tfm) +{ + struct cmh_mldsa_tfm_ctx *ctx =3D cmh_mldsa_ctx(tfm); + + memset(ctx, 0, sizeof(*ctx)); + ctx->mode =3D ML_DSA_MODE_65; + ctx->mode_idx =3D 1; + return 0; +} + +static int cmh_mldsa_87_init(struct crypto_sig *tfm) +{ + struct cmh_mldsa_tfm_ctx *ctx =3D cmh_mldsa_ctx(tfm); + + memset(ctx, 0, sizeof(*ctx)); + ctx->mode =3D ML_DSA_MODE_87; + ctx->mode_idx =3D 2; + return 0; +} + +static void cmh_mldsa_exit(struct crypto_sig *tfm) +{ + struct cmh_mldsa_tfm_ctx *ctx =3D cmh_mldsa_ctx(tfm); + + cmh_key_destroy(&ctx->key); + kfree(ctx->pub_key); + ctx->pub_key =3D NULL; +} + +/* + * Priority 5001: the kernel's software ML-DSA (crypto/mldsa.c) registers + * at priority 5000 but only implements verify -- sign returns -EOPNOTSUPP. + * We provide full HW-accelerated sign + verify, so we must override. + */ +static struct sig_alg cmh_mldsa_algs[] =3D { + { + .sign =3D cmh_mldsa_sign, + .verify =3D cmh_mldsa_verify, + .set_pub_key =3D cmh_mldsa_set_pub_key, + .set_priv_key =3D cmh_mldsa_set_priv_key, + .key_size =3D cmh_mldsa_key_size, + .max_size =3D cmh_mldsa_max_size, + .init =3D cmh_mldsa_44_init, + .exit =3D cmh_mldsa_exit, + .base =3D { + .cra_name =3D "mldsa44", + .cra_driver_name =3D "rambus-cmh-mldsa44", + .cra_priority =3D 5001, + .cra_module =3D THIS_MODULE, + .cra_ctxsize =3D sizeof(struct cmh_mldsa_tfm_ctx), + }, + }, + { + .sign =3D cmh_mldsa_sign, + .verify =3D cmh_mldsa_verify, + .set_pub_key =3D cmh_mldsa_set_pub_key, + .set_priv_key =3D cmh_mldsa_set_priv_key, + .key_size =3D cmh_mldsa_key_size, + .max_size =3D cmh_mldsa_max_size, + .init =3D cmh_mldsa_65_init, + .exit =3D cmh_mldsa_exit, + .base =3D { + .cra_name =3D "mldsa65", + .cra_driver_name =3D "rambus-cmh-mldsa65", + .cra_priority =3D 5001, + .cra_module =3D THIS_MODULE, + .cra_ctxsize =3D sizeof(struct cmh_mldsa_tfm_ctx), + }, + }, + { + .sign =3D cmh_mldsa_sign, + .verify =3D cmh_mldsa_verify, + .set_pub_key =3D cmh_mldsa_set_pub_key, + .set_priv_key =3D cmh_mldsa_set_priv_key, + .key_size =3D cmh_mldsa_key_size, + .max_size =3D cmh_mldsa_max_size, + .init =3D cmh_mldsa_87_init, + .exit =3D cmh_mldsa_exit, + .base =3D { + .cra_name =3D "mldsa87", + .cra_driver_name =3D "rambus-cmh-mldsa87", + .cra_priority =3D 5001, + .cra_module =3D THIS_MODULE, + .cra_ctxsize =3D sizeof(struct cmh_mldsa_tfm_ctx), + }, + }, +}; + +/** + * cmh_pqc_mldsa_register() - Register ML-DSA akcipher algorithms with the= crypto framework + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_pqc_mldsa_register(void) +{ + int ret, i; + + if (!cmh_core_present(CMH_CORE_QSE)) + return 0; + + for (i =3D 0; i < ARRAY_SIZE(cmh_mldsa_algs); i++) { + ret =3D crypto_register_sig(&cmh_mldsa_algs[i]); + if (ret) { + dev_err(cmh_dev(), "cmh: failed to register %s (%d)\n", + cmh_mldsa_algs[i].base.cra_name, ret); + goto err_unregister; + } + } + + return 0; + +err_unregister: + while (i--) + crypto_unregister_sig(&cmh_mldsa_algs[i]); + return ret; +} + +/** + * cmh_pqc_mldsa_unregister() - Unregister ML-DSA akcipher algorithms from= the crypto framework + */ +void cmh_pqc_mldsa_unregister(void) +{ + int i =3D ARRAY_SIZE(cmh_mldsa_algs); + + if (!cmh_core_present(CMH_CORE_QSE)) + return; + + while (i--) + crypto_unregister_sig(&cmh_mldsa_algs[i]); +} diff --git a/drivers/crypto/cmh/cmh_pqc_sizes.c b/drivers/crypto/cmh/cmh_pq= c_sizes.c new file mode 100644 index 000000000000..39e3d56f4312 --- /dev/null +++ b/drivers/crypto/cmh/cmh_pqc_sizes.c @@ -0,0 +1,39 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- PQC Algorithm Size Tables + * + * Centralised ML-DSA and SLH-DSA parameter-size arrays. Declared + * extern in cmh_qse_abi.h / cmh_hcq_abi.h, defined here once to + * avoid per-TU duplication. + */ + +#include +#include +#include + +#include "cmh_qse_abi.h" +#include "cmh_hcq_abi.h" + +/* ML-DSA size tables (indexed by ml_dsa_mode_idx()) */ +const u32 ml_dsa_pk_size[3] =3D { 1312U, 1952U, 2592U }; +const u32 ml_dsa_sk_size[3] =3D { 2560U, 4032U, 4896U }; +const u32 ml_dsa_sk_size_masked[3] =3D { 3360U, 5472U, 6368U }; +const u32 ml_dsa_sig_size[3] =3D { 2420U, 3309U, 4627U }; + +static_assert(ARRAY_SIZE(ml_dsa_pk_size) =3D=3D ARRAY_SIZE(ml_dsa_sk_size)= ); +static_assert(ARRAY_SIZE(ml_dsa_pk_size) =3D=3D ARRAY_SIZE(ml_dsa_sk_size_= masked)); +static_assert(ARRAY_SIZE(ml_dsa_pk_size) =3D=3D ARRAY_SIZE(ml_dsa_sig_size= )); + +/* SLH-DSA n-values and signature sizes (indexed by param_set - 1) */ +const u32 slhdsa_n[12] =3D { + 16, 16, 24, 24, 32, 32, /* SHAKE 128s/f, 192s/f, 256s/f */ + 16, 16, 24, 24, 32, 32, /* SHA2 128s/f, 192s/f, 256s/f */ +}; + +const u32 slhdsa_sig_size[12] =3D { + 7856, 17088, 16224, 35664, 29792, 49856, /* SHAKE */ + 7856, 17088, 16224, 35664, 29792, 49856, /* SHA2 */ +}; + +static_assert(ARRAY_SIZE(slhdsa_n) =3D=3D ARRAY_SIZE(slhdsa_sig_size)); diff --git a/drivers/crypto/cmh/cmh_qse.c b/drivers/crypto/cmh/cmh_qse.c new file mode 100644 index 000000000000..257dc3ee29a8 --- /dev/null +++ b/drivers/crypto/cmh/cmh_qse.c @@ -0,0 +1,211 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- QSE Core VCQ Builders + * + * VCQ builder functions for ML-KEM and ML-DSA commands (plain and masked). + * Each function populates a single vcq_cmd slot. Callers assemble + * complete VCQs with header + command(s) + flush, then submit via + * cmh_tm_submit_sync(). + */ + +#include + +#include "cmh_sys.h" + +/* -- QSE flush -- */ + +/** + * vcq_add_qse_flush() - Build a QSE flush VCQ command + * @slot: VCQ command slot to populate + * @core_id: Hardware core ID for dispatch + */ +void vcq_add_qse_flush(struct vcq_cmd *slot, u32 core_id) +{ + vcq_add_flush(slot, core_id); +} + +/* -- ML-KEM -- */ + +/** + * vcq_add_qse_ml_kem_keygen() - Build an ML-KEM key generation VCQ command + * @slot: VCQ command slot to populate + * @core_id: Hardware core ID for dispatch + * @k: ML-KEM security parameter (k =3D 2, 3, or 4) + * @flags: Command flags + * @seed: DMA address of seed input buffer + * @z: DMA address of implicit rejection value buffer + * @ek: DMA address of encapsulation key output buffer + * @dk: DMA address of decapsulation key output buffer + * @dk_type: Decapsulation key datastore type + * @masked: Use masked (side-channel protected) variant + */ +void vcq_add_qse_ml_kem_keygen(struct vcq_cmd *slot, u32 core_id, u32 k, u= 32 flags, + u64 seed, u64 z, u64 ek, u64 dk, u32 dk_type, + bool masked) +{ + u32 cmd_id =3D masked ? QSE_CMD_ML_KEM_KEYGEN_MASKED + : QSE_CMD_ML_KEM_KEYGEN; + + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, cmd_id); + slot->hwc.qse.cmd_ml_kem_keygen.k =3D k; + slot->hwc.qse.cmd_ml_kem_keygen.flags =3D flags; + slot->hwc.qse.cmd_ml_kem_keygen.seed =3D seed; + slot->hwc.qse.cmd_ml_kem_keygen.z =3D z; + slot->hwc.qse.cmd_ml_kem_keygen.ek =3D ek; + slot->hwc.qse.cmd_ml_kem_keygen.dk =3D dk; + slot->hwc.qse.cmd_ml_kem_keygen.dk_type =3D dk_type; +} + +/** + * vcq_add_qse_ml_kem_enc() - Build an ML-KEM encapsulation VCQ command + * @slot: VCQ command slot to populate + * @core_id: Hardware core ID for dispatch + * @k: ML-KEM security parameter (k =3D 2, 3, or 4) + * @flags: Command flags + * @coin: DMA address of encapsulation coin/randomness buffer + * @ek: DMA address of encapsulation key input buffer + * @ct: DMA address of ciphertext output buffer + * @ss: DMA address of shared secret output buffer + * @ss_type: Shared secret datastore type + * @masked: Use masked (side-channel protected) variant + */ +void vcq_add_qse_ml_kem_enc(struct vcq_cmd *slot, u32 core_id, u32 k, u32 = flags, + u64 coin, u64 ek, u64 ct, u64 ss, u32 ss_type, + bool masked) +{ + u32 cmd_id =3D masked ? QSE_CMD_ML_KEM_ENC_MASKED + : QSE_CMD_ML_KEM_ENC; + + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, cmd_id); + slot->hwc.qse.cmd_ml_kem_enc.k =3D k; + slot->hwc.qse.cmd_ml_kem_enc.flags =3D flags; + slot->hwc.qse.cmd_ml_kem_enc.coin =3D coin; + slot->hwc.qse.cmd_ml_kem_enc.ek =3D ek; + slot->hwc.qse.cmd_ml_kem_enc.ct =3D ct; + slot->hwc.qse.cmd_ml_kem_enc.ss =3D ss; + slot->hwc.qse.cmd_ml_kem_enc.ss_type =3D ss_type; +} + +/** + * vcq_add_qse_ml_kem_dec() - Build an ML-KEM decapsulation VCQ command + * @slot: VCQ command slot to populate + * @core_id: Hardware core ID for dispatch + * @k: ML-KEM security parameter (k =3D 2, 3, or 4) + * @flags: Command flags + * @ct: DMA address of ciphertext input buffer + * @dk: DMA address of decapsulation key input buffer + * @ss: DMA address of shared secret output buffer + * @ss_type: Shared secret datastore type + * @masked: Use masked (side-channel protected) variant + */ +void vcq_add_qse_ml_kem_dec(struct vcq_cmd *slot, u32 core_id, u32 k, u32 = flags, + u64 ct, u64 dk, u64 ss, u32 ss_type, + bool masked) +{ + u32 cmd_id =3D masked ? QSE_CMD_ML_KEM_DEC_MASKED + : QSE_CMD_ML_KEM_DEC; + + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, cmd_id); + slot->hwc.qse.cmd_ml_kem_dec.k =3D k; + slot->hwc.qse.cmd_ml_kem_dec.flags =3D flags; + slot->hwc.qse.cmd_ml_kem_dec.ct =3D ct; + slot->hwc.qse.cmd_ml_kem_dec.dk =3D dk; + slot->hwc.qse.cmd_ml_kem_dec.ss =3D ss; + slot->hwc.qse.cmd_ml_kem_dec.ss_type =3D ss_type; +} + +/* -- ML-DSA -- */ + +/** + * vcq_add_qse_ml_dsa_keygen() - Build an ML-DSA key generation VCQ command + * @slot: VCQ command slot to populate + * @core_id: Hardware core ID for dispatch + * @mode: ML-DSA mode (44, 65, or 87) + * @flags: Command flags + * @seed: DMA address of seed input buffer + * @pk: DMA address of public key output buffer + * @sk: DMA address of secret key output buffer + * @sk_type: Secret key datastore type + * @masked: Use masked (side-channel protected) variant + */ +void vcq_add_qse_ml_dsa_keygen(struct vcq_cmd *slot, u32 core_id, u32 mode= , u32 flags, + u64 seed, u64 pk, u64 sk, u32 sk_type, + bool masked) +{ + u32 cmd_id =3D masked ? QSE_CMD_ML_DSA_KEYGEN_MASKED + : QSE_CMD_ML_DSA_KEYGEN; + + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, cmd_id); + slot->hwc.qse.cmd_ml_dsa_keygen.mode =3D mode; + slot->hwc.qse.cmd_ml_dsa_keygen.flags =3D flags; + slot->hwc.qse.cmd_ml_dsa_keygen.seed =3D seed; + slot->hwc.qse.cmd_ml_dsa_keygen.pk =3D pk; + slot->hwc.qse.cmd_ml_dsa_keygen.sk =3D sk; + slot->hwc.qse.cmd_ml_dsa_keygen.sk_type =3D sk_type; +} + +/** + * vcq_add_qse_ml_dsa_sign() - Build an ML-DSA signing VCQ command + * @slot: VCQ command slot to populate + * @core_id: Hardware core ID for dispatch + * @mode: ML-DSA mode (44, 65, or 87) + * @flags: Command flags + * @rnd: DMA address of signing randomness buffer + * @m: DMA address of message buffer + * @sk: DMA address of secret key buffer + * @sig: DMA address of signature output buffer + * @mlen: Length of message in bytes + * @masked: Use masked (side-channel protected) variant + */ +void vcq_add_qse_ml_dsa_sign(struct vcq_cmd *slot, u32 core_id, u32 mode, = u32 flags, + u64 rnd, u64 m, u64 sk, u64 sig, u32 mlen, + bool masked) +{ + u32 cmd_id =3D masked ? QSE_CMD_ML_DSA_SIGN_MASKED + : QSE_CMD_ML_DSA_SIGN; + + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, cmd_id); + slot->hwc.qse.cmd_ml_dsa_sign.mode =3D mode; + slot->hwc.qse.cmd_ml_dsa_sign.flags =3D flags; + slot->hwc.qse.cmd_ml_dsa_sign.rnd =3D rnd; + slot->hwc.qse.cmd_ml_dsa_sign.m =3D m; + slot->hwc.qse.cmd_ml_dsa_sign.sk =3D sk; + slot->hwc.qse.cmd_ml_dsa_sign.sig =3D sig; + slot->hwc.qse.cmd_ml_dsa_sign.mlen =3D mlen; +} + +/** + * vcq_add_qse_ml_dsa_verify() - Build an ML-DSA signature verify VCQ comm= and + * @slot: VCQ command slot to populate + * @core_id: Hardware core ID for dispatch + * @mode: ML-DSA mode (44, 65, or 87) + * @flags: Command flags + * @m: DMA address of message buffer + * @pk: DMA address of public key buffer + * @sig: DMA address of signature buffer to verify + * @mlen: Length of message in bytes + */ +void vcq_add_qse_ml_dsa_verify(struct vcq_cmd *slot, u32 core_id, u32 mode= , u32 flags, + u64 m, u64 pk, u64 sig, u32 mlen) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, QSE_CMD_ML_DSA_VERIFY); + slot->hwc.qse.cmd_ml_dsa_verify.mode =3D mode; + slot->hwc.qse.cmd_ml_dsa_verify.flags =3D flags; + slot->hwc.qse.cmd_ml_dsa_verify.m =3D m; + slot->hwc.qse.cmd_ml_dsa_verify.pk =3D pk; + slot->hwc.qse.cmd_ml_dsa_verify.sig =3D sig; + slot->hwc.qse.cmd_ml_dsa_verify.mlen =3D mlen; +} --=20 2.43.7 From nobody Fri Sep 25 01:20:34 2026 Received: from DM1PR04CU001.outbound.protection.outlook.com (mail-centralusazon11020125.outbound.protection.outlook.com [52.101.61.125]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6B2A14A6CDA; Thu, 17 Sep 2026 23:00:45 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.61.125 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789686054; cv=fail; b=tXmgAMXL73ulOsGTSrrMtdPwjmb+04Vk2aZ2U7mf5nTJ5wV3DHRQMP5kJSdd4ERbCk4UoGRuqxNe9Xma6fI6GzPzJPaQFCgsOg5JS2TxEKpnArOkNNdvtfFNuYYAY8n6nThh2e1LAXehBVbSOA/+yRs9hRVxrDePjit2TG/Su8A= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789686054; c=relaxed/simple; bh=68I1sXe+HtIKbu02GjegNXZQO8GLkzOwWqtVXSWmXBk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=j4e5QHKtTQKG776fzH5xb9QIKpeGKoNzvDQY6AP8vSpKtfxl8v464IGNtIGSVU1bAhIBG0cbYSoamQVBBzoDtdY9k7hITD8svG9hKhhVdlOkpi9MoLYLoYHXpaQnG9WGtzkp3JgALujMB9a4ruGpd4X07aWSK/XNKhVzfv483SQ= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com; spf=fail smtp.mailfrom=rambus.com; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b=42V5z6no; arc=fail smtp.client-ip=52.101.61.125 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=rambus.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b="42V5z6no" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=NCrSpzMdOZMwF/v72jXrcK5CYL51DcEuknbVPDyELZ10sYQPcRJkd/680U1L+QZ7h1GZ2y4Qaa0TAViJ6mnW/h9yVjM1g8HJVeyFHBWvwuWZUGENCU+WJImo3lSKYU02cfnmIMh+iI32avdSOCD62dJeR6XEgGpWMC2hXLATXGDGXlJhH98rVtAoFJ1QPKCKGqkn8aqdYSIO3OtHrO2FfU2HWlTxJqKp1p49sfgCxRK4tlDyv8APgkPj9kVIhXfNCWetW0LRR1g5Nel/Io4KKMUrSceN4kIJ8/RPF+b+q8oacBQLxoebML07yL7DgrqB7549xziafGW1vwHhjXSToQ== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=KR8e/pz0axY+p4O4mtYS1qvloVTWC2KmdGSC6wXAb/k=; b=AHQd49JhB9pguci+MbHGUBncXBGHEaX5wUkYQd7hzBzd8MwmvNi72PzIfDO36j1Gmikp9yfaGo23ATsph7PrvlGvuCcatejCTR795NRwR/+NgmQsczGlWPtJheHCp/r/buS4G5TdyDhS45ElSAwTpG9JZC7KlYKHsfv3CmJyLPu/3UiRHgXCVquCe9/kL+POIbv3m3dfHz8MhHWLzZTgjscxGA85BDe0GzRlCA8+Glcsj9i67N+uTLxtXsn9keImQGJQ12TKQ0O+abNW5386Cwe8M43/IPWwe3zUkXvhY+qpJHfIBihZxfPm1Swq4GhvoxlnayAS0NlWhcw3KYQknA== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 192.86.86.210) smtp.rcpttodomain=cryptography.com smtp.mailfrom=rambus.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=rambus.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=rambus.com; s=selector1; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=KR8e/pz0axY+p4O4mtYS1qvloVTWC2KmdGSC6wXAb/k=; b=42V5z6noFaE05ABFqis8AHuxa3lxUfF12Pfh2OXehoQCEYNgBzyBHdxvTF8hnDXLUEo8qDfp7hhkdzDCvEZlalq3frLaAz1nDuZht4nEsoH+bl0QfZtZNAaR7/SdCMpp7zSLXugRRt474QaEoEmEfG/EW2kdhC4VV6TfD1XPzWDy4kTpuBtzJU0QNRx5Hh3N7Wf9qFsgfhK4z6iHMYnisKHHJa7nXG3YYqxIej7giUayS5jTlXchK8C48zSYPT4GSLYgFTSYGyC40FIb8yy3rb4DaMMdW2GsfmihaTzzyNizhlp4x9flk/bWgW4Wdv2O4SRRDU51VVCi+ymNQEr9Vw== Received: from BN9PR03CA0221.namprd03.prod.outlook.com (2603:10b6:408:f8::16) by DS2PR04MB993776.namprd04.prod.outlook.com (2603:10b6:8:338::6) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.428.13; Thu, 17 Sep 2026 22:59:36 +0000 Received: from BN7PEPF00000090.namprd03.prod.outlook.com (2603:10b6:408:f8:cafe::2a) by BN9PR03CA0221.outlook.office365.com (2603:10b6:408:f8::16) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.406.12 via Frontend Transport; Thu, 17 Sep 2026 22:59:36 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 192.86.86.210) smtp.mailfrom=rambus.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=rambus.com; Received-SPF: Pass (protection.outlook.com: domain of rambus.com designates 192.86.86.210 as permitted sender) receiver=protection.outlook.com; client-ip=192.86.86.210; helo=hqxsv-psmtppxy02.rambus.com; pr=C Received: from hqxsv-psmtppxy02.rambus.com (192.86.86.210) by BN7PEPF00000090.mail.protection.outlook.com (10.167.245.68) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.451.8 via Frontend Transport; Thu, 17 Sep 2026 22:59:34 +0000 Received: from hqxsv-cmdev3-aousherovitch.rambus.com (hqn-lb-int-float.rambus.com [10.12.20.20]) by hqxsv-psmtppxy02.rambus.com (Postfix) with ESMTP id 848FD180176F; Thu, 17 Sep 2026 22:59:32 +0000 (UTC) From: Alex Ousherovitch To: Albert Ou , Alex Ousherovitch , Conor Dooley , "David S. Miller" , Herbert Xu , Jonathan Corbet , Krzysztof Kozlowski , Palmer Dabbelt , Paul Walmsley , Rob Herring , Saravanakrishnan Krishnamoorthy , Shuah Khan Cc: Alexandre Ghiti , devicetree@vger.kernel.org, Joel Wittenauer , linux-api@vger.kernel.org, linux-crypto@vger.kernel.org, linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-riscv@lists.infradead.org, Shuah Khan , Thi Nguyen Subject: [PATCH v5 16/19] crypto: cmh - add SLH-DSA/LMS/XMSS (HCQ) Date: Thu, 17 Sep 2026 15:59:25 -0700 Message-ID: <20260917225929.2494111-17-aousherovitch@rambus.com> X-Mailer: git-send-email 2.43.7 In-Reply-To: <20260917225929.2494111-1-aousherovitch@rambus.com> References: <20260917225929.2494111-1-aousherovitch@rambus.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: BN7PEPF00000090:EE_|DS2PR04MB993776:EE_ X-MS-Office365-Filtering-Correlation-Id: f438711e-519c-4484-113f-08df150f57d0 X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|376014|7416014|82310400026|1800799024|23010399003|36860700016|921020|3023799007|6133799003|10067099003|11063799006|56012099006|22082099003|18002099003; X-Microsoft-Antispam-Message-Info: y2VHzeI22U3NQnUXXDoYRIqIOrpYtYtxYXwgHbmiSPDV+EThQKqeCXKkSCKQ+m93o/nvNCQpm7Auq6o5/L9s+uiSyEY5oPTeaFNSZeIt6ioGJd1AAfCgEZ7o8XQ0/BEC2laVIPT0hH79W9tdvSUbbFSAXhoM4KwTfrkOyzZvnJKjyKtRGELBsfo/odAG31xNzaZ9kpkUUdD+cyVtwnp0htkvzzcTy9mGqWWTgB0fb0d9J9dJv9lNJcmvG4UzBV/SYvFK0QJaQydgPAq/rjguPsd6nQObfUyciAF9rw0PfZxvQ0NP+d6UxhNiEUPa1u3IM5cvSCzAsMNFPfLZzoXqVceqSTahbDOsSrAap/KtlZq9XrLhIANPEeipPPaekbSFdGJfzK/73TZQ/Lhc6UnBVzNVOAb3YnvMks2iJPt5uEJl8Bu7DdOE8P3uzZWsExfZxfIH4zh5XaU05E2uiEVs0s+FD7XZMaLxnxURgIlCme4UVrMYiwV+5tn7Lb5xuVevdff+B+k2NAjUkewNZJ+d0ktUnV/Yt7SS/y4LYGQBUSL7CWvNVS/wOcuXg5de2YMDppu785dWXiUBNg1qWflBV1uRsKIeMGPEN9WZ4U7yJDizGFuo/Fhqt6SSHrnnNhWUyHzNQF44IWnsrQs48/4C03FwNQCn5b8pQl5yj4b9otmkxN0/3kfbSmwHZi/IV1olJllL7LHhd84fllnSO5+0ePwuIUKNI8n1VNm1U6nywIx4IcYexmcXzTucn92pBQU9 X-Forefront-Antispam-Report: CIP:192.86.86.210;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:hqxsv-psmtppxy02.rambus.com;PTR:InfoDomainNonexistent;CAT:NONE;SFS:(13230040)(376014)(7416014)(82310400026)(1800799024)(23010399003)(36860700016)(921020)(3023799007)(6133799003)(10067099003)(11063799006)(56012099006)(22082099003)(18002099003);DIR:OUT;SFP:1102; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: qc0XloygO3Lct9x9OxtVb0uBFb7lRuO7LkW/6+/+e/dhT6eA3KpwUxPLVKNUhrYh7FpG1Kh0yv3qiHD9vJHcc2H3SdnjO0l2VphTD07KrEsGshChi/4o7XiHpGSfn9w2YeGYCyfgmReCWE8Zo82dkI1xvzrsn8owxwhHDO9xBJ5nKfQWNHdydpPK0IdrqubwbjCDdChOkQQS56UZDZBnDUHjSG6QjL6ZTEK+M64Y7JozKkAIrjsNHDEoXZhGWlJ57gjn3UO+qwqI7BrpGAgIwfzFauU3tbi2KaJV/TguSfWfInn/gRMXEGDW4GDBqs/rdgTc5vDb/xJxIsP8fJCIax/MFEKoDMql0Ocv7Rj/Tc6fvGnDgqa1gJGey2hihSPhC7/vRrNwuUB7qgc/RHwNomgT0TBC5hf9FQ+H3SqN2/5VCMPcWY3q5HNsSkDcqzkD X-OriginatorOrg: rambus.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 17 Sep 2026 22:59:34.9534 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: f438711e-519c-4484-113f-08df150f57d0 X-MS-Exchange-CrossTenant-Id: bd0ba799-c2b9-413c-9c56-5d1731c4827c X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=bd0ba799-c2b9-413c-9c56-5d1731c4827c;Ip=[192.86.86.210];Helo=[hqxsv-psmtppxy02.rambus.com] X-MS-Exchange-CrossTenant-AuthSource: BN7PEPF00000090.namprd03.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: DS2PR04MB993776 Content-Type: text/plain; charset="utf-8" Register SLH-DSA, LMS, LMS-HSS, XMSS, and XMSS-MT algorithms using the CMH HCQ core (core ID 0x08). SLH-DSA is registered as a sig algorithm with sign and verify support. LMS, LMS-HSS, XMSS, and XMSS-MT are registered as verify-only sig algorithms: their stateful signing semantics (one-time-key private state the signer must track) are not modeled by the kernel crypto API, so only verification is exposed. Signed-off-by: Alex Ousherovitch Co-developed-by: Saravanakrishnan Krishnamoorthy Signed-off-by: Saravanakrishnan Krishnamoorthy --- drivers/crypto/cmh/Makefile | 6 +- drivers/crypto/cmh/cmh_hcq.c | 313 +++++++++++++++++++++++ drivers/crypto/cmh/cmh_main.c | 24 ++ drivers/crypto/cmh/cmh_pqc_lms.c | 244 ++++++++++++++++++ drivers/crypto/cmh/cmh_pqc_slhdsa.c | 382 ++++++++++++++++++++++++++++ drivers/crypto/cmh/cmh_pqc_xmss.c | 244 ++++++++++++++++++ 6 files changed, 1212 insertions(+), 1 deletion(-) create mode 100644 drivers/crypto/cmh/cmh_hcq.c create mode 100644 drivers/crypto/cmh/cmh_pqc_lms.c create mode 100644 drivers/crypto/cmh/cmh_pqc_slhdsa.c create mode 100644 drivers/crypto/cmh/cmh_pqc_xmss.c diff --git a/drivers/crypto/cmh/Makefile b/drivers/crypto/cmh/Makefile index 7196d992487e..a78311bf3195 100644 --- a/drivers/crypto/cmh/Makefile +++ b/drivers/crypto/cmh/Makefile @@ -36,7 +36,11 @@ cmh-y :=3D \ cmh_pke_ecdh.o \ cmh_qse.o \ cmh_pqc_mldsa.o \ - cmh_pqc_sizes.o + cmh_pqc_sizes.o \ + cmh_hcq.o \ + cmh_pqc_slhdsa.o \ + cmh_pqc_lms.o \ + cmh_pqc_xmss.o =20 # Management ioctl device (/dev/cmh_mgmt): key lifecycle, PKE, PQC ioctls. cmh-$(CONFIG_CRYPTO_DEV_CMH_MGMT) +=3D \ diff --git a/drivers/crypto/cmh/cmh_hcq.c b/drivers/crypto/cmh/cmh_hcq.c new file mode 100644 index 000000000000..8fc3a5cb0f9f --- /dev/null +++ b/drivers/crypto/cmh/cmh_hcq.c @@ -0,0 +1,313 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- HCQ Core VCQ Builders + * + * VCQ builder functions for SLH-DSA, LMS, and XMSS commands. + * Each function populates a single vcq_cmd slot. Callers assemble + * complete VCQs with header + command(s) + flush, then submit via + * cmh_tm_submit_sync(). + */ + +#include + +#include "cmh_sys.h" + +/* -- HCQ flush -- */ + +/** + * vcq_add_hcq_flush() - Build an HCQ flush VCQ command + * @slot: VCQ command slot to populate + * @core_id: Hardware core ID for dispatch + */ +void vcq_add_hcq_flush(struct vcq_cmd *slot, u32 core_id) +{ + vcq_add_flush(slot, core_id); +} + +/* -- SLH-DSA -- */ + +/** + * vcq_add_hcq_slhdsa_keygen() - Build an SLH-DSA key generation VCQ comma= nd + * @slot: VCQ command slot to populate + * @core_id: Hardware core ID for dispatch + * @param_set: SLH-DSA parameter set identifier + * @seed_len: Length of seed buffer in bytes + * @pk_len: Length of public key buffer in bytes + * @sk_len: Length of secret key buffer in bytes + * @seed: DMA address of seed input buffer + * @pk: DMA address of public key output buffer + * @sk: DMA address of secret key output buffer + */ +void vcq_add_hcq_slhdsa_keygen(struct vcq_cmd *slot, u32 core_id, u32 para= m_set, + u32 seed_len, u32 pk_len, u32 sk_len, + u64 seed, u64 pk, u64 sk) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, HCQ_CMD_SLHDSA_KEYGEN); + slot->hwc.hcq.cmd_slhdsa_keygen.parameter_set =3D param_set; + slot->hwc.hcq.cmd_slhdsa_keygen.seed_len =3D seed_len; + slot->hwc.hcq.cmd_slhdsa_keygen.pk_len =3D pk_len; + slot->hwc.hcq.cmd_slhdsa_keygen.sk_len =3D sk_len; + slot->hwc.hcq.cmd_slhdsa_keygen.seed =3D seed; + slot->hwc.hcq.cmd_slhdsa_keygen.pk =3D pk; + slot->hwc.hcq.cmd_slhdsa_keygen.sk =3D sk; +} + +/** + * vcq_add_hcq_slhdsa_sign() - Build an SLH-DSA signing VCQ command + * @slot: VCQ command slot to populate + * @core_id: Hardware core ID for dispatch + * @param_set: SLH-DSA parameter set identifier + * @msg_len: Length of message buffer in bytes + * @ctx_len: Length of context string in bytes + * @add_random: DMA address of additional randomness buffer + * @msg: DMA address of message buffer + * @ctx: DMA address of context string buffer + * @sk: DMA address of secret key buffer + * @sig: DMA address of signature output buffer + */ +void vcq_add_hcq_slhdsa_sign(struct vcq_cmd *slot, u32 core_id, u32 param_= set, + u32 msg_len, u32 ctx_len, + u64 add_random, u64 msg, u64 ctx, + u64 sk, u64 sig) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, HCQ_CMD_SLHDSA_SIGN); + slot->hwc.hcq.cmd_slhdsa_sign.parameter_set =3D param_set; + slot->hwc.hcq.cmd_slhdsa_sign.message_len =3D msg_len; + slot->hwc.hcq.cmd_slhdsa_sign.add_random =3D add_random; + slot->hwc.hcq.cmd_slhdsa_sign.message =3D msg; + slot->hwc.hcq.cmd_slhdsa_sign.context =3D ctx; + slot->hwc.hcq.cmd_slhdsa_sign.sk =3D sk; + slot->hwc.hcq.cmd_slhdsa_sign.sig =3D sig; + slot->hwc.hcq.cmd_slhdsa_sign.context_len =3D ctx_len; +} + +/** + * vcq_add_hcq_slhdsa_sign_internal() - Build an SLH-DSA internal signing = VCQ command + * @slot: VCQ command slot to populate + * @core_id: Hardware core ID for dispatch + * @param_set: SLH-DSA parameter set identifier + * @msg_len: Length of message buffer in bytes + * @add_random: DMA address of additional randomness buffer + * @msg: DMA address of message buffer + * @sk: DMA address of secret key buffer + * @sig: DMA address of signature output buffer + */ +void vcq_add_hcq_slhdsa_sign_internal(struct vcq_cmd *slot, u32 core_id, u= 32 param_set, + u32 msg_len, u64 add_random, + u64 msg, u64 sk, u64 sig) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, HCQ_CMD_SLHDSA_SIGN_INTERNAL); + slot->hwc.hcq.cmd_slhdsa_sign_internal.parameter_set =3D param_set; + slot->hwc.hcq.cmd_slhdsa_sign_internal.message_len =3D msg_len; + slot->hwc.hcq.cmd_slhdsa_sign_internal.add_random =3D add_random; + slot->hwc.hcq.cmd_slhdsa_sign_internal.message =3D msg; + slot->hwc.hcq.cmd_slhdsa_sign_internal.sk =3D sk; + slot->hwc.hcq.cmd_slhdsa_sign_internal.sig =3D sig; +} + +/** + * vcq_add_hcq_slhdsa_verify() - Build an SLH-DSA verification VCQ command + * @slot: VCQ command slot to populate + * @core_id: Hardware core ID for dispatch + * @param_set: SLH-DSA parameter set identifier + * @msg_len: Length of message buffer in bytes + * @ctx_len: Length of context string in bytes + * @msg: DMA address of message buffer + * @ctx: DMA address of context string buffer + * @pk: DMA address of public key buffer + * @sig: DMA address of signature buffer to verify + */ +void vcq_add_hcq_slhdsa_verify(struct vcq_cmd *slot, u32 core_id, u32 para= m_set, + u32 msg_len, u32 ctx_len, + u64 msg, u64 ctx, u64 pk, u64 sig) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, HCQ_CMD_SLHDSA_VERIFY); + slot->hwc.hcq.cmd_slhdsa_verify.parameter_set =3D param_set; + slot->hwc.hcq.cmd_slhdsa_verify.message_len =3D msg_len; + slot->hwc.hcq.cmd_slhdsa_verify.message =3D msg; + slot->hwc.hcq.cmd_slhdsa_verify.context =3D ctx; + slot->hwc.hcq.cmd_slhdsa_verify.pk =3D pk; + slot->hwc.hcq.cmd_slhdsa_verify.sig =3D sig; + slot->hwc.hcq.cmd_slhdsa_verify.context_len =3D ctx_len; +} + +/** + * vcq_add_hcq_slhdsa_sign_prehash() - Build an SLH-DSA prehash signing VC= Q command + * @slot: VCQ command slot to populate + * @core_id: Hardware core ID for dispatch + * @cmd: VCQ command ID (sign-prehash variant) + * @param_set: SLH-DSA parameter set identifier + * @prehash_algo: Prehash algorithm identifier + * @msg_len: Length of message buffer in bytes + * @ctx_len: Length of context string in bytes + * @add_random: DMA address of additional randomness buffer + * @msg: DMA address of message buffer + * @ctx: DMA address of context string buffer + * @sk: DMA address of secret key buffer + * @sig: DMA address of signature output buffer + */ +void vcq_add_hcq_slhdsa_sign_prehash(struct vcq_cmd *slot, u32 core_id, + u32 cmd, u32 param_set, u32 prehash_algo, + u32 msg_len, u32 ctx_len, + u64 add_random, u64 msg, u64 ctx, + u64 sk, u64 sig) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, cmd); + slot->hwc.hcq.cmd_slhdsa_sign_prehash.parameter_set =3D param_set; + slot->hwc.hcq.cmd_slhdsa_sign_prehash.prehash_algo =3D prehash_algo; + slot->hwc.hcq.cmd_slhdsa_sign_prehash.message_len =3D msg_len; + slot->hwc.hcq.cmd_slhdsa_sign_prehash.context_len =3D ctx_len; + slot->hwc.hcq.cmd_slhdsa_sign_prehash.add_random =3D add_random; + slot->hwc.hcq.cmd_slhdsa_sign_prehash.message =3D msg; + slot->hwc.hcq.cmd_slhdsa_sign_prehash.context =3D ctx; + slot->hwc.hcq.cmd_slhdsa_sign_prehash.sk =3D sk; + slot->hwc.hcq.cmd_slhdsa_sign_prehash.sig =3D sig; +} + +/** + * vcq_add_hcq_slhdsa_verify_prehash() - Build an SLH-DSA prehash verify V= CQ command + * @slot: VCQ command slot to populate + * @core_id: Hardware core ID for dispatch + * @cmd: VCQ command ID (verify-prehash variant) + * @param_set: SLH-DSA parameter set identifier + * @prehash_algo: Prehash algorithm identifier + * @msg_len: Length of message buffer in bytes + * @ctx_len: Length of context string in bytes + * @msg: DMA address of message buffer + * @ctx: DMA address of context string buffer + * @pk: DMA address of public key buffer + * @sig: DMA address of signature buffer to verify + */ +void vcq_add_hcq_slhdsa_verify_prehash(struct vcq_cmd *slot, u32 core_id, + u32 cmd, u32 param_set, u32 prehash_algo, + u32 msg_len, u32 ctx_len, + u64 msg, u64 ctx, u64 pk, u64 sig) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, cmd); + slot->hwc.hcq.cmd_slhdsa_verify_prehash.parameter_set =3D param_set; + slot->hwc.hcq.cmd_slhdsa_verify_prehash.prehash_algo =3D prehash_algo; + slot->hwc.hcq.cmd_slhdsa_verify_prehash.message_len =3D msg_len; + slot->hwc.hcq.cmd_slhdsa_verify_prehash.context_len =3D ctx_len; + slot->hwc.hcq.cmd_slhdsa_verify_prehash.message =3D msg; + slot->hwc.hcq.cmd_slhdsa_verify_prehash.context =3D ctx; + slot->hwc.hcq.cmd_slhdsa_verify_prehash.pk =3D pk; + slot->hwc.hcq.cmd_slhdsa_verify_prehash.sig =3D sig; +} + +/** + * vcq_add_hcq_slhdsa_verify_internal() - Build an SLH-DSA internal verify= VCQ command + * @slot: VCQ command slot to populate + * @core_id: Hardware core ID for dispatch + * @param_set: SLH-DSA parameter set identifier + * @msg_len: Length of message buffer in bytes + * @msg: DMA address of message buffer + * @pk: DMA address of public key buffer + * @sig: DMA address of signature buffer to verify + */ +void vcq_add_hcq_slhdsa_verify_internal(struct vcq_cmd *slot, u32 core_id,= u32 param_set, + u32 msg_len, u64 msg, u64 pk, u64 sig) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, + HCQ_CMD_SLHDSA_VERIFY_INTERNAL); + slot->hwc.hcq.cmd_slhdsa_verify_internal.parameter_set =3D param_set; + slot->hwc.hcq.cmd_slhdsa_verify_internal.message_len =3D msg_len; + slot->hwc.hcq.cmd_slhdsa_verify_internal.message =3D msg; + slot->hwc.hcq.cmd_slhdsa_verify_internal.pk =3D pk; + slot->hwc.hcq.cmd_slhdsa_verify_internal.sig =3D sig; +} + +/** + * vcq_add_hcq_slhdsa_pubgen() - Build an SLH-DSA public key generation VC= Q command + * @slot: VCQ command slot to populate + * @core_id: Hardware core ID for dispatch + * @param_set: SLH-DSA parameter set identifier + * @sk_len: Length of secret key buffer in bytes + * @sk: DMA address of secret key input buffer + * @pk: DMA address of public key output buffer + */ +void vcq_add_hcq_slhdsa_pubgen(struct vcq_cmd *slot, u32 core_id, u32 para= m_set, + u32 sk_len, u64 sk, u64 pk) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, HCQ_CMD_SLHDSA_PUBGEN); + slot->hwc.hcq.cmd_slhdsa_pubgen.parameter_set =3D param_set; + slot->hwc.hcq.cmd_slhdsa_pubgen.sk_len =3D sk_len; + slot->hwc.hcq.cmd_slhdsa_pubgen.sk =3D sk; + slot->hwc.hcq.cmd_slhdsa_pubgen.pk =3D pk; +} + +/* -- LMS -- */ + +/** + * vcq_add_hcq_lms_verify() - Build an LMS/HSS signature verify VCQ command + * @slot: VCQ command slot to populate + * @core_id: Hardware core ID for dispatch + * @lms_hss: LMS/HSS mode flag (0 =3D LMS, 1 =3D HSS) + * @pk_len: Length of public key buffer in bytes + * @sig_len: Length of signature buffer in bytes + * @dig_len: Length of digest buffer in bytes + * @pk: DMA address of public key buffer + * @sig: DMA address of signature buffer + * @dig: DMA address of digest buffer + */ +void vcq_add_hcq_lms_verify(struct vcq_cmd *slot, u32 core_id, u32 lms_hss, + u32 pk_len, u32 sig_len, u32 dig_len, + u64 pk, u64 sig, u64 dig) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, HCQ_CMD_LMS_VERIFY); + slot->hwc.hcq.cmd_lms_verify.lms_hss =3D lms_hss; + slot->hwc.hcq.cmd_lms_verify.pk_len =3D pk_len; + slot->hwc.hcq.cmd_lms_verify.sig_len =3D sig_len; + slot->hwc.hcq.cmd_lms_verify.dig_len =3D dig_len; + slot->hwc.hcq.cmd_lms_verify.pk =3D pk; + slot->hwc.hcq.cmd_lms_verify.sig =3D sig; + slot->hwc.hcq.cmd_lms_verify.dig =3D dig; +} + +/* -- XMSS -- */ + +/** + * vcq_add_hcq_xmss_verify() - Build an XMSS/XMSS^MT signature verify VCQ = command + * @slot: VCQ command slot to populate + * @core_id: Hardware core ID for dispatch + * @xmss_mt: XMSS/XMSS^MT mode flag (0 =3D XMSS, 1 =3D XMSS^MT) + * @pk_len: Length of public key buffer in bytes + * @sig_len: Length of signature buffer in bytes + * @dig_len: Length of digest buffer in bytes + * @pk: DMA address of public key buffer + * @sig: DMA address of signature buffer + * @dig: DMA address of digest buffer + */ +void vcq_add_hcq_xmss_verify(struct vcq_cmd *slot, u32 core_id, u32 xmss_m= t, + u32 pk_len, u32 sig_len, u32 dig_len, + u64 pk, u64 sig, u64 dig) +{ + memset(slot, 0, sizeof(*slot)); + slot->magic =3D VCQ_CMD_MAGIC; + slot->id =3D VCQ_CMD_ID(core_id, 0, 1, HCQ_CMD_XMSS_VERIFY); + slot->hwc.hcq.cmd_xmss_verify.xmss_mt =3D xmss_mt; + slot->hwc.hcq.cmd_xmss_verify.pk_len =3D pk_len; + slot->hwc.hcq.cmd_xmss_verify.sig_len =3D sig_len; + slot->hwc.hcq.cmd_xmss_verify.dig_len =3D dig_len; + slot->hwc.hcq.cmd_xmss_verify.pk =3D pk; + slot->hwc.hcq.cmd_xmss_verify.sig =3D sig; + slot->hwc.hcq.cmd_xmss_verify.dig =3D dig; +} diff --git a/drivers/crypto/cmh/cmh_main.c b/drivers/crypto/cmh/cmh_main.c index 7dfcdf395175..bed18321f535 100644 --- a/drivers/crypto/cmh/cmh_main.c +++ b/drivers/crypto/cmh/cmh_main.c @@ -308,6 +308,21 @@ static int cmh_probe(struct platform_device *pdev) if (ret) goto err_pqc_mldsa_register; =20 + /* Register PQC SLH-DSA */ + ret =3D cmh_pqc_slhdsa_register(); + if (ret) + goto err_pqc_slhdsa_register; + + /* Register PQC LMS */ + ret =3D cmh_pqc_lms_register(); + if (ret) + goto err_pqc_lms_register; + + /* Register PQC XMSS */ + ret =3D cmh_pqc_xmss_register(); + if (ret) + goto err_pqc_xmss_register; + /* Register key management device (/dev/cmh_mgmt) */ ret =3D cmh_mgmt_register(); if (ret) @@ -318,6 +333,12 @@ static int cmh_probe(struct platform_device *pdev) return 0; =20 err_mgmt_register: + cmh_pqc_xmss_unregister(); +err_pqc_xmss_register: + cmh_pqc_lms_unregister(); +err_pqc_lms_register: + cmh_pqc_slhdsa_unregister(); +err_pqc_slhdsa_register: cmh_pqc_mldsa_unregister(); err_pqc_mldsa_register: cmh_pke_ecdh_unregister(); @@ -382,6 +403,9 @@ static void cmh_remove(struct platform_device *pdev) cfg =3D &dev->config; =20 cmh_mgmt_unregister(); + cmh_pqc_xmss_unregister(); + cmh_pqc_lms_unregister(); + cmh_pqc_slhdsa_unregister(); cmh_pqc_mldsa_unregister(); cmh_pke_ecdh_unregister(); cmh_pke_ecdsa_unregister(); diff --git a/drivers/crypto/cmh/cmh_pqc_lms.c b/drivers/crypto/cmh/cmh_pqc_= lms.c new file mode 100644 index 000000000000..882142229981 --- /dev/null +++ b/drivers/crypto/cmh/cmh_pqc_lms.c @@ -0,0 +1,244 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- LMS/HSS Signature Driver (verify-only, sig_alg, synchronous) + * + * Registers "lms" and "lms-hss" sig algorithms with verify-only + * callbacks. Sign is not supported (stateful signature -- key + * management must happen externally). + * + * Verify: src =3D raw signature, digest =3D message bytes + * Public key: raw pk bytes (variable length, set via set_pub_key) + */ + +#include +#include +#include +#include +#include + +#include "cmh_sys.h" +#include "cmh_hcq_abi.h" +#include "cmh_txn.h" +#include "cmh_dma.h" +#include "cmh_pqc.h" + +#define LMS_VCQ_CMDS 3 /* header + cmd + flush */ + +struct cmh_lms_tfm_ctx { + u8 *pub_key; + u32 pub_key_len; + u32 lms_hss; /* 0 =3D LMS, 1 =3D LMS-HSS */ +}; + +static inline struct cmh_lms_tfm_ctx *cmh_lms_ctx(struct crypto_sig *tfm) +{ + return crypto_sig_ctx(tfm); +} + +/* + * LMS/HSS verify (synchronous sig_alg) + * + * @src: raw signature + * @slen: signature length + * @digest: message bytes + * @dlen: message length + * + * Returns 0 on successful verification, negative errno on failure. + */ +static int cmh_lms_verify(struct crypto_sig *tfm, + const void *src, unsigned int slen, + const void *digest, unsigned int dlen) +{ + struct cmh_lms_tfm_ctx *ctx =3D cmh_lms_ctx(tfm); + struct core_dispatch d =3D cmh_core_select_instance(CMH_CORE_HCQ); + struct vcq_cmd vcq[LMS_VCQ_CMDS]; + u8 *sig_buf =3D NULL, *m_buf =3D NULL, *pk_buf =3D NULL; + dma_addr_t sig_dma =3D DMA_MAPPING_ERROR; + dma_addr_t m_dma =3D DMA_MAPPING_ERROR; + dma_addr_t pk_dma =3D DMA_MAPPING_ERROR; + int ret; + + if (!ctx->pub_key) + return -EINVAL; + if (!slen || slen > LMS_MAX_SIG_LEN) + return -EINVAL; + if (!dlen || dlen > LMS_MAX_MSG_LEN) + return -EINVAL; + + sig_buf =3D kmemdup(src, slen, GFP_KERNEL); + m_buf =3D kmemdup(digest, dlen, GFP_KERNEL); + pk_buf =3D kmemdup(ctx->pub_key, ctx->pub_key_len, GFP_KERNEL); + if (!sig_buf || !m_buf || !pk_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + sig_dma =3D cmh_dma_map_single(sig_buf, slen, DMA_TO_DEVICE); + m_dma =3D cmh_dma_map_single(m_buf, dlen, DMA_TO_DEVICE); + pk_dma =3D cmh_dma_map_single(pk_buf, ctx->pub_key_len, DMA_TO_DEVICE); + + if (cmh_dma_map_error(sig_dma) || cmh_dma_map_error(m_dma) || + cmh_dma_map_error(pk_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + vcq_set_header(&vcq[0], LMS_VCQ_CMDS); + vcq_add_hcq_lms_verify(&vcq[1], d.core_id, ctx->lms_hss, + ctx->pub_key_len, slen, dlen, + pk_dma, sig_dma, m_dma); + vcq_add_hcq_flush(&vcq[2], d.core_id); + + /* LMS verify traverses Merkle hash chains -- inherently slow */ + ret =3D cmh_tm_submit_sync_tmo(vcq, LMS_VCQ_CMDS, 1, d.mbx_idx, + cmh_tm_slow_op_timeout_jiffies()); + +out_unmap: + if (!cmh_dma_map_error(pk_dma)) + cmh_dma_unmap_single(pk_dma, ctx->pub_key_len, DMA_TO_DEVICE); + if (!cmh_dma_map_error(m_dma)) + cmh_dma_unmap_single(m_dma, dlen, DMA_TO_DEVICE); + if (!cmh_dma_map_error(sig_dma)) + cmh_dma_unmap_single(sig_dma, slen, DMA_TO_DEVICE); + +out_free: + kfree(pk_buf); + kfree(m_buf); + kfree(sig_buf); + return ret; +} + +static int cmh_lms_set_pub_key(struct crypto_sig *tfm, + const void *key, unsigned int keylen) +{ + struct cmh_lms_tfm_ctx *ctx =3D cmh_lms_ctx(tfm); + + if (!keylen || keylen > LMS_MAX_PK_LEN) + return -EINVAL; + + kfree(ctx->pub_key); + ctx->pub_key =3D NULL; + ctx->pub_key_len =3D 0; + + ctx->pub_key =3D kmemdup(key, keylen, GFP_KERNEL); + if (!ctx->pub_key) + return -ENOMEM; + + ctx->pub_key_len =3D keylen; + return 0; +} + +static unsigned int cmh_lms_key_size(struct crypto_sig *tfm) +{ + struct cmh_lms_tfm_ctx *ctx =3D cmh_lms_ctx(tfm); + + return ctx->pub_key_len * 8; +} + +static int cmh_lms_init(struct crypto_sig *tfm) +{ + struct cmh_lms_tfm_ctx *ctx =3D cmh_lms_ctx(tfm); + + memset(ctx, 0, sizeof(*ctx)); + return 0; +} + +static int cmh_lms_hss_init(struct crypto_sig *tfm) +{ + struct cmh_lms_tfm_ctx *ctx =3D cmh_lms_ctx(tfm); + + memset(ctx, 0, sizeof(*ctx)); + ctx->lms_hss =3D 1; + return 0; +} + +static void cmh_lms_exit(struct crypto_sig *tfm) +{ + struct cmh_lms_tfm_ctx *ctx =3D cmh_lms_ctx(tfm); + + kfree(ctx->pub_key); + ctx->pub_key =3D NULL; +} + +static unsigned int cmh_lms_max_size(struct crypto_sig *tfm) +{ + /* Verify-only; report the max signature length for API parity. */ + return LMS_MAX_SIG_LEN; +} + +static struct sig_alg cmh_lms_algs[] =3D { + { + .verify =3D cmh_lms_verify, + .set_pub_key =3D cmh_lms_set_pub_key, + .key_size =3D cmh_lms_key_size, + .max_size =3D cmh_lms_max_size, + .init =3D cmh_lms_init, + .exit =3D cmh_lms_exit, + .base =3D { + .cra_name =3D "lms", + .cra_driver_name =3D "rambus-cmh-lms", + .cra_priority =3D 300, + .cra_module =3D THIS_MODULE, + .cra_ctxsize =3D sizeof(struct cmh_lms_tfm_ctx), + }, + }, + { + .verify =3D cmh_lms_verify, + .set_pub_key =3D cmh_lms_set_pub_key, + .key_size =3D cmh_lms_key_size, + .max_size =3D cmh_lms_max_size, + .init =3D cmh_lms_hss_init, + .exit =3D cmh_lms_exit, + .base =3D { + .cra_name =3D "lms-hss", + .cra_driver_name =3D "rambus-cmh-lms-hss", + .cra_priority =3D 300, + .cra_module =3D THIS_MODULE, + .cra_ctxsize =3D sizeof(struct cmh_lms_tfm_ctx), + }, + }, +}; + +/** + * cmh_pqc_lms_register() - Register LMS/LMS-HSS sig algorithms with the c= rypto framework + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_pqc_lms_register(void) +{ + int ret, i; + + if (!cmh_core_present(CMH_CORE_HCQ)) + return 0; + + for (i =3D 0; i < ARRAY_SIZE(cmh_lms_algs); i++) { + ret =3D crypto_register_sig(&cmh_lms_algs[i]); + if (ret) { + dev_err(cmh_dev(), "cmh: failed to register %s (%d)\n", + cmh_lms_algs[i].base.cra_name, ret); + goto err_unregister; + } + } + + return 0; + +err_unregister: + while (i--) + crypto_unregister_sig(&cmh_lms_algs[i]); + return ret; +} + +/** + * cmh_pqc_lms_unregister() - Unregister LMS/LMS-HSS sig algorithms from t= he crypto framework + */ +void cmh_pqc_lms_unregister(void) +{ + int i =3D ARRAY_SIZE(cmh_lms_algs); + + if (!cmh_core_present(CMH_CORE_HCQ)) + return; + + while (i--) + crypto_unregister_sig(&cmh_lms_algs[i]); +} diff --git a/drivers/crypto/cmh/cmh_pqc_slhdsa.c b/drivers/crypto/cmh/cmh_p= qc_slhdsa.c new file mode 100644 index 000000000000..1766871a2cc4 --- /dev/null +++ b/drivers/crypto/cmh/cmh_pqc_slhdsa.c @@ -0,0 +1,382 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- SLH-DSA Signature Driver (sig_alg, synchronous) + * + * Registers SLH-DSA sig algorithms for all 12 parameter sets + * (SHAKE/SHA2 x 128/192/256 x s/f) with sign and verify callbacks. + * + * Key format: + * Public key =3D raw pk bytes (2*n bytes) + * Private key =3D raw sk bytes (4*n) + * + * Sign: src =3D message, dst =3D raw signature + * Verify: src =3D raw signature, digest =3D message bytes + * + * Private keys are raw (written to SYS_REF_TEMP per-operation). + */ + +#include +#include +#include +#include +#include +#include + +#include "cmh_sys.h" +#include "cmh_qse_abi.h" +#include "cmh_hcq_abi.h" +#include "cmh_txn.h" +#include "cmh_dma.h" +#include "cmh_key.h" +#include "cmh_pqc.h" + +struct cmh_slhdsa_tfm_ctx { + struct cmh_key_ctx key; /* private key (raw sk bytes) */ + u8 *pub_key; + u32 pub_key_len; + u32 param_set; /* HCQ_SLHDSA_SHAKE_128S .. SHA2_256F */ +}; + +static inline struct cmh_slhdsa_tfm_ctx *cmh_slhdsa_ctx(struct crypto_sig = *tfm) +{ + return crypto_sig_ctx(tfm); +} + +/* + * SLH-DSA sign (synchronous sig_alg) + * + * @src: message bytes + * @slen: message length + * @dst: signature output buffer + * @dlen: output buffer length + * + * Returns signature length on success, negative errno on failure. + * Uses raw private keys written to SYS_REF_TEMP per-operation. + */ +static int cmh_slhdsa_sign(struct crypto_sig *tfm, + const void *src, unsigned int slen, + void *dst, unsigned int dlen) +{ + struct cmh_slhdsa_tfm_ctx *ctx =3D cmh_slhdsa_ctx(tfm); + u32 sig_sz =3D slhdsa_get_sig_size(ctx->param_set); + u32 sk_sz =3D slhdsa_sk_size(ctx->param_set); + struct vcq_cmd vcq[HCQ_VCQ_CMDS_MAX]; /* raw: hdr+write+sign+flush */ + struct core_dispatch d; + u32 vcq_count; + u8 *m_buf =3D NULL, *sig_buf =3D NULL, *sk_buf =3D NULL; + dma_addr_t m_dma =3D DMA_MAPPING_ERROR; + dma_addr_t sig_dma =3D DMA_MAPPING_ERROR; + dma_addr_t sk_dma =3D DMA_MAPPING_ERROR; + int ret, idx; + + if (ctx->key.mode =3D=3D CMH_KEY_NONE) + return -EINVAL; + if (dlen < sig_sz) + return -EINVAL; + if (!slen || slen > SLHDSA_MAX_MSG_LEN) + return -EINVAL; + + m_buf =3D kmemdup(src, slen, GFP_KERNEL); + sig_buf =3D kzalloc(sig_sz, GFP_KERNEL); + if (!m_buf || !sig_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + m_dma =3D cmh_dma_map_single(m_buf, slen, DMA_TO_DEVICE); + sig_dma =3D cmh_dma_map_single(sig_buf, sig_sz, DMA_FROM_DEVICE); + if (cmh_dma_map_error(m_dma) || cmh_dma_map_error(sig_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + sk_dma =3D DMA_MAPPING_ERROR; + idx =3D 0; + + d =3D cmh_core_select_instance(CMH_CORE_HCQ); + + if (ctx->key.raw.len !=3D sk_sz) { + ret =3D -EINVAL; + goto out_unmap; + } + sk_buf =3D kmemdup(ctx->key.raw.data, ctx->key.raw.len, + GFP_KERNEL); + if (!sk_buf) { + ret =3D -ENOMEM; + goto out_unmap; + } + sk_dma =3D cmh_dma_map_single(sk_buf, sk_sz, DMA_TO_DEVICE); + if (cmh_dma_map_error(sk_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + vcq_count =3D HCQ_VCQ_CMDS_MIN + 1; + vcq_set_header(&vcq[idx++], vcq_count); + vcq_add_sys_write(&vcq[idx++], SYS_REF_TEMP, sk_dma, + SYS_REF_NONE, sk_sz, + ctx->key.raw.sys_type); + vcq_add_hcq_slhdsa_sign_internal(&vcq[idx++], d.core_id, + ctx->param_set, + slen, 0, + m_dma, SYS_REF_TEMP, + sig_dma); + vcq_add_hcq_flush(&vcq[idx++], d.core_id); + + ret =3D cmh_tm_submit_sync_tmo(vcq, vcq_count, 1, d.mbx_idx, + cmh_tm_slow_op_timeout_jiffies()); + + if (!ret) { + /* Sync bounce buffer so CPU sees the DMA-written signature */ + cmh_dma_sync_for_cpu(sig_dma, sig_sz, DMA_FROM_DEVICE); + memcpy(dst, sig_buf, sig_sz); + ret =3D sig_sz; + } + +out_unmap: + if (sk_buf) { + if (!cmh_dma_map_error(sk_dma)) + cmh_dma_unmap_single(sk_dma, sk_sz, DMA_TO_DEVICE); + kfree_sensitive(sk_buf); + } + if (!cmh_dma_map_error(sig_dma)) + cmh_dma_unmap_single(sig_dma, sig_sz, DMA_FROM_DEVICE); + if (!cmh_dma_map_error(m_dma)) + cmh_dma_unmap_single(m_dma, slen, DMA_TO_DEVICE); + +out_free: + kfree(sig_buf); + kfree(m_buf); + return ret; +} + +/* + * SLH-DSA verify (synchronous sig_alg) + * + * @src: raw signature + * @slen: signature length + * @digest: message bytes + * @dlen: message length + * + * Returns 0 on successful verification, negative errno on failure. + */ +static int cmh_slhdsa_verify(struct crypto_sig *tfm, + const void *src, unsigned int slen, + const void *digest, unsigned int dlen) +{ + struct cmh_slhdsa_tfm_ctx *ctx =3D cmh_slhdsa_ctx(tfm); + u32 sig_sz =3D slhdsa_get_sig_size(ctx->param_set); + u32 pk_sz =3D slhdsa_pk_size(ctx->param_set); + struct core_dispatch d =3D cmh_core_select_instance(CMH_CORE_HCQ); + struct vcq_cmd vcq[HCQ_VCQ_CMDS_MIN]; + u8 *sig_buf =3D NULL, *m_buf =3D NULL, *pk_buf =3D NULL; + dma_addr_t sig_dma =3D DMA_MAPPING_ERROR; + dma_addr_t m_dma =3D DMA_MAPPING_ERROR; + dma_addr_t pk_dma =3D DMA_MAPPING_ERROR; + int ret; + + if (!ctx->pub_key) + return -EINVAL; + if (slen !=3D sig_sz) + return -EINVAL; + if (!dlen || dlen > SLHDSA_MAX_MSG_LEN) + return -EINVAL; + + sig_buf =3D kmemdup(src, slen, GFP_KERNEL); + m_buf =3D kmemdup(digest, dlen, GFP_KERNEL); + pk_buf =3D kmemdup(ctx->pub_key, pk_sz, GFP_KERNEL); + if (!sig_buf || !m_buf || !pk_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + sig_dma =3D cmh_dma_map_single(sig_buf, sig_sz, DMA_TO_DEVICE); + m_dma =3D cmh_dma_map_single(m_buf, dlen, DMA_TO_DEVICE); + pk_dma =3D cmh_dma_map_single(pk_buf, pk_sz, DMA_TO_DEVICE); + + if (cmh_dma_map_error(sig_dma) || cmh_dma_map_error(m_dma) || + cmh_dma_map_error(pk_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + vcq_set_header(&vcq[0], HCQ_VCQ_CMDS_MIN); + vcq_add_hcq_slhdsa_verify_internal(&vcq[1], d.core_id, ctx->param_set, + dlen, m_dma, pk_dma, sig_dma); + vcq_add_hcq_flush(&vcq[2], d.core_id); + + /* SLH-DSA verify recomputes hyper-tree hashes -- inherently slow */ + ret =3D cmh_tm_submit_sync_tmo(vcq, HCQ_VCQ_CMDS_MIN, 1, d.mbx_idx, + cmh_tm_slow_op_timeout_jiffies()); + +out_unmap: + if (!cmh_dma_map_error(pk_dma)) + cmh_dma_unmap_single(pk_dma, pk_sz, DMA_TO_DEVICE); + if (!cmh_dma_map_error(m_dma)) + cmh_dma_unmap_single(m_dma, dlen, DMA_TO_DEVICE); + if (!cmh_dma_map_error(sig_dma)) + cmh_dma_unmap_single(sig_dma, sig_sz, DMA_TO_DEVICE); + +out_free: + kfree(pk_buf); + kfree(m_buf); + kfree(sig_buf); + return ret; +} + +static int cmh_slhdsa_set_pub_key(struct crypto_sig *tfm, + const void *key, unsigned int keylen) +{ + struct cmh_slhdsa_tfm_ctx *ctx =3D cmh_slhdsa_ctx(tfm); + u32 expected =3D slhdsa_pk_size(ctx->param_set); + + if (keylen !=3D expected) + return -EINVAL; + + kfree(ctx->pub_key); + ctx->pub_key =3D NULL; + ctx->pub_key_len =3D 0; + + ctx->pub_key =3D kmemdup(key, keylen, GFP_KERNEL); + if (!ctx->pub_key) + return -ENOMEM; + + ctx->pub_key_len =3D keylen; + return 0; +} + +static int cmh_slhdsa_set_priv_key(struct crypto_sig *tfm, + const void *key, unsigned int keylen) +{ + struct cmh_slhdsa_tfm_ctx *ctx =3D cmh_slhdsa_ctx(tfm); + + /* Raw sk (4*n bytes) */ + if (keylen !=3D slhdsa_sk_size(ctx->param_set)) + return -EINVAL; + + return cmh_key_setkey_raw(&ctx->key, key, keylen, CORE_ID_HCQ); +} + +static unsigned int cmh_slhdsa_key_size(struct crypto_sig *tfm) +{ + struct cmh_slhdsa_tfm_ctx *ctx =3D cmh_slhdsa_ctx(tfm); + + /* crypto_sig_keysize() returns bits, not bytes */ + return slhdsa_pk_size(ctx->param_set) * 8; +} + +static unsigned int cmh_slhdsa_max_size(struct crypto_sig *tfm) +{ + struct cmh_slhdsa_tfm_ctx *ctx =3D cmh_slhdsa_ctx(tfm); + + return slhdsa_get_sig_size(ctx->param_set); +} + +static void cmh_slhdsa_exit(struct crypto_sig *tfm) +{ + struct cmh_slhdsa_tfm_ctx *ctx =3D cmh_slhdsa_ctx(tfm); + + cmh_key_destroy(&ctx->key); + kfree(ctx->pub_key); + ctx->pub_key =3D NULL; +} + +/* Generate init functions for all 12 parameter sets */ +#define SLHDSA_INIT(ps_val) \ + static int cmh_slhdsa_init_##ps_val(struct crypto_sig *tfm) \ + { \ + struct cmh_slhdsa_tfm_ctx *ctx =3D cmh_slhdsa_ctx(tfm); \ + memset(ctx, 0, sizeof(*ctx)); \ + ctx->param_set =3D ps_val; \ + return 0; \ + } + +SLHDSA_INIT(HCQ_SLHDSA_SHAKE_128S) +SLHDSA_INIT(HCQ_SLHDSA_SHAKE_128F) +SLHDSA_INIT(HCQ_SLHDSA_SHAKE_192S) +SLHDSA_INIT(HCQ_SLHDSA_SHAKE_192F) +SLHDSA_INIT(HCQ_SLHDSA_SHAKE_256S) +SLHDSA_INIT(HCQ_SLHDSA_SHAKE_256F) +SLHDSA_INIT(HCQ_SLHDSA_SHA2_128S) +SLHDSA_INIT(HCQ_SLHDSA_SHA2_128F) +SLHDSA_INIT(HCQ_SLHDSA_SHA2_192S) +SLHDSA_INIT(HCQ_SLHDSA_SHA2_192F) +SLHDSA_INIT(HCQ_SLHDSA_SHA2_256S) +SLHDSA_INIT(HCQ_SLHDSA_SHA2_256F) + +#define SLHDSA_ALG(name, drv, ps_val) { \ + .sign =3D cmh_slhdsa_sign, \ + .verify =3D cmh_slhdsa_verify, \ + .set_pub_key =3D cmh_slhdsa_set_pub_key, \ + .set_priv_key =3D cmh_slhdsa_set_priv_key, \ + .key_size =3D cmh_slhdsa_key_size, \ + .max_size =3D cmh_slhdsa_max_size, \ + .init =3D cmh_slhdsa_init_##ps_val, \ + .exit =3D cmh_slhdsa_exit, \ + .base =3D { \ + .cra_name =3D name, \ + .cra_driver_name =3D drv, \ + .cra_priority =3D 300, \ + .cra_module =3D THIS_MODULE, \ + .cra_ctxsize =3D sizeof(struct cmh_slhdsa_tfm_ctx), \ + }, \ + } + +static struct sig_alg cmh_slhdsa_algs[] =3D { + SLHDSA_ALG("slh-dsa-shake-128s", "rambus-cmh-slh-dsa-shake-128s", HCQ_SLH= DSA_SHAKE_128S), + SLHDSA_ALG("slh-dsa-shake-128f", "rambus-cmh-slh-dsa-shake-128f", HCQ_SLH= DSA_SHAKE_128F), + SLHDSA_ALG("slh-dsa-shake-192s", "rambus-cmh-slh-dsa-shake-192s", HCQ_SLH= DSA_SHAKE_192S), + SLHDSA_ALG("slh-dsa-shake-192f", "rambus-cmh-slh-dsa-shake-192f", HCQ_SLH= DSA_SHAKE_192F), + SLHDSA_ALG("slh-dsa-shake-256s", "rambus-cmh-slh-dsa-shake-256s", HCQ_SLH= DSA_SHAKE_256S), + SLHDSA_ALG("slh-dsa-shake-256f", "rambus-cmh-slh-dsa-shake-256f", HCQ_SLH= DSA_SHAKE_256F), + SLHDSA_ALG("slh-dsa-sha2-128s", "rambus-cmh-slh-dsa-sha2-128s", HCQ_SLH= DSA_SHA2_128S), + SLHDSA_ALG("slh-dsa-sha2-128f", "rambus-cmh-slh-dsa-sha2-128f", HCQ_SLH= DSA_SHA2_128F), + SLHDSA_ALG("slh-dsa-sha2-192s", "rambus-cmh-slh-dsa-sha2-192s", HCQ_SLH= DSA_SHA2_192S), + SLHDSA_ALG("slh-dsa-sha2-192f", "rambus-cmh-slh-dsa-sha2-192f", HCQ_SLH= DSA_SHA2_192F), + SLHDSA_ALG("slh-dsa-sha2-256s", "rambus-cmh-slh-dsa-sha2-256s", HCQ_SLH= DSA_SHA2_256S), + SLHDSA_ALG("slh-dsa-sha2-256f", "rambus-cmh-slh-dsa-sha2-256f", HCQ_SLH= DSA_SHA2_256F), +}; + +/** + * cmh_pqc_slhdsa_register() - Register SLH-DSA akcipher algorithms with t= he crypto framework + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_pqc_slhdsa_register(void) +{ + int ret, i; + + if (!cmh_core_present(CMH_CORE_HCQ)) + return 0; + + for (i =3D 0; i < ARRAY_SIZE(cmh_slhdsa_algs); i++) { + ret =3D crypto_register_sig(&cmh_slhdsa_algs[i]); + if (ret) { + dev_err(cmh_dev(), "cmh: failed to register %s (%d)\n", + cmh_slhdsa_algs[i].base.cra_name, ret); + goto err_unregister; + } + } + + return 0; + +err_unregister: + while (i--) + crypto_unregister_sig(&cmh_slhdsa_algs[i]); + return ret; +} + +/** + * cmh_pqc_slhdsa_unregister() - Unregister SLH-DSA akcipher algorithms fr= om the crypto framework + */ +void cmh_pqc_slhdsa_unregister(void) +{ + int i =3D ARRAY_SIZE(cmh_slhdsa_algs); + + if (!cmh_core_present(CMH_CORE_HCQ)) + return; + + while (i--) + crypto_unregister_sig(&cmh_slhdsa_algs[i]); +} diff --git a/drivers/crypto/cmh/cmh_pqc_xmss.c b/drivers/crypto/cmh/cmh_pqc= _xmss.c new file mode 100644 index 000000000000..4e6c464521d6 --- /dev/null +++ b/drivers/crypto/cmh/cmh_pqc_xmss.c @@ -0,0 +1,244 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Cryptography Research, Inc. (CRI). + * CMH LKM -- XMSS/XMSS-MT Signature Driver (verify-only, sig_alg, synchro= nous) + * + * Registers "xmss" and "xmss-mt" sig algorithms with verify-only + * callbacks. Sign is not supported (stateful signature -- key + * management must happen externally). + * + * Verify: src =3D raw signature, digest =3D message bytes + * Public key: raw pk bytes (variable length, set via set_pub_key) + */ + +#include +#include +#include +#include +#include + +#include "cmh_sys.h" +#include "cmh_hcq_abi.h" +#include "cmh_txn.h" +#include "cmh_dma.h" +#include "cmh_pqc.h" + +#define XMSS_VCQ_CMDS 3 /* header + cmd + flush */ + +struct cmh_xmss_tfm_ctx { + u8 *pub_key; + u32 pub_key_len; + u32 xmss_mt; /* 0 =3D XMSS, 1 =3D XMSS-MT */ +}; + +static inline struct cmh_xmss_tfm_ctx *cmh_xmss_ctx(struct crypto_sig *tfm) +{ + return crypto_sig_ctx(tfm); +} + +/* + * XMSS/XMSS-MT verify (synchronous sig_alg) + * + * @src: raw signature + * @slen: signature length + * @digest: message bytes + * @dlen: message length + * + * Returns 0 on successful verification, negative errno on failure. + */ +static int cmh_xmss_verify(struct crypto_sig *tfm, + const void *src, unsigned int slen, + const void *digest, unsigned int dlen) +{ + struct cmh_xmss_tfm_ctx *ctx =3D cmh_xmss_ctx(tfm); + struct core_dispatch d =3D cmh_core_select_instance(CMH_CORE_HCQ); + struct vcq_cmd vcq[XMSS_VCQ_CMDS]; + u8 *sig_buf =3D NULL, *m_buf =3D NULL, *pk_buf =3D NULL; + dma_addr_t sig_dma =3D DMA_MAPPING_ERROR; + dma_addr_t m_dma =3D DMA_MAPPING_ERROR; + dma_addr_t pk_dma =3D DMA_MAPPING_ERROR; + int ret; + + if (!ctx->pub_key) + return -EINVAL; + if (!slen || slen > XMSS_MAX_SIG_LEN) + return -EINVAL; + if (!dlen || dlen > XMSS_MAX_MSG_LEN) + return -EINVAL; + + sig_buf =3D kmemdup(src, slen, GFP_KERNEL); + m_buf =3D kmemdup(digest, dlen, GFP_KERNEL); + pk_buf =3D kmemdup(ctx->pub_key, ctx->pub_key_len, GFP_KERNEL); + if (!sig_buf || !m_buf || !pk_buf) { + ret =3D -ENOMEM; + goto out_free; + } + + sig_dma =3D cmh_dma_map_single(sig_buf, slen, DMA_TO_DEVICE); + m_dma =3D cmh_dma_map_single(m_buf, dlen, DMA_TO_DEVICE); + pk_dma =3D cmh_dma_map_single(pk_buf, ctx->pub_key_len, DMA_TO_DEVICE); + + if (cmh_dma_map_error(sig_dma) || cmh_dma_map_error(m_dma) || + cmh_dma_map_error(pk_dma)) { + ret =3D -ENOMEM; + goto out_unmap; + } + + vcq_set_header(&vcq[0], XMSS_VCQ_CMDS); + vcq_add_hcq_xmss_verify(&vcq[1], d.core_id, ctx->xmss_mt, + ctx->pub_key_len, slen, dlen, + pk_dma, sig_dma, m_dma); + vcq_add_hcq_flush(&vcq[2], d.core_id); + + /* XMSS verify traverses Merkle hash chains -- inherently slow */ + ret =3D cmh_tm_submit_sync_tmo(vcq, XMSS_VCQ_CMDS, 1, d.mbx_idx, + cmh_tm_slow_op_timeout_jiffies()); + +out_unmap: + if (!cmh_dma_map_error(pk_dma)) + cmh_dma_unmap_single(pk_dma, ctx->pub_key_len, DMA_TO_DEVICE); + if (!cmh_dma_map_error(m_dma)) + cmh_dma_unmap_single(m_dma, dlen, DMA_TO_DEVICE); + if (!cmh_dma_map_error(sig_dma)) + cmh_dma_unmap_single(sig_dma, slen, DMA_TO_DEVICE); + +out_free: + kfree(pk_buf); + kfree(m_buf); + kfree(sig_buf); + return ret; +} + +static int cmh_xmss_set_pub_key(struct crypto_sig *tfm, + const void *key, unsigned int keylen) +{ + struct cmh_xmss_tfm_ctx *ctx =3D cmh_xmss_ctx(tfm); + + if (!keylen || keylen > XMSS_MAX_PK_LEN) + return -EINVAL; + + kfree(ctx->pub_key); + ctx->pub_key =3D NULL; + ctx->pub_key_len =3D 0; + + ctx->pub_key =3D kmemdup(key, keylen, GFP_KERNEL); + if (!ctx->pub_key) + return -ENOMEM; + + ctx->pub_key_len =3D keylen; + return 0; +} + +static unsigned int cmh_xmss_key_size(struct crypto_sig *tfm) +{ + struct cmh_xmss_tfm_ctx *ctx =3D cmh_xmss_ctx(tfm); + + return ctx->pub_key_len * 8; +} + +static int cmh_xmss_init(struct crypto_sig *tfm) +{ + struct cmh_xmss_tfm_ctx *ctx =3D cmh_xmss_ctx(tfm); + + memset(ctx, 0, sizeof(*ctx)); + return 0; +} + +static int cmh_xmss_mt_init(struct crypto_sig *tfm) +{ + struct cmh_xmss_tfm_ctx *ctx =3D cmh_xmss_ctx(tfm); + + memset(ctx, 0, sizeof(*ctx)); + ctx->xmss_mt =3D 1; + return 0; +} + +static void cmh_xmss_exit(struct crypto_sig *tfm) +{ + struct cmh_xmss_tfm_ctx *ctx =3D cmh_xmss_ctx(tfm); + + kfree(ctx->pub_key); + ctx->pub_key =3D NULL; +} + +static unsigned int cmh_xmss_max_size(struct crypto_sig *tfm) +{ + /* Verify-only; report the max signature length for API parity. */ + return XMSS_MAX_SIG_LEN; +} + +static struct sig_alg cmh_xmss_algs[] =3D { + { + .verify =3D cmh_xmss_verify, + .set_pub_key =3D cmh_xmss_set_pub_key, + .key_size =3D cmh_xmss_key_size, + .max_size =3D cmh_xmss_max_size, + .init =3D cmh_xmss_init, + .exit =3D cmh_xmss_exit, + .base =3D { + .cra_name =3D "xmss", + .cra_driver_name =3D "rambus-cmh-xmss", + .cra_priority =3D 300, + .cra_module =3D THIS_MODULE, + .cra_ctxsize =3D sizeof(struct cmh_xmss_tfm_ctx), + }, + }, + { + .verify =3D cmh_xmss_verify, + .set_pub_key =3D cmh_xmss_set_pub_key, + .key_size =3D cmh_xmss_key_size, + .max_size =3D cmh_xmss_max_size, + .init =3D cmh_xmss_mt_init, + .exit =3D cmh_xmss_exit, + .base =3D { + .cra_name =3D "xmss-mt", + .cra_driver_name =3D "rambus-cmh-xmss-mt", + .cra_priority =3D 300, + .cra_module =3D THIS_MODULE, + .cra_ctxsize =3D sizeof(struct cmh_xmss_tfm_ctx), + }, + }, +}; + +/** + * cmh_pqc_xmss_register() - Register XMSS/XMSS-MT sig algorithms with the= crypto framework + * + * Return: 0 on success, negative errno on failure. + */ +int cmh_pqc_xmss_register(void) +{ + int ret, i; + + if (!cmh_core_present(CMH_CORE_HCQ)) + return 0; + + for (i =3D 0; i < ARRAY_SIZE(cmh_xmss_algs); i++) { + ret =3D crypto_register_sig(&cmh_xmss_algs[i]); + if (ret) { + dev_err(cmh_dev(), "cmh: failed to register %s (%d)\n", + cmh_xmss_algs[i].base.cra_name, ret); + goto err_unregister; + } + } + + return 0; + +err_unregister: + while (i--) + crypto_unregister_sig(&cmh_xmss_algs[i]); + return ret; +} + +/** + * cmh_pqc_xmss_unregister() - Unregister XMSS/XMSS-MT sig algorithms from= the crypto framework + */ +void cmh_pqc_xmss_unregister(void) +{ + int i =3D ARRAY_SIZE(cmh_xmss_algs); + + if (!cmh_core_present(CMH_CORE_HCQ)) + return; + + while (i--) + crypto_unregister_sig(&cmh_xmss_algs[i]); +} --=20 2.43.7 From nobody Fri Sep 25 01:20:34 2026 Received: from DM1PR04CU001.outbound.protection.outlook.com (mail-centralusazon11020120.outbound.protection.outlook.com [52.101.61.120]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 445CB4BF92E; Thu, 17 Sep 2026 22:59:52 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.61.120 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789686000; cv=fail; b=DiCM412XgJ0oVNADe+2yP68bneTrUAO16hu9Ww5DpaualXSJCSzlxni4huC3SAJIqnzJ16Yww/Qidol/Sdny5f/Oh/UjlsrCXU35M69pNz36aesWFYecJgeVq91d6cArbf7HEwUFmLuIY3/dh+yQUpKGnGfAf6HWIhbq4QiXVrQ= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789686000; c=relaxed/simple; bh=thUvqgeI0nznJk3tprE+iRjfoB5mJIFmRP9qlqAxgwc=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=mTg78PFTADikHctUD4pSBwv/So0g7GBDqqhz+5Eu4ma5g9VoTtPwFKVlGFmcg1zxb4CXXbIUyv3lKpcHKLGYnE0Tm2WeD88hNZhGaOphxyzt+6Rv811BKAaHLbwwPECOfa4vZ/nqiVp5g30rvai+E6ZdJdGvddLe+Dnqt5XWS0c= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com; spf=fail smtp.mailfrom=rambus.com; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b=ha/cf6Ls; arc=fail smtp.client-ip=52.101.61.120 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=rambus.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b="ha/cf6Ls" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=q+LCvpHnTi67QkPqxEPAgCbei5ZRnyg1Bjq09qKCq+ZZuJyTv3TK+PdGHRzLPUdD+hHs7ew8Oes+dJO0mpAkP2dK9QvxmHQ0/BtoXytsmTmwOwLFDTrfDzaSeSGseBlZNU70kFhcym0fcBbjlOpleNk9mFI2NgzTHhdbS9/Zl0y2VSSV0PPkzNDa0zJT4AZb/uc5TTLFKiNTMkpYTK8VGIVq0EB1yQdUx8meWNQ+FfD3reyCbt2TgdPPVACU3o1wIwiyBNS33X0ui+F4f5RqLGJqQ5pCj4b0DzK09kxdKRFHB3l8Y/YCsE6yhX7vb17kab6U2+liWjtorUQ3JnsVxw== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=zyaWpSQnjrYz1Ld2zGUVB1ax9Z9dF+S+WyW4plx64aI=; b=PKbdEiTFt1dn1RGbm6lgByghXRSXHR/YH0VS3QYjXrxflTnMLNZ318k4cgPCQQAQckX5qYyqZUlw+7OpzeMNoP/CGIE+k8c8m10cLDYg4vRxPawu7HXdvPrnPi5N1Anbatho65FigiX5Jfb9bu0tn0M+TtYE1CWQLHn5eb1Wd6KO2bd9GgTBY0HEVmPh4E8SUEnYoXgHr7HytflZGiCWH+nuF1g7SWs16Gi3SDFPhGlhFtIeHvJbMSc10wziRXAbMT1Jl13CFHz7EJBibqvoZ1HGBGCKfckiUJJrmcXFap/15WJcYfCwNESje8YuA7M52SfkwNjADKMXQrGkB+5cPw== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 192.86.86.210) smtp.rcpttodomain=cryptography.com smtp.mailfrom=rambus.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=rambus.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=rambus.com; s=selector1; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=zyaWpSQnjrYz1Ld2zGUVB1ax9Z9dF+S+WyW4plx64aI=; b=ha/cf6Ls1ZWALQSrPCw+7W+wUTMewcCa7d62xmqV87Kx3bv/SRyNtDRgSDRvAg1RPu8eYmctGSaDnDD8oJh0TEGyimQhNC8Es9O3j1/tXv2vp/TsF5nomDT9JbQWrz0YiFereyW24PjecrKgkegtICbk/KzxzCAnp7MyEKx4/UOiiUPokR4e0Q+qG3mbSaLoTcPQGFHC3FqOuxElCI2Wh+2f7ZgpiIm7vJw0MocDVKGp5bQQr6W7ML2AFqqFskOpmX/b7HMmAIEUPHWSeluMsyyuXkpBPs89eunoUlBbSkLkMgNLfULhGIi+URy5QrsmHfEkGWAm5rC/dTUUjiamQQ== Received: from MW4PR04CA0336.namprd04.prod.outlook.com (2603:10b6:303:8a::11) by MN2PR04MB6463.namprd04.prod.outlook.com (2603:10b6:208:1a3::15) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.428.12; Thu, 17 Sep 2026 22:59:35 +0000 Received: from SJ5PEPF000001CF.namprd05.prod.outlook.com (2603:10b6:303:8a:cafe::1b) by MW4PR04CA0336.outlook.office365.com (2603:10b6:303:8a::11) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.428.12 via Frontend Transport; Thu, 17 Sep 2026 22:59:35 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 192.86.86.210) smtp.mailfrom=rambus.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=rambus.com; Received-SPF: Pass (protection.outlook.com: domain of rambus.com designates 192.86.86.210 as permitted sender) receiver=protection.outlook.com; client-ip=192.86.86.210; helo=hqxsv-psmtppxy02.rambus.com; pr=C Received: from hqxsv-psmtppxy02.rambus.com (192.86.86.210) by SJ5PEPF000001CF.mail.protection.outlook.com (10.167.242.43) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.451.8 via Frontend Transport; Thu, 17 Sep 2026 22:59:34 +0000 Received: from hqxsv-cmdev3-aousherovitch.rambus.com (hqn-lb-int-float.rambus.com [10.12.20.20]) by hqxsv-psmtppxy02.rambus.com (Postfix) with ESMTP id 93E6F1801770; Thu, 17 Sep 2026 22:59:32 +0000 (UTC) From: Alex Ousherovitch To: Albert Ou , Alex Ousherovitch , Conor Dooley , "David S. Miller" , Herbert Xu , Jonathan Corbet , Krzysztof Kozlowski , Palmer Dabbelt , Paul Walmsley , Rob Herring , Saravanakrishnan Krishnamoorthy , Shuah Khan Cc: Alexandre Ghiti , devicetree@vger.kernel.org, Joel Wittenauer , linux-api@vger.kernel.org, linux-crypto@vger.kernel.org, linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-riscv@lists.infradead.org, Shuah Khan , Thi Nguyen Subject: [PATCH v5 17/19] Documentation: ioctl: add CMH ioctl documentation and register 'J' Date: Thu, 17 Sep 2026 15:59:26 -0700 Message-ID: <20260917225929.2494111-18-aousherovitch@rambus.com> X-Mailer: git-send-email 2.43.7 In-Reply-To: <20260917225929.2494111-1-aousherovitch@rambus.com> References: <20260917225929.2494111-1-aousherovitch@rambus.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: SJ5PEPF000001CF:EE_|MN2PR04MB6463:EE_ X-MS-Office365-Filtering-Correlation-Id: 11a5ae9b-3a96-4dff-743d-08df150f577a X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|7416014|36860700016|376014|82310400026|23010399003|1800799024|921020|18002099003|22082099003|6133799003|56012099006|11063799006|10067099003|3023799007; X-Microsoft-Antispam-Message-Info: njQzSuyF/bCFWHyQ5BNkV3+1TJbOmw4Q5cF6Oa5B7Qe7MuMbZzyPcbFZYPQRSSk7fwgqMWmIbE1FoyTVsiuDWgctomc+yrvWBNpAtuabtiO8Di9HFWuBFW7CR+nzpjL/47JtOEXcXglE/PPemu8fEQvtEOl7ode8QP4KzF35XQt4O8mn2/qvTpnBpgALpQP0B9O7WfMVTZm2iRGoWdo7L3VKdAzUlcNBHTqvD1khXLC/DhzQ33s3SOly1ogUndJAE7yDz+p9GgDowqGiQAz8upiM0Sqlc/qBKujC5ifwV94uGVlYTx/0dFLxLolNye3BV/4RFOTM7TnbKdFTyHTApxcLvkXtHow7vzaPcbgmxYRNRTc4kt0RGzdNq/BZXUkSsO1PF1mS9I9xswd2m51CFYlB3UJ17AYQNaErEFzpHdPeYgY1BotaLqpLBp5y8ntv8LLAwtyP2M8TFpIItAbfeyPlBnnD5o0zPU6SyIhtCj/Mm0Npu7/boMzVRpHzDHJcIVAnbwEkcLQZkkKqJN8jX/jBYIxtARfVSEL1qPL5BSXMLzhz/eVqGCqUyEuhKmM/3+yg0h/E3EyK7cjT/YOhmvmPYKtdokncpRY3FJ5rQJgTiL6Gqg6oLDPRX3FaL9Aut+sosFzmfsMigWUrYDbLJXP/dr3QwwYGUvpJUQJu7EskQ6GhG/YLjgNIHiacwKWUwkqMLvEW1ucqNQ1qj0RFJYJY4dVt8rQ0HyjMRxH/aWUH9yrl7z+yzxgsLWGNf8Lm X-Forefront-Antispam-Report: CIP:192.86.86.210;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:hqxsv-psmtppxy02.rambus.com;PTR:InfoDomainNonexistent;CAT:NONE;SFS:(13230040)(7416014)(36860700016)(376014)(82310400026)(23010399003)(1800799024)(921020)(18002099003)(22082099003)(6133799003)(56012099006)(11063799006)(10067099003)(3023799007);DIR:OUT;SFP:1102; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: bMYg/jFqc5ceuCPZ7idpIoAi0bdMrJcMsXQnEHDVa2VlYS5LrFEIdKcvg/vx3GnkojWIYtqmcZBgUXrzrEzNO0TOf3q6One5Rt7NLnrUFAq99wV8F5J6FB80XX+vZtsbTrllNKfP8MdTfMp011ycFB+zeayMcT8hlxSRwqpkMxBQ7SpobVeZej8f6Ibn5zRnag7Sj3wZrc35ECyfKZ0ocMwNHNDACeVrkdo6vqXE+yaTIJB8+QazCHRCjaRsLi8x5LTN4sd/SPQJF76a1YnIxI50++ZI89EZ/dIMZyowXHNPKTbX/rDMM1/G3fqz4I4TCOog+kb/pEggXqNcmTP9fbSEZAhdo+gTxhZasFooY6aHDlUEF5RyP9HrvSjdQftnxJdEIhCKkawraiCoWK5iMfzGrjXYlxzXkRjYRkWSf9nxbds/Z5F/q2kznoNQZ0QA X-OriginatorOrg: rambus.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 17 Sep 2026 22:59:34.5775 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: 11a5ae9b-3a96-4dff-743d-08df150f577a X-MS-Exchange-CrossTenant-Id: bd0ba799-c2b9-413c-9c56-5d1731c4827c X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=bd0ba799-c2b9-413c-9c56-5d1731c4827c;Ip=[192.86.86.210];Helo=[hqxsv-psmtppxy02.rambus.com] X-MS-Exchange-CrossTenant-AuthSource: SJ5PEPF000001CF.namprd05.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: MN2PR04MB6463 Content-Type: text/plain; charset="utf-8" Add Documentation/userspace-api/ioctl/cmh_mgmt.rst documenting the ioctl commands on the /dev/cmh_mgmt misc device for the Rambus CryptoManager Hub (CMH) hardware crypto accelerator driver. Covers key management, KIC key derivation, PKE (RSA, ECDSA, ECDH, EdDSA), PQC (ML-KEM, ML-DSA, SLH-DSA), SM2, EAC, and DRBG. Link the page into the userspace-api/ioctl index toctree. Register ioctl magic number 'J' (0x4A) in ioctl-number.rst. The driver uses ioctls 0x01-0x40. Signed-off-by: Alex Ousherovitch Co-developed-by: Saravanakrishnan Krishnamoorthy Signed-off-by: Saravanakrishnan Krishnamoorthy --- .../userspace-api/ioctl/cmh_mgmt.rst | 1292 +++++++++++++++++ Documentation/userspace-api/ioctl/index.rst | 1 + .../userspace-api/ioctl/ioctl-number.rst | 1 + 3 files changed, 1294 insertions(+) create mode 100644 Documentation/userspace-api/ioctl/cmh_mgmt.rst diff --git a/Documentation/userspace-api/ioctl/cmh_mgmt.rst b/Documentation= /userspace-api/ioctl/cmh_mgmt.rst new file mode 100644 index 000000000000..c02bac128787 --- /dev/null +++ b/Documentation/userspace-api/ioctl/cmh_mgmt.rst @@ -0,0 +1,1292 @@ +.. SPDX-License-Identifier: GPL-2.0 + +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D +CMH Key Management ioctl Interface (cmh_mgmt) +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +:Author: Cryptography Research, Inc. (CRI) +:Maintainer: linux-crypto@vger.kernel.org + +Introduction +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +The ``/dev/cmh_mgmt`` character device provides user-space access to key +management, key derivation, public-key, and post-quantum cryptographic +operations on the CryptoManager Hub (CMH) hardware accelerator. + +The device is created by the ``cmh`` kernel module as a ``misc_device``. +All operations are synchronous -- the ioctl blocks until the hardware +completes. Opening the device requires ``CAP_SYS_ADMIN``. + +All ioctl argument structures are versioned: user space sets the +``version`` field to ``CMH_MGMT_V1`` (currently 1). This allows the +driver to extend structures in the future without breaking the ABI. + +Data types and ioctl numbers are defined in +````. The ioctl type letter is ``'J'`` +(0x4A). + +Error Handling +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +Unless otherwise noted, all ioctls return 0 on success and a negative +errno on failure. Common error codes: + +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D +``EINVAL`` Invalid ``version`` field, unsupported parameter, or + out-of-range length. +``EFAULT`` Failed to copy data to/from user space. +``ENOMEM`` Kernel memory allocation failed. +``EIO`` Hardware returned an error (eSW command failure). +``ENOENT`` Key not found (``KEY_FIND``, ``KEY_LIST``). +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +Datastore Concepts +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +The CMH hardware maintains an embedded datastore managed by the eSW +firmware. Objects in the datastore are identified by a 64-bit reference +(``ref``) and optionally by a 64-bit Content ID (``cid``). + +Two storage classes exist: + +**Temporary (SYS_REF_TEMP)** + Lifetime is scoped to a single mailbox slot. The eSW firmware + reclaims the object when the slot is reused. Used for raw-key + provisioning via ``KEY_NEW`` + ``KEY_WRITE``. + +**Persistent (SYS_REF_PERSIST)** + Survives across mailbox slots. Requires explicit deletion via + ``KEY_DELETE``. Identified by CID; resolved to a per-mailbox ref + via ``KEY_FIND``. + +Mailbox Dispatch +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +All ``/dev/cmh_mgmt`` ioctls are submitted on a single management +mailbox. This is a structural requirement of the eSW datastore model, +not a tunable: + +* Datastore access control is **per-mailbox**. ``KEY_NEW`` grants the + creating mailbox read/write/execute access; other mailboxes have none + until granted. The returned 64-bit ``ref`` encodes a randomised + offset and does **not** carry the owning mailbox, so an operation that + receives only a ``ref`` (``KEY_GRANT``, ``KEY_READ``, ``KEY_DELETE``, + ``DS_EXPORT``) cannot itself determine which mailbox owns the object. + Using one fixed management mailbox guarantees that a key's create, + modify, grant, read and hardware-held-key compute steps all share the + mailbox that holds its access rights, without exposing mailbox + identity in the UABI. User space may still widen a key's access to + additional mailboxes via ``KEY_GRANT``. + +* The eSW ``SYS_REF_TEMP`` scratch store is per-mailbox and persists + across ioctl calls, so a multi-step flow that derives into + ``SYS_REF_TEMP`` (for example a ``KIC_*`` derivation) and later + consumes it (``DS_EXPORT`` with ``wrap_key =3D SYS_REF_TEMP``) requires + both calls to use the same mailbox. + +Per-mailbox ``rambus,cores`` device-tree affinity applies to the *stateles= s* +in-kernel crypto API path, which carries no datastore state between +requests and is balanced across mailboxes by the driver. + +Key Types +=3D=3D=3D=3D=3D=3D=3D=3D=3D + +The ``ds_type`` field in ``KEY_NEW`` and ``KEY_WRITE`` selects the +datastore object type. Values are defined as ``CMH_DS_*`` constants: + +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D +Constant Value Description +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D +``CMH_DS_RAW_VALUE`` 1 Raw byte array +``CMH_DS_AES_KEY`` 2 AES key (128/192/256-bit) +``CMH_DS_AES_XTS_KEY`` 3 AES-XTS key (256/512-bit) +``CMH_DS_HMAC_KEY`` 4 HMAC key +``CMH_DS_KMAC_KEY`` 5 KMAC key +``CMH_DS_SM4_KEY`` 6 SM4 key (128-bit) +``CMH_DS_CHACHA20_KEY`` 7 ChaCha20 key (256-bit) +``CMH_DS_RSA_PRIV_KEY`` 10 RSA private key +``CMH_DS_RSA_PUB_KEY`` 11 RSA public key +``CMH_DS_RSA_CRT_KEY`` 12 RSA CRT private key +``CMH_DS_ECDSA_PRIV_KEY`` 13 ECDSA private key +``CMH_DS_ECDSA_PUB_KEY`` 14 ECDSA public key +``CMH_DS_ECDH_PRIV_KEY`` 15 ECDH private key +``CMH_DS_EDDSA_PRIV_KEY`` 16 EdDSA private key +``CMH_DS_SHARED_SECRET`` 17 Shared secret +``CMH_DS_SM2_PRIV_KEY`` 18 SM2 private key +``CMH_DS_ML_KEM_DK`` 20 ML-KEM decapsulation key +``CMH_DS_ML_DSA_SK`` 21 ML-DSA secret key +``CMH_DS_SLHDSA_SK`` 25 SLH-DSA secret key +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +Key Flags +=3D=3D=3D=3D=3D=3D=3D=3D=3D + +The ``flags`` field in ``KEY_NEW`` and ``KEY_WRITE`` is a bitmask: + +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D +Flag Bit Description +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D +``CMH_FLAG_PT`` 16 Key can be read as plaintext +``CMH_FLAG_XC`` 17 Key can be exported over XC bus +``CMH_FLAG_SCA`` 18 SCA key stored in 2 shares +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +Elliptic Curve IDs +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +Curve identifiers for PKE operations (``curve`` field): + +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D =3D=3D=3D=3D=3D +Constant Value +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D =3D=3D=3D=3D=3D +``CMH_CURVE_P192`` 0x01 +``CMH_CURVE_P224`` 0x02 +``CMH_CURVE_P256`` 0x03 +``CMH_CURVE_P384`` 0x04 +``CMH_CURVE_P521`` 0x05 +``CMH_CURVE_SECP256K1`` 0x07 +``CMH_CURVE_BP192R1`` 0x11 +``CMH_CURVE_BP224R1`` 0x12 +``CMH_CURVE_BP256R1`` 0x13 +``CMH_CURVE_BP320R1`` 0x14 +``CMH_CURVE_BP384R1`` 0x15 +``CMH_CURVE_BP512R1`` 0x16 +``CMH_CURVE_SM2`` 0x18 +``CMH_CURVE_25519`` 0x21 +``CMH_CURVE_448`` 0x22 +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D =3D=3D=3D=3D=3D + +Key Management ioctls +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +CMH_IOCTL_KEY_NEW +----------------- + +Create a new empty datastore object. + +:Direction: ``_IOWR`` +:Number: 0x01 +:Argument: ``struct cmh_ioctl_key_new`` + +:: + + struct cmh_ioctl_key_new { + __u32 version; /* must be CMH_MGMT_V1 */ + __u32 ds_type; /* CMH_DS_* key type */ + __u32 len; /* key length in bytes */ + __u32 flags; /* CMH_FLAG_* */ + __u64 cid; /* caller ID (name) for the key */ + __u64 ref; /* [out] key reference */ + }; + +The returned ``ref`` is used in subsequent ``KEY_WRITE``, ``KEY_READ``, +and crypto operation ioctls. + +CMH_IOCTL_KEY_NEW_RANDOM +------------------------ + +Create a new datastore object filled with hardware-generated random data. + +:Direction: ``_IOWR`` +:Number: 0x0B +:Argument: ``struct cmh_ioctl_key_new`` + +Same structure as ``KEY_NEW``. The hardware DRBG fills the object with +``len`` random bytes. + +CMH_IOCTL_KEY_WRITE +------------------- + +Write key material into a previously created datastore object. + +:Direction: ``_IOW`` +:Number: 0x02 +:Argument: ``struct cmh_ioctl_key_write`` + +:: + + struct cmh_ioctl_key_write { + __u32 version; + __u32 len; /* key data length */ + __u32 ds_type; /* CMH_DS_* key type */ + __u32 flags; /* CMH_FLAG_* */ + __u64 ref; /* key reference from KEY_NEW */ + __u64 wrap_key; /* wrapping key ref (CMH_REF_NONE =3D plaintext) = */ + __u64 data; /* user-space pointer to key material */ + }; + +If ``wrap_key`` is ``CMH_REF_NONE`` (0), key material is written in +plaintext. Otherwise, the data is unwrapped using the specified +wrapping key. + +CMH_IOCTL_KEY_READ +------------------ + +Read key material from a datastore object. + +:Direction: ``_IOWR`` +:Number: 0x03 +:Argument: ``struct cmh_ioctl_key_read`` + +:: + + struct cmh_ioctl_key_read { + __u32 version; + __u32 len; /* buffer length */ + __u64 ref; /* key reference */ + __u64 wrap_key; /* wrapping key ref (CMH_REF_NONE =3D plaintext) = */ + __u64 data; /* user-space pointer to output buffer */ + __u32 out_len; /* [out] actual bytes written */ + __u32 __reserved; + }; + +Plaintext reads require the ``CMH_FLAG_PT`` attribute on the key. +The eSW prepends a 16-byte header (``CMH_SYS_WRAP_HDR_SIZE``) even +for plaintext reads; the output buffer must accommodate this. The +output overhead is ``CMH_DS_EXPORT_OVERHEAD_PLAIN`` (16 bytes) for +plaintext reads and ``CMH_DS_EXPORT_OVERHEAD_WRAPPED`` (48 bytes: +16-byte header + 16-byte nonce + 16-byte tag) for wrapped reads. + +CMH_IOCTL_KEY_FIND +------------------ + +Resolve a Content ID to a datastore reference. + +:Direction: ``_IOWR`` +:Number: 0x04 +:Argument: ``struct cmh_ioctl_key_find`` + +:: + + struct cmh_ioctl_key_find { + __u32 version; + __u32 __reserved; + __u64 cid; /* caller ID to search for */ + __u64 ref; /* [out] resolved key reference */ + __u32 len; /* [out] key length */ + __u32 type; /* [out] key type */ + }; + +Returns ``-ENOENT`` if no object with the given CID exists. + +CMH_IOCTL_KEY_LIST +------------------ + +Iterate datastore objects. + +:Direction: ``_IOWR`` +:Number: 0x0E +:Argument: ``struct cmh_ioctl_key_list`` + +:: + + struct cmh_ioctl_key_list { + __u32 version; + __u32 __reserved; + __u64 start_ref; /* starting DS reference (0 =3D first) */ + __u64 ref; /* [out] object reference */ + __u64 cid; /* [out] caller ID */ + __u32 len; /* [out] object length */ + __u32 type; /* [out] object type */ + }; + +Pass ``start_ref=3D0`` to begin from the first object. On return, pass +the returned ``ref`` as ``start_ref`` in the next call. Iteration ends +when ``ref =3D=3D 0``. + +CMH_IOCTL_KEY_GRANT +------------------- + +Set per-mailbox access permissions on a datastore object. + +:Direction: ``_IOW`` +:Number: 0x05 +:Argument: ``struct cmh_ioctl_key_grant`` + +:: + + struct cmh_ioctl_key_grant { + __u32 version; + __u32 __reserved; + __u64 ref; /* key reference */ + __u64 read; /* per-MBX read permission bitfield */ + __u64 write; /* per-MBX write permission bitfield */ + __u64 execute; /* per-MBX execute permission bitfield */ + }; + +CMH_IOCTL_KEY_DELETE +-------------------- + +Delete a datastore object (persistent keys only). + +:Direction: ``_IOW`` +:Number: 0x06 +:Argument: ``struct cmh_ioctl_key_grant`` + +Uses the same structure as ``KEY_GRANT``; only the ``ref`` field is +used. + +Datastore Export/Import ioctls +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D + +CMH_IOCTL_DS_EXPORT +------------------- + +Export the entire datastore as an encrypted blob. + +:Direction: ``_IOWR`` +:Number: 0x07 +:Argument: ``struct cmh_ioctl_ds_export`` + +:: + + struct cmh_ioctl_ds_export { + __u32 version; + __u32 len; /* buffer length */ + __u64 cid; /* caller ID for response tagging */ + __u64 wrap_key; /* wrapping key ref (CMH_REF_NONE =3D plaintext) = */ + __u64 data; /* user-space pointer to output buffer */ + __u32 out_len; /* [out] actual bytes written */ + __u32 __reserved; + }; + +CMH_IOCTL_DS_IMPORT +------------------- + +Import a previously exported datastore blob. + +:Direction: ``_IOW`` +:Number: 0x08 +:Argument: ``struct cmh_ioctl_ds_import`` + +:: + + struct cmh_ioctl_ds_import { + __u32 version; + __u32 len; /* blob length */ + __u64 wrap_key; /* wrapping key ref (CMH_REF_NONE =3D plaintext) = */ + __u64 data; /* user-space pointer to import blob */ + }; + +Key Derivation ioctls (KIC) +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D + +The Key Initialization Core (KIC) provides hardware key derivation from +OTP-provisioned base keys. Up to 8 base keys are available +(``CMH_KIC_KEY1`` through ``CMH_KIC_KEY8``). + +CMH_IOCTL_KIC_HKDF1 +-------------------- + +HKDF-based key derivation (single-step, label only). + +:Direction: ``_IOWR`` +:Number: 0x09 +:Argument: ``struct cmh_ioctl_kic_hkdf1`` + +:: + + struct cmh_ioctl_kic_hkdf1 { + __u32 version; + __u32 key_len; /* output key length */ + __u64 base_key; /* KIC base key reference */ + __u64 cid; /* CID for the new DS entry */ + __u64 label; /* user-space pointer to label data */ + __u32 label_len; /* label length in bytes */ + __u32 flags; /* CMH_KIC_FLAG_* */ + __u64 ref; /* [out] derived key reference */ + }; + +If ``CMH_KIC_FLAG_TEMP`` is set, the result is stored in the temporary +datastore (not persistent). + +CMH_IOCTL_KIC_HKDF2 +-------------------- + +HKDF-based key derivation (two-step, with salt key). + +:Direction: ``_IOWR`` +:Number: 0x0A +:Argument: ``struct cmh_ioctl_kic_hkdf2`` + +:: + + struct cmh_ioctl_kic_hkdf2 { + __u32 version; + __u32 key_len; + __u64 base_key; + __u64 salt_key; /* salt key reference (CMH_REF_NONE =3D no salt) = */ + __u64 cid; + __u64 label; + __u32 label_len; + __u32 flags; + __u64 ref; /* [out] derived key reference */ + }; + +CMH_IOCTL_KIC_AES_CMAC_KDF +--------------------------- + +AES-CMAC-based key derivation (NIST SP 800-108). + +:Direction: ``_IOWR`` +:Number: 0x0C +:Argument: ``struct cmh_ioctl_kic_aes_cmac_kdf`` + +:: + + struct cmh_ioctl_kic_aes_cmac_kdf { + __u32 version; + __u32 key_len; /* base & output key length (must be 32) */ + __u64 base_key; + __u64 cid; + __u64 label; + __u32 label_len; + __u32 flags; + __u64 ref; /* [out] derived key reference */ + }; + +CMH_IOCTL_KIC_DKEK_DERIVE +-------------------------- + +Derive a Device Key Encryption Key (DKEK) for secure key export. + +:Direction: ``_IOWR`` +:Number: 0x0D +:Argument: ``struct cmh_ioctl_kic_dkek_derive`` + +:: + + struct cmh_ioctl_kic_dkek_derive { + __u32 version; + __u32 host_id; /* target host ID (0 =3D caller's own) */ + __u64 base_key; + __u64 cid; + __u64 metadata; /* user-space pointer to metadata */ + __u32 metadata_len; + __u32 flags; + __u64 ref; /* [out] derived KEK reference */ + }; + +PKE (Public Key Engine) ioctls +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D + +RSA Operations +-------------- + +CMH_IOCTL_PKE_RSA_ENC +~~~~~~~~~~~~~~~~~~~~~~ + +RSA public-key encryption. + +:Direction: ``_IOWR`` +:Number: 0x10 +:Argument: ``struct cmh_ioctl_pke_rsa_enc`` + +:: + + struct cmh_ioctl_pke_rsa_enc { + __u32 version; + __u32 bits; /* RSA key size in bits (512-4096) */ + __u64 e; /* user-space pointer to public exponent */ + __u32 e_len; /* exponent length in bytes */ + __u32 __reserved; + __u64 n; /* user-space pointer to modulus */ + __u64 input; /* user-space pointer to input data */ + __u64 output; /* user-space pointer to output buffer */ + }; + +The public key (e, n) is passed as raw user-space buffers. + +CMH_IOCTL_PKE_RSA_DEC +~~~~~~~~~~~~~~~~~~~~~~ + +RSA private-key decryption using a datastore key reference. + +:Direction: ``_IOWR`` +:Number: 0x11 +:Argument: ``struct cmh_ioctl_pke_rsa_dec`` + +:: + + struct cmh_ioctl_pke_rsa_dec { + __u32 version; + __u32 bits; + __u64 e; /* public exponent */ + __u32 e_len; + __u32 __reserved; + __u64 n; /* modulus */ + __u64 input; /* ciphertext */ + __u64 output; /* plaintext output */ + __u64 key_ref; /* private key DS reference */ + }; + +CMH_IOCTL_PKE_RSA_CRT_DEC +~~~~~~~~~~~~~~~~~~~~~~~~~~ + +RSA CRT private-key decryption (faster, uses CRT key format). + +:Direction: ``_IOWR`` +:Number: 0x12 +:Argument: ``struct cmh_ioctl_pke_rsa_crt_dec`` + +:: + + struct cmh_ioctl_pke_rsa_crt_dec { + __u32 version; + __u32 bits; + __u64 e; + __u32 e_len; + __u32 __reserved; + __u64 n; + __u64 input; + __u64 output; + __u64 crt_ref; /* CRT key DS reference */ + }; + +CMH_IOCTL_PKE_RSA_KEYGEN +~~~~~~~~~~~~~~~~~~~~~~~~~ + +Generate an RSA key pair in hardware. + +:Direction: ``_IOWR`` +:Number: 0x13 +:Argument: ``struct cmh_ioctl_pke_rsa_keygen`` + +:: + + struct cmh_ioctl_pke_rsa_keygen { + __u32 version; + __u32 bits; /* key size in bits */ + __u64 e; /* user-space pointer to public exponent */ + __u32 e_len; + __u32 flags; /* CMH_FLAG_* */ + __u64 n; /* [out] user-space pointer to modulus buffer */ + __u64 d_cid; /* CID for private key DS entry */ + __u64 d_ref; /* [out] private key reference */ + __u64 crt_cid; /* CID for CRT key DS entry (0 =3D skip CRT) */ + __u64 crt_ref; /* [out] CRT key reference */ + }; + +Returns private key and optional CRT key as datastore references. +The modulus is written back to user space. + +ECDSA Operations +---------------- + +CMH_IOCTL_PKE_ECDSA_SIGN +~~~~~~~~~~~~~~~~~~~~~~~~~ + +ECDSA signature generation using a datastore private key. + +:Direction: ``_IOWR`` +:Number: 0x14 +:Argument: ``struct cmh_ioctl_pke_ecdsa_sign`` + +:: + + struct cmh_ioctl_pke_ecdsa_sign { + __u32 version; + __u32 curve; /* ABI curve ID (e.g. 0x03 =3D P-256) */ + __u64 digest; /* user-space pointer to hash digest */ + __u32 digest_len; /* digest length in bytes */ + __u32 __reserved; + __u64 signature; /* [out] user-space pointer to (r,s) */ + __u64 key_ref; /* private key DS reference */ + }; + +CMH_IOCTL_PKE_ECDH +~~~~~~~~~~~~~~~~~~~ + +Compute ECDH shared secret from a peer public key and a datastore +private key. + +:Direction: ``_IOWR`` +:Number: 0x16 +:Argument: ``struct cmh_ioctl_pke_ecdh`` + +:: + + struct cmh_ioctl_pke_ecdh { + __u32 version; + __u32 curve; + __u64 peer_key_x; /* user-space pointer to peer public key X */ + __u64 key_ref; /* private key DS reference */ + __u32 flags; /* reserved, must be 0 */ + __u32 __reserved; + __u64 __reserved2; /* reserved, must be 0 */ + __u64 output; /* [out] raw shared secret */ + }; + +CMH_IOCTL_PKE_ECDH_KEYGEN +~~~~~~~~~~~~~~~~~~~~~~~~~~ + +Derive a public key from a datastore private key. + +:Direction: ``_IOWR`` +:Number: 0x17 +:Argument: ``struct cmh_ioctl_pke_ecdh_keygen`` + +:: + + struct cmh_ioctl_pke_ecdh_keygen { + __u32 version; + __u32 curve; + __u64 key_ref; /* private key DS reference */ + __u64 public_key_x; /* [out] user-space pointer to public key X */ + }; + +EdDSA Operations +---------------- + +CMH_IOCTL_PKE_EDDSA_SIGN +~~~~~~~~~~~~~~~~~~~~~~~~~ + +EdDSA (Ed25519/Ed448) signature generation. + +:Direction: ``_IOWR`` +:Number: 0x18 +:Argument: ``struct cmh_ioctl_pke_eddsa_sign`` + +:: + + struct cmh_ioctl_pke_eddsa_sign { + __u32 version; + __u32 curve; /* CURVE_25519 or CURVE_448 */ + __u64 digest; /* user-space ptr to message (not digest) */ + __u32 digest_len; + __u32 __reserved; + __u64 signature; /* [out] user-space pointer to signature */ + __u64 key_ref; /* private key DS reference */ + }; + +Note: the ``digest`` field is the full message (pure EdDSA), not a +pre-computed hash. + +CMH_IOCTL_PKE_EDDSA_VERIFY +~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +EdDSA signature verification. + +:Direction: ``_IOW`` +:Number: 0x19 +:Argument: ``struct cmh_ioctl_pke_eddsa_verify`` + +:: + + struct cmh_ioctl_pke_eddsa_verify { + __u32 version; + __u32 curve; + __u64 digest; + __u32 digest_len; + __u32 __reserved; + __u64 signature; + __u64 public_key_y; /* user-space pointer to public key Y */ + }; + +EC Key Management +----------------- + +CMH_IOCTL_PKE_EC_KEYGEN +~~~~~~~~~~~~~~~~~~~~~~~~ + +Generate an EC private key in the hardware datastore. + +:Direction: ``_IOWR`` +:Number: 0x1A +:Argument: ``struct cmh_ioctl_pke_ec_keygen`` + +:: + + struct cmh_ioctl_pke_ec_keygen { + __u32 version; + __u32 curve; + __u32 flags; /* CMH_FLAG_* */ + __u32 __reserved; + __u64 cid; /* CID for the new key DS entry */ + __u64 ref; /* [out] private key reference */ + }; + +CMH_IOCTL_PKE_EC_PUBGEN +~~~~~~~~~~~~~~~~~~~~~~~~ + +Derive the public key from a datastore private key. + +:Direction: ``_IOWR`` +:Number: 0x1B +:Argument: ``struct cmh_ioctl_pke_ec_pubgen`` + +:: + + struct cmh_ioctl_pke_ec_pubgen { + __u32 version; + __u32 curve; + __u64 key_ref; /* private key DS reference */ + __u64 public_key; /* [out] user-space pointer to public key */ + }; + +CMH_IOCTL_PKE_EDDSA_KEYGEN_SCA +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +Generate a 2-share SCA-protected Ed448 private key. + +:Direction: ``_IOWR`` +:Number: 0x1C +:Argument: ``struct cmh_ioctl_pke_eddsa_keygen_sca`` + +:: + + struct cmh_ioctl_pke_eddsa_keygen_sca { + __u32 version; + __u32 curve; /* must be CURVE_448 */ + __u64 key_ref; /* input: normal Ed448 private key DS ref */ + __u64 cid; /* CID for the new SCA key DS entry */ + __u64 sca_ref; /* [out] SCA private key reference */ + }; + +Post-Quantum Cryptography (PQC) ioctls +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +PQC operations support the following flags in the ``flags`` field: + +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D =3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D +Flag Bit Description +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D =3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D +``CMH_QSE_FLAG_MASKED`` 0 Use masked (SCA-resistant) HW path +``CMH_QSE_FLAG_DS_REF`` 1 Store key output in DS, return ref +``CMH_QSE_FLAG_HW_RNG`` 2 Use HW RNG for seed/randomness +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D =3D=3D=3D=3D =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +ML-KEM (FIPS 203) +----------------- + +CMH_IOCTL_ML_KEM_KEYGEN +~~~~~~~~~~~~~~~~~~~~~~~~ + +Generate an ML-KEM key pair. + +:Direction: ``_IOWR`` +:Number: 0x20 +:Argument: ``struct cmh_ioctl_ml_kem_keygen`` + +:: + + struct cmh_ioctl_ml_kem_keygen { + __u32 version; + __u32 k; /* security parameter: 2/3/4 */ + __u32 flags; /* CMH_QSE_FLAG_* */ + __u32 __reserved; + __u64 seed; /* user-space pointer to seed (or 0 for HW RNG) */ + __u64 z; /* user-space pointer to z (or 0 for HW RNG) */ + __u64 ek; /* [out] user-space pointer to encapsulation key = */ + __u64 dk; /* [out] user-space pointer to decapsulation key, + * or [out] DS ref if CMH_QSE_FLAG_DS_REF */ + __u64 dk_cid; /* CID for DS entry (if DS_REF) */ + __u64 dk_ref; /* [out] dk DS reference (if DS_REF) */ + }; + +Security parameter ``k`` selects the strength: 2 (ML-KEM-512), +3 (ML-KEM-768), or 4 (ML-KEM-1024). + +CMH_IOCTL_ML_KEM_ENC +~~~~~~~~~~~~~~~~~~~~~ + +ML-KEM encapsulation. Produces ciphertext and shared secret. + +:Direction: ``_IOWR`` +:Number: 0x21 +:Argument: ``struct cmh_ioctl_ml_kem_enc`` + +:: + + struct cmh_ioctl_ml_kem_enc { + __u32 version; + __u32 k; + __u32 flags; /* CMH_QSE_FLAG_* */ + __u32 __reserved; + __u64 coin; /* user-space pointer to random coin (or 0) */ + __u64 ek; /* user-space pointer to encapsulation key */ + __u64 ct; /* [out] user-space pointer to ciphertext */ + __u64 ss; /* [out] user-space pointer to shared secret */ + __u64 __reserved2[2]; /* reserved for future use */ + }; + +CMH_IOCTL_ML_KEM_DEC +~~~~~~~~~~~~~~~~~~~~~ + +ML-KEM decapsulation. Recovers shared secret from ciphertext. + +:Direction: ``_IOWR`` +:Number: 0x22 +:Argument: ``struct cmh_ioctl_ml_kem_dec`` + +:: + + struct cmh_ioctl_ml_kem_dec { + __u32 version; + __u32 k; + __u32 flags; /* CMH_QSE_FLAG_* */ + __u32 __reserved; + __u64 ct; /* user-space pointer to ciphertext */ + __u64 dk; /* user-space pointer to dk or DS ref */ + __u64 ss; /* [out] user-space pointer to shared secret */ + __u64 __reserved2[2]; /* reserved for future use */ + }; + +ML-DSA (FIPS 204) +----------------- + +CMH_IOCTL_ML_DSA_KEYGEN +~~~~~~~~~~~~~~~~~~~~~~~~ + +Generate an ML-DSA key pair. + +:Direction: ``_IOWR`` +:Number: 0x23 +:Argument: ``struct cmh_ioctl_ml_dsa_keygen`` + +:: + + struct cmh_ioctl_ml_dsa_keygen { + __u32 version; + __u32 mode; /* security parameter: 2/3/5 */ + __u32 flags; /* CMH_QSE_FLAG_* */ + __u32 __reserved; + __u64 seed; /* user-space pointer to seed (or 0 for HW RNG) */ + __u64 pk; /* [out] user-space pointer to public key */ + __u64 sk; /* [out] user-space pointer to secret key, + * or [out] DS ref if CMH_QSE_FLAG_DS_REF */ + __u64 sk_cid; /* CID for DS entry (if DS_REF) */ + __u64 sk_ref; /* [out] sk DS reference (if DS_REF) */ + }; + +Security parameter ``mode`` selects the strength: 2 (ML-DSA-44), +3 (ML-DSA-65), or 5 (ML-DSA-87). + +.. note:: + + When ``CMH_QSE_FLAG_DS_REF`` keeps the secret key in the datastore, + the public key returned in ``pk`` is the only copy: there is no + operation to derive the public key from the secret-key reference + for ML-DSA. User space must persist ``pk`` at keygen time. + +CMH_IOCTL_ML_DSA_SIGN +~~~~~~~~~~~~~~~~~~~~~~ + +ML-DSA signature generation. + +:Direction: ``_IOWR`` +:Number: 0x24 +:Argument: ``struct cmh_ioctl_ml_dsa_sign`` + +:: + + struct cmh_ioctl_ml_dsa_sign { + __u32 version; + __u32 mode; + __u32 flags; /* CMH_QSE_FLAG_* */ + __u32 mlen; /* message length in bytes */ + __u64 m; /* user-space pointer to message */ + __u64 sk; /* user-space pointer to sk or DS ref */ + __u64 sig; /* [out] user-space pointer to signature */ + __u64 rnd; /* user-space pointer to randomness (or 0) */ + }; + +If ``mlen`` is set to ``CMH_ML_DSA_MLEN_EXTERNAL_MU`` (0xFFFFFFFF), +the ``m`` pointer is interpreted as a 64-byte pre-hashed mu value +(ExternalMu mode). + +CMH_IOCTL_SLHDSA_KEYGEN +~~~~~~~~~~~~~~~~~~~~~~~~ + +Generate an SLH-DSA key pair. + +:Direction: ``_IOWR`` +:Number: 0x28 +:Argument: ``struct cmh_ioctl_slhdsa_keygen`` + +:: + + struct cmh_ioctl_slhdsa_keygen { + __u32 version; + __u32 parameter_set; /* HCQ_SLHDSA_SHAKE_128S .. SHA2_256F */ + __u32 flags; /* CMH_QSE_FLAG_DS_REF */ + __u32 __reserved; + __u64 seed; /* user-space pointer to seed */ + __u64 pk; /* [out] user-space pointer to public key */ + __u64 sk; /* [out] user-space pointer to secret key, + * or [out] DS ref if CMH_QSE_FLAG_DS_RE= F */ + __u64 sk_cid; /* CID for DS entry (if DS_REF) */ + __u64 sk_ref; /* [out] sk DS reference (if DS_REF) */ + }; + +.. note:: + + When ``CMH_QSE_FLAG_DS_REF`` keeps the secret key in the datastore, + the public key returned in ``pk`` is the only copy: there is no + operation to derive the public key from the secret-key reference + for SLH-DSA. User space must persist ``pk`` at keygen time. + +CMH_IOCTL_SLHDSA_SIGN +~~~~~~~~~~~~~~~~~~~~~~ + +SLH-DSA signature generation (pure mode). + +:Direction: ``_IOWR`` +:Number: 0x29 +:Argument: ``struct cmh_ioctl_slhdsa_sign`` + +:: + + struct cmh_ioctl_slhdsa_sign { + __u32 version; + __u32 parameter_set; + __u32 msg_len; + __u32 ctx_len; + __u64 msg; /* user-space pointer to message */ + __u64 ctx; /* user-space pointer to context (or 0) */ + __u64 sk; /* DS ref for secret key */ + __u64 sig; /* [out] user-space pointer to signature */ + __u64 add_random; /* user-space pointer to addl. randomness (or 0) = */ + }; + +CMH_IOCTL_SLHDSA_SIGN_PREHASH +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +SLH-DSA pre-hash signature generation. + +:Direction: ``_IOWR`` +:Number: 0x2D +:Argument: ``struct cmh_ioctl_slhdsa_sign_prehash`` + +:: + + struct cmh_ioctl_slhdsa_sign_prehash { + __u32 version; + __u32 parameter_set; + __u32 prehash_algo; /* CMH_SLHDSA_PREHASH_* */ + __u32 digest; /* 0 =3D raw msg (eSW hashes), 1 =3D pre-comput= ed */ + __u32 msg_len; + __u32 ctx_len; + __u64 msg; /* user-space pointer to message/digest */ + __u64 ctx; /* user-space pointer to context (or 0) */ + __u64 sk; /* DS ref for secret key */ + __u64 sig; /* [out] user-space pointer to signature */ + __u64 add_random; /* user-space pointer to addl. randomness (or 0= ) */ + }; + +The ``prehash_algo`` field selects the hash algorithm +(``CMH_SLHDSA_PREHASH_SHA256``, etc.). + +CMH_IOCTL_SM2_ENC_POINT +~~~~~~~~~~~~~~~~~~~~~~~~ + +:Direction: ``_IOWR`` +:Number: 0x33 +:Argument: ``struct cmh_ioctl_sm2_enc_point`` + +:: + + struct cmh_ioctl_sm2_enc_point { + __u32 version; + __u32 nonce_len; /* 0 =3D HW generates, 32 =3D caller provides */ + __u64 nonce; /* user-space pointer to nonce (or 0) */ + __u64 public_key; /* user-space pointer to public key (64B) */ + __u64 ciphertext; /* [out] user-space pointer to C1 (64B) */ + __u64 enc_point; /* [out] user-space pointer to enc point (64B) */ + }; + +CMH_IOCTL_SM2_ENC_HASH +~~~~~~~~~~~~~~~~~~~~~~~ + +:Direction: ``_IOWR`` +:Number: 0x37 +:Argument: ``struct cmh_ioctl_sm2_enc_hash`` + +:: + + struct cmh_ioctl_sm2_enc_hash { + __u32 version; + __u32 message_len; /* message length (1..32) */ + __u64 message; /* user-space pointer to plaintext */ + __u64 enc_point; /* user-space pointer to enc point (64B) */ + __u64 ciphertext; /* [out] user-space pointer to ciphertext */ + }; + +CMH_IOCTL_SM2_DEC_POINT +~~~~~~~~~~~~~~~~~~~~~~~~ + +:Direction: ``_IOWR`` +:Number: 0x32 +:Argument: ``struct cmh_ioctl_sm2_dec_point`` + +:: + + struct cmh_ioctl_sm2_dec_point { + __u32 version; + __u32 ciphertext_len; /* total ciphertext length (97..128) */ + __u64 ciphertext; /* user-space pointer to ciphertext (64B: C1)= */ + __u64 dec_point; /* [out] user-space pointer to dec point (64B= ) */ + __u64 key_ref; /* private key DS reference */ + }; + +CMH_IOCTL_SM2_DEC_HASH +~~~~~~~~~~~~~~~~~~~~~~~ + +:Direction: ``_IOWR`` +:Number: 0x36 +:Argument: ``struct cmh_ioctl_sm2_dec_hash`` + +:: + + struct cmh_ioctl_sm2_dec_hash { + __u32 version; + __u32 ciphertext_len; /* ciphertext length (97..128) */ + __u64 ciphertext; /* user-space pointer to full ciphertext */ + __u64 dec_point; /* user-space pointer to dec point (64B) */ + __u64 plaintext; /* [out] user-space pointer to plaintext */ + }; + +SM2 Key Exchange (GM/T 0003.3) +------------------------------ + +The key exchange protocol is a multi-step flow: + +1. ``EC_KEYGEN(CMH_CURVE_SM2)`` -- generate a long-lived private key. +2. ``EC_PUBGEN`` -- derive the public key. +3. ``SM2_ID_DIGEST`` -- compute the SM3 identity digest (ZA). +4. ``SM2_ECDH_KEYGEN`` -- generate an ephemeral session key. +5. Exchange session keys with the peer. +6. ``SM2_ECDH`` -- compute the shared point. +7. ``SM2_ECDH_HASH`` -- derive the shared key from the shared point + and both parties' ZA digests. + +CMH_IOCTL_SM2_ECDH_KEYGEN +~~~~~~~~~~~~~~~~~~~~~~~~~~ + +:Direction: ``_IOWR`` +:Number: 0x30 +:Argument: ``struct cmh_ioctl_sm2_ecdh_keygen`` + +:: + + struct cmh_ioctl_sm2_ecdh_keygen { + __u32 version; + __u32 nonce_len; /* 0 =3D HW generates r (written back), 32 =3D c= aller */ + __u64 nonce; /* [in/out] user-space pointer to nonce buffer (= 32B) */ + __u64 session_key; /* [out] user-space pointer to R=3Dr*G (64B) */ + }; + +``nonce_len`` must be 0 or 32. If ``nonce_len=3D0``, the hardware +generates the ephemeral scalar and writes it back to the ``nonce`` +buffer. + +CMH_IOCTL_SM2_ECDH +~~~~~~~~~~~~~~~~~~~ + +:Direction: ``_IOWR`` +:Number: 0x31 +:Argument: ``struct cmh_ioctl_sm2_ecdh`` + +:: + + struct cmh_ioctl_sm2_ecdh { + __u32 version; + __u32 nonce_len; /* 0 =3D HW generates, 32 =3D caller provid= es */ + __u64 nonce; /* [in/out] user-space pointer to nonce r (= 32B) */ + __u64 peer_public_key; /* user-space pointer to peer pub key (64B)= */ + __u64 peer_session_key; /* user-space pointer to peer session key (= 64B) */ + __u64 key_ref; /* private key DS reference */ + __u64 shared_point; /* [out] user-space pointer to shared point= (64B) */ + __u64 shared_point_ref; /* [in/out] 0 =3D read-back; &ref =3D keep = DS */ + }; + +If ``shared_point_ref`` points to a non-zero value, the shared point +is kept in the datastore for use by ``SM2_ECDH_HASH``. + +CMH_IOCTL_SM2_ID_DIGEST +~~~~~~~~~~~~~~~~~~~~~~~~ + +Compute the SM3 identity digest (ZA) for a public key and identity +string. + +:Direction: ``_IOWR`` +:Number: 0x34 +:Argument: ``struct cmh_ioctl_sm2_id_digest`` + +:: + + struct cmh_ioctl_sm2_id_digest { + __u32 version; + __u32 id_len; /* identity length in bytes (<=3D32) */ + __u64 id; /* user-space pointer to identity string */ + __u64 public_key; /* user-space pointer to public key (64B) */ + __u64 digest; /* [out] user-space pointer to ZA digest (32B) */ + }; + +CMH_IOCTL_SM2_ECDH_HASH +~~~~~~~~~~~~~~~~~~~~~~~~ + +Derive the shared key from the shared point and ZA digests. + +:Direction: ``_IOWR`` +:Number: 0x35 +:Argument: ``struct cmh_ioctl_sm2_ecdh_hash`` + +:: + + struct cmh_ioctl_sm2_ecdh_hash { + __u32 version; + __u32 __reserved; + __u64 peer_id_digest; /* ptr to Z_A -- initiator's digest (32B) */ + __u64 id_digest; /* ptr to Z_B -- responder's digest (32B) */ + __u64 shared_point_ref; /* DS reference from SM2_ECDH */ + __u64 shared_key; /* [out] ptr to shared key (16B) */ + }; + +.. important:: + + The digest fields use **absolute** ordering per GM/T 0003.3, not + relative own/peer ordering. Both parties must pass: + + - ``peer_id_digest`` =3D Z_A (initiator's digest) -- hashed first + - ``id_digest`` =3D Z_B (responder's digest) -- hashed second + +Hardware Management ioctls +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D + +CMH_IOCTL_EAC_READ +------------------- + +Read and clear the hardware Error and Alarm Controller registers. + +:Direction: ``_IOWR`` +:Number: 0x0F +:Argument: ``struct cmh_ioctl_eac_read`` + +:: + + struct cmh_ioctl_eac_read { + __u32 version; + __u32 __reserved; + __u64 mailbox_notification; + __u32 hw_error; + __u32 hw_nmi; + __u32 hw_panic; + __u32 safety_fatal; + __u32 safety_notification; + __u32 sw_info0; + __u32 sw_info1; + __u32 sram_bank_errors[4]; + __u32 __pad; + }; + +The eSW atomically reads and clears the registers on each call. +Successive reads show only new events since the last read. + +CMH_IOCTL_DRBG_CONFIG +---------------------- + +Configure the hardware DRBG before first use. + +:Direction: ``_IOW`` +:Number: 0x40 +:Argument: ``struct cmh_ioctl_drbg_config`` + +:: + + struct cmh_ioctl_drbg_config { + __u32 version; + __u32 entropy_ratio; /* CMH_DRBG_RATIO_* */ + __u32 security_strength; /* CMH_DRBG_STRENGTH_* */ + __u32 __reserved; + }; + +This is a management operation normally performed once at system +startup. Must be called before any ``hwrng`` reads or DRBG generate +operations. + +ioctl Number Summary +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D =3D=3D=3D=3D =3D=3D=3D=3D =3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D +ioctl Dir Seq Argument +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D =3D=3D=3D=3D =3D=3D=3D=3D =3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D +``CMH_IOCTL_KEY_NEW`` IOWR 0x01 ``cmh_ioctl_key_new`` +``CMH_IOCTL_KEY_WRITE`` IOW 0x02 ``cmh_ioctl_key_write`` +``CMH_IOCTL_KEY_READ`` IOWR 0x03 ``cmh_ioctl_key_read`` +``CMH_IOCTL_KEY_FIND`` IOWR 0x04 ``cmh_ioctl_key_find`` +``CMH_IOCTL_KEY_GRANT`` IOW 0x05 ``cmh_ioctl_key_grant`` +``CMH_IOCTL_KEY_DELETE`` IOW 0x06 ``cmh_ioctl_key_grant`` +``CMH_IOCTL_DS_EXPORT`` IOWR 0x07 ``cmh_ioctl_ds_export`` +``CMH_IOCTL_DS_IMPORT`` IOW 0x08 ``cmh_ioctl_ds_import`` +``CMH_IOCTL_KIC_HKDF1`` IOWR 0x09 ``cmh_ioctl_kic_hkdf1`` +``CMH_IOCTL_KIC_HKDF2`` IOWR 0x0A ``cmh_ioctl_kic_hkdf2`` +``CMH_IOCTL_KEY_NEW_RANDOM`` IOWR 0x0B ``cmh_ioctl_key_new`` +``CMH_IOCTL_KIC_AES_CMAC_KDF`` IOWR 0x0C ``cmh_ioctl_kic_aes_cm= ac_kdf`` +``CMH_IOCTL_KIC_DKEK_DERIVE`` IOWR 0x0D ``cmh_ioctl_kic_dkek_d= erive`` +``CMH_IOCTL_KEY_LIST`` IOWR 0x0E ``cmh_ioctl_key_list`` +``CMH_IOCTL_EAC_READ`` IOWR 0x0F ``cmh_ioctl_eac_read`` +``CMH_IOCTL_PKE_RSA_ENC`` IOWR 0x10 ``cmh_ioctl_pke_rsa_en= c`` +``CMH_IOCTL_PKE_RSA_DEC`` IOWR 0x11 ``cmh_ioctl_pke_rsa_de= c`` +``CMH_IOCTL_PKE_RSA_CRT_DEC`` IOWR 0x12 ``cmh_ioctl_pke_rsa_cr= t_dec`` +``CMH_IOCTL_PKE_RSA_KEYGEN`` IOWR 0x13 ``cmh_ioctl_pke_rsa_ke= ygen`` +``CMH_IOCTL_PKE_ECDSA_SIGN`` IOWR 0x14 ``cmh_ioctl_pke_ecdsa_= sign`` +``CMH_IOCTL_PKE_ECDH`` IOWR 0x16 ``cmh_ioctl_pke_ecdh`` +``CMH_IOCTL_PKE_ECDH_KEYGEN`` IOWR 0x17 ``cmh_ioctl_pke_ecdh_k= eygen`` +``CMH_IOCTL_PKE_EDDSA_SIGN`` IOWR 0x18 ``cmh_ioctl_pke_eddsa_= sign`` +``CMH_IOCTL_PKE_EDDSA_VERIFY`` IOW 0x19 ``cmh_ioctl_pke_eddsa_= verify`` +``CMH_IOCTL_PKE_EC_KEYGEN`` IOWR 0x1A ``cmh_ioctl_pke_ec_key= gen`` +``CMH_IOCTL_PKE_EC_PUBGEN`` IOWR 0x1B ``cmh_ioctl_pke_ec_pub= gen`` +``CMH_IOCTL_PKE_EDDSA_KEYGEN_SCA`` IOWR 0x1C ``cmh_ioctl_pke_eddsa_= keygen_sca`` +``CMH_IOCTL_ML_KEM_KEYGEN`` IOWR 0x20 ``cmh_ioctl_ml_kem_key= gen`` +``CMH_IOCTL_ML_KEM_ENC`` IOWR 0x21 ``cmh_ioctl_ml_kem_enc= `` +``CMH_IOCTL_ML_KEM_DEC`` IOWR 0x22 ``cmh_ioctl_ml_kem_dec= `` +``CMH_IOCTL_ML_DSA_KEYGEN`` IOWR 0x23 ``cmh_ioctl_ml_dsa_key= gen`` +``CMH_IOCTL_ML_DSA_SIGN`` IOWR 0x24 ``cmh_ioctl_ml_dsa_sig= n`` +``CMH_IOCTL_SLHDSA_KEYGEN`` IOWR 0x28 ``cmh_ioctl_slhdsa_key= gen`` +``CMH_IOCTL_SLHDSA_SIGN`` IOWR 0x29 ``cmh_ioctl_slhdsa_sig= n`` +``CMH_IOCTL_SLHDSA_SIGN_PREHASH`` IOWR 0x2D ``cmh_ioctl_slhdsa_sig= n_prehash`` +``CMH_IOCTL_SM2_ECDH_KEYGEN`` IOWR 0x30 ``cmh_ioctl_sm2_ecdh_k= eygen`` +``CMH_IOCTL_SM2_ECDH`` IOWR 0x31 ``cmh_ioctl_sm2_ecdh`` +``CMH_IOCTL_SM2_DEC_POINT`` IOWR 0x32 ``cmh_ioctl_sm2_dec_po= int`` +``CMH_IOCTL_SM2_ENC_POINT`` IOWR 0x33 ``cmh_ioctl_sm2_enc_po= int`` +``CMH_IOCTL_SM2_ID_DIGEST`` IOWR 0x34 ``cmh_ioctl_sm2_id_dig= est`` +``CMH_IOCTL_SM2_ECDH_HASH`` IOWR 0x35 ``cmh_ioctl_sm2_ecdh_h= ash`` +``CMH_IOCTL_SM2_DEC_HASH`` IOWR 0x36 ``cmh_ioctl_sm2_dec_ha= sh`` +``CMH_IOCTL_SM2_ENC_HASH`` IOWR 0x37 ``cmh_ioctl_sm2_enc_ha= sh`` +``CMH_IOCTL_DRBG_CONFIG`` IOW 0x40 ``cmh_ioctl_drbg_confi= g`` +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D =3D=3D=3D=3D =3D=3D=3D=3D =3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +Relationship to the in-kernel crypto API +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +The main reason these operations are exposed as ioctls, rather than +through the standard in-kernel crypto API, is the CMH datastore key +model: an ioctl can operate on a *datastore-referenced* (hardware-held) +key, identified only by a ``ref`` or CID, whose raw bytes the CPU never +sees. The standard crypto API cannot express this -- every +``.setkey()`` takes raw key material -- so hardware key lifecycle +(create, import, derive, grant, destroy) and compute-on-hardware-held-key +operations have no crypto API equivalent and are only reachable here. + +These ioctls remain supported. Where an operation can *also* be +expressed through a standard kernel abstraction, additional in-kernel +crypto API bindings may be added over time; when such a binding lands the +driver registers through it, and the ioctl continues to be maintained for +backward compatibility: + +- **EdDSA** (``CMH_IOCTL_PKE_EDDSA_*``): a kernel ``sig`` binding may be + added once ed25519/ed448 algorithm types are accepted upstream. + +- **ML-KEM** (``CMH_IOCTL_ML_KEM_*``): a kernel KEM binding may be added + once the in-flight KEM subsystem series lands. + +- **Key lifecycle** (``CMH_IOCTL_KEY_*``): integration with the kernel + KEYS subsystem (trusted-keys / encrypted-keys) may be evaluated as a + follow-up series. + +Operations that are inherently vendor-specific (EAC Chip Authentication, +KIC key derivation, SM2 key exchange, DRBG configuration, datastore +export/import) have no corresponding kernel abstraction and are expected +to remain ioctl-only. diff --git a/Documentation/userspace-api/ioctl/index.rst b/Documentation/us= erspace-api/ioctl/index.rst index 475675eae086..bf88bb6b9a6f 100644 --- a/Documentation/userspace-api/ioctl/index.rst +++ b/Documentation/userspace-api/ioctl/index.rst @@ -12,4 +12,5 @@ IOCTLs ioctl-decoding =20 cdrom + cmh_mgmt hdio diff --git a/Documentation/userspace-api/ioctl/ioctl-number.rst b/Documenta= tion/userspace-api/ioctl/ioctl-number.rst index 2fc53093752d..25f020c7e367 100644 --- a/Documentation/userspace-api/ioctl/ioctl-number.rst +++ b/Documentation/userspace-api/ioctl/ioctl-number.rst @@ -170,6 +170,7 @@ Code Seq# Include File = Comments 'I' all linux/isdn.h con= flict! 'I' 00-0F drivers/isdn/divert/isdn_divert.h con= flict! 'I' 40-4F linux/mISDNif.h con= flict! +'J' 01-40 uapi/linux/cmh_mgmt_ioctl.h Ram= bus CryptoManager Hub (CMH) 'K' all linux/kd.h 'L' 00-1F linux/loop.h con= flict! 'L' 10-1F drivers/scsi/mpt3sas/mpt3sas_ctl.h con= flict! --=20 2.43.7 From nobody Fri Sep 25 01:20:34 2026 Received: from DM5PR21CU001.outbound.protection.outlook.com (mail-centralusazon11021115.outbound.protection.outlook.com [52.101.62.115]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D1B484B4864; Thu, 17 Sep 2026 22:59:47 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.62.115 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789685994; cv=fail; b=FxOBARCS6rprGc2A9bE/yb3I4ShjulVBrbx3i/8lJYUwQp2DGEfvd2rLKmzWQBn4HTwiazWY80t87nZMCk6ENcU0AkbGD5Htsu6DUrRxU0jHS2Yjmjm2q+9xwhasbCmkuxl1RvHeyo396XOPLLupJ8ob7T6lyWrJcDydD340678= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789685994; c=relaxed/simple; bh=zmm8d7v6SuOSeaEZ2OQrsdvrrSc+3yTFtZpg5jDML8w=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=A8dt/kiBCjdD5kC/GfC+UtTEi4rf3bhHHb/ammFSGKmdw4qCAni8vNO4Fo+HaATzu7uqA3ENHiOWUZlH7KbL+32NY0d+WGWREpwM88Ai6ao8b9QLvwbUwTmpJJ0/D1ofn1mo2O8nzvAo6I3whFJhqJD0Kc0PlDPzp/WAkTVTyhU= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com; spf=fail smtp.mailfrom=rambus.com; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b=Z7mepCfO; arc=fail smtp.client-ip=52.101.62.115 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=rambus.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b="Z7mepCfO" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=fgf6e23A1xYebIq0zHSiSbiQz4dOp9hBRNC+OreS2w0c8kNq97ArvMIUaqBgtqlSZJ5ABG2d95d/mHzYMLhRhlnXX47mIC714k+aKxxxDuUNtWG7g41ym5nlFaB/jiNwWqDAvyOzQV1M3FmjrOFSU18Xt7CcCzPq/BpRSYCcL50did8S5ymCj3KOyYYVHW39snFFb3AOty4xBOMH0eqHXe4I+VQ4q1ry8T/Lm5UCFCqRwmiKowF+3zZnOtsV9pmpyfEzAn4xWcTxziviApmU5KH315gGZTFlwbD19cffQYYYPbVom/PnnFSm8uk3cihQ8sw+IdI380HPbk+InsB5RA== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=kcFfijZeV2dayzaTQy1HddPt3o0QadNbDV9JFToJgbg=; b=A94/z5TRqzWnDzqWCCRcMngUJ29EfVpxAstgEYLz9Gmpo1dR2b5h8KjzmMYr+bwFF4dj2O3c63gi6qBp4UmIUeJ8OPijcvwkX+mm+oeWFGZkyjoGtAi+8JjBkbaCLwiIn0n95BviymDLpF7ZcRl6bDqYv03RfjiSqbmwYwQFasyGve0WCe9aRlWPoBLo+41Znh5tVXmfDqDO1uYqVSPqkM73Ueuk5q8Bb3VbRVgbLzN8WfsGYxmPk8pMwQEDkE5g7SvPViQY4L2yuRZJPkoiVet3lAbrCh0CLDVfvEGv0/46eKPr0pNxTjQFKyEsKY7d/nzTl6vpRpo+cDMf5xvPQg== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 192.86.86.210) smtp.rcpttodomain=cryptography.com smtp.mailfrom=rambus.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=rambus.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=rambus.com; s=selector1; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=kcFfijZeV2dayzaTQy1HddPt3o0QadNbDV9JFToJgbg=; b=Z7mepCfOJV34+zQ8wfOSb4u+Tc90oQdS1ghg6A7ui9OW+QzmZ6+f5CIv9c2fBuH7kKQI5dyCkAZ/ROE6t5S4I9yHd0etxmSWCu+fMMAyMXJcQFIMvLOAbzG9QBUVHtC2FH1U7/fhc5Yfmq5E15a3KVYR7AK8xMBCD7JFdssIE0OexGMVyRHyiw33vLlZSMv3bwFhb/t67um+sauMGadO3N6oR0N0T+h/0EIEl5BSj5PNAj3cdPvAQmD9HkjtPMNyUbAUOdwWF1XSupL6JvCaX7J2Cw84rCYDl80RmYQrKaQIVG0ePGjhpN2f+8fvda3p7apvR35o8L41X8j7ZLmX9w== Received: from BY5PR04CA0027.namprd04.prod.outlook.com (2603:10b6:a03:1d0::37) by DM6PR04MB6778.namprd04.prod.outlook.com (2603:10b6:5:24c::15) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.406.14; Thu, 17 Sep 2026 22:59:35 +0000 Received: from SJ5PEPF000001C8.namprd05.prod.outlook.com (2603:10b6:a03:1d0:cafe::5e) by BY5PR04CA0027.outlook.office365.com (2603:10b6:a03:1d0::37) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.428.11 via Frontend Transport; Thu, 17 Sep 2026 22:59:34 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 192.86.86.210) smtp.mailfrom=rambus.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=rambus.com; Received-SPF: Pass (protection.outlook.com: domain of rambus.com designates 192.86.86.210 as permitted sender) receiver=protection.outlook.com; client-ip=192.86.86.210; helo=hqxsv-psmtppxy02.rambus.com; pr=C Received: from hqxsv-psmtppxy02.rambus.com (192.86.86.210) by SJ5PEPF000001C8.mail.protection.outlook.com (10.167.242.36) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.451.8 via Frontend Transport; Thu, 17 Sep 2026 22:59:34 +0000 Received: from hqxsv-cmdev3-aousherovitch.rambus.com (hqn-lb-int-float.rambus.com [10.12.20.20]) by hqxsv-psmtppxy02.rambus.com (Postfix) with ESMTP id A045E1801771; Thu, 17 Sep 2026 22:59:32 +0000 (UTC) From: Alex Ousherovitch To: Albert Ou , Alex Ousherovitch , Conor Dooley , "David S. Miller" , Herbert Xu , Jonathan Corbet , Krzysztof Kozlowski , Palmer Dabbelt , Paul Walmsley , Rob Herring , Saravanakrishnan Krishnamoorthy , Shuah Khan Cc: Alexandre Ghiti , devicetree@vger.kernel.org, Joel Wittenauer , linux-api@vger.kernel.org, linux-crypto@vger.kernel.org, linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-riscv@lists.infradead.org, Shuah Khan , Thi Nguyen Subject: [PATCH v5 18/19] selftests: crypto: cmh - add kselftest for management ioctl Date: Thu, 17 Sep 2026 15:59:27 -0700 Message-ID: <20260917225929.2494111-19-aousherovitch@rambus.com> X-Mailer: git-send-email 2.43.7 In-Reply-To: <20260917225929.2494111-1-aousherovitch@rambus.com> References: <20260917225929.2494111-1-aousherovitch@rambus.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: SJ5PEPF000001C8:EE_|DM6PR04MB6778:EE_ X-MS-Office365-Filtering-Correlation-Id: 2df82056-0781-49b1-61ad-08df150f5779 X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|36860700016|7416014|23010399003|82310400026|376014|1800799024|6133799003|921020|18002099003|22082099003|11063799006|56012099006|10067099003|3023799007; X-Microsoft-Antispam-Message-Info: DlognZQEuK92YC0QLvHriAGiVtnBoCSwSWOfbUS0Yf9RiFnhRjLsIeW9DPDFUuCUNKZ04WgY2vIFfCiR4W7EBjiWSRivOWJta1Me/zB0azGGkfOCi59FGadmgkSdFtJ3+75eSpxqsj+XJZA6+uUab3V5aAzqIfzhQipi7oHwOD1fi6yPgBD+C9PqcDcGfw5Q1ZEW4CM70JaBrJBXbNRnkE8gKHInEGlIVH158McdF2h+OxAFWBsRoAOxg1QziEp19qrf9l6hGS9fJ1cIVIsKccLugelaCF+3FPS9Ls21DtfHL8Q1JGv1RVYRXE4qU+B3SARmCwkTa9jdjxFLahXeeLBMS8L3mmaJW9ilcr9CYdZ6MM9TALNVAkeRqI3T0nPPh+mXAhepGkjJu+4M1pFFogLOl0XsL91z6buWMHd7oQ+W2gnlbnKryVcyw6JTRJcpUNN1jJAMZwwefdi60iS4HBEsDLRKMGVy4Z4m1VxaDBP4taMyW36aHtZi3f4hbV2TYE3AdgJigNh/fGO+nDxWvJAgIhruLlvaeTh8E1Qb34DrmDVzs7AjBiYSOzUK7Vs4C73/e0ZJujYmrPN2teQEyn1QRU62G4rXUTyPAEn80gmXpGOB6nMEfXAKHf+A8NpTdwCtqs4wvdsY4B1myVwqboPS/3Ri9qOAFSFLjWA4u2G0imVWRD7Hb3cA1bGJhLKJL/f2i72qjSpCCfhapqXSizASL1uocclRScmlzv26+N5a5FJ25BO7iw4ivYL44g0o X-Forefront-Antispam-Report: CIP:192.86.86.210;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:hqxsv-psmtppxy02.rambus.com;PTR:InfoDomainNonexistent;CAT:NONE;SFS:(13230040)(36860700016)(7416014)(23010399003)(82310400026)(376014)(1800799024)(6133799003)(921020)(18002099003)(22082099003)(11063799006)(56012099006)(10067099003)(3023799007);DIR:OUT;SFP:1102; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: IzQ+yxG1aQsgyugswv3/T1l7YW9hPu/iKqV76J2OboUZ5mS/rfE7KBBZ3O5XOQjz25wKCMoMVk8TQ+7I5N/N/Jd5cYzs1WAoWQ7AHG71my5b7Xu0exnRkocMheY9M3gmt/MOHEHXqezAc82UonXa5yQEdbbYMS+TrF6pUOd7svZnoNs19653BECcxZlTNRPOLY+vzJ3hcQVy0Y6zQIg6vMuRVXnmBDBj18bhLf7umEfRPVdrQ2WkzTmlFp+4KO6owJIxi4cgVd4OPSxme+HIvlDd9Y1osaYzcacnGOfvRRGsjdHyT3Azc3RUHQai9X/WjO10HBedYiUkRhd3L5rKsgpmmY+uyrccG7XUK2Q5ZKtznVvVpMysdAorlzi3RLy4AIManyTBWkmmDnvBtWt1GEN93KE5760xxII6KBTkgPrB/rNYbMRF/mJW/8fZrUTk X-OriginatorOrg: rambus.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 17 Sep 2026 22:59:34.5710 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: 2df82056-0781-49b1-61ad-08df150f5779 X-MS-Exchange-CrossTenant-Id: bd0ba799-c2b9-413c-9c56-5d1731c4827c X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=bd0ba799-c2b9-413c-9c56-5d1731c4827c;Ip=[192.86.86.210];Helo=[hqxsv-psmtppxy02.rambus.com] X-MS-Exchange-CrossTenant-AuthSource: SJ5PEPF000001C8.namprd05.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: DM6PR04MB6778 Content-Type: text/plain; charset="utf-8" Add a minimal kselftest exercising the /dev/cmh_mgmt ioctl interface: - open/close the device node - invalid ioctl returns -ENOTTY - bad version field returns -EINVAL - KEY_NEW + KEY_DELETE lifecycle - KIC HKDF1 key derivation - ML-KEM-768 keygen via hardware RNG Tests use the kselftest_harness.h fixture framework and output TAP. Tests that require hardware features not present on the device under test are gracefully skipped (SKIP). Signed-off-by: Alex Ousherovitch Co-developed-by: Saravanakrishnan Krishnamoorthy Signed-off-by: Saravanakrishnan Krishnamoorthy --- tools/testing/selftests/Makefile | 1 + .../selftests/drivers/crypto/cmh/Makefile | 6 + .../drivers/crypto/cmh/cmh_mgmt_test.c | 183 ++++++++++++++++++ .../selftests/drivers/crypto/cmh/config | 1 + 4 files changed, 191 insertions(+) create mode 100644 tools/testing/selftests/drivers/crypto/cmh/Makefile create mode 100644 tools/testing/selftests/drivers/crypto/cmh/cmh_mgmt_tes= t.c create mode 100644 tools/testing/selftests/drivers/crypto/cmh/config diff --git a/tools/testing/selftests/Makefile b/tools/testing/selftests/Mak= efile index 2d960626750e..71eed3f8b591 100644 --- a/tools/testing/selftests/Makefile +++ b/tools/testing/selftests/Makefile @@ -19,6 +19,7 @@ TARGETS +=3D dax TARGETS +=3D devices/error_logs TARGETS +=3D devices/probe TARGETS +=3D dmabuf-heaps +TARGETS +=3D drivers/crypto/cmh TARGETS +=3D drivers/dma-buf TARGETS +=3D drivers/ntsync TARGETS +=3D drivers/s390x/uvdevice diff --git a/tools/testing/selftests/drivers/crypto/cmh/Makefile b/tools/te= sting/selftests/drivers/crypto/cmh/Makefile new file mode 100644 index 000000000000..86cb63839b27 --- /dev/null +++ b/tools/testing/selftests/drivers/crypto/cmh/Makefile @@ -0,0 +1,6 @@ +# SPDX-License-Identifier: GPL-2.0 +TEST_GEN_PROGS :=3D cmh_mgmt_test + +CFLAGS +=3D -Wall -Wno-misleading-indentation -O2 $(KHDR_INCLUDES) + +include ../../../lib.mk diff --git a/tools/testing/selftests/drivers/crypto/cmh/cmh_mgmt_test.c b/t= ools/testing/selftests/drivers/crypto/cmh/cmh_mgmt_test.c new file mode 100644 index 000000000000..8c2269a3cd13 --- /dev/null +++ b/tools/testing/selftests/drivers/crypto/cmh/cmh_mgmt_test.c @@ -0,0 +1,183 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Kselftest for /dev/cmh_mgmt ioctl interface. + * + * Tests basic ioctl operations on the Rambus CryptoManager Hub management + * device. Requires the cmh module loaded on real or emulated hardware. + * + * Run: ./cmh_mgmt_test + * Output: TAP format (compatible with kselftest harness) + */ + +#include +#include +#include +#include +#include +#include + +#include "kselftest_harness.h" +#include + +#define CMH_DEV "/dev/cmh_mgmt" + +FIXTURE(cmh_mgmt) +{ + int fd; +}; + +FIXTURE_SETUP(cmh_mgmt) +{ + self->fd =3D open(CMH_DEV, O_RDWR); + if (self->fd < 0 && errno =3D=3D ENOENT) + SKIP(return, "Device " CMH_DEV " not present (module not loaded?)"); + if (self->fd < 0 && errno =3D=3D EACCES) + SKIP(return, "Permission denied -- run as root or with CAP_SYS_ADMIN"); + ASSERT_GE(self->fd, 0); +} + +FIXTURE_TEARDOWN(cmh_mgmt) +{ + if (self->fd >=3D 0) + close(self->fd); +} + +/* + * Test 1: open and close succeed. + * If we get here, FIXTURE_SETUP already validated the open. + */ +TEST_F(cmh_mgmt, open_close) +{ + ASSERT_GE(self->fd, 0); +} + +/* + * Test 2: invalid ioctl number returns -ENOTTY. + */ +TEST_F(cmh_mgmt, invalid_ioctl) +{ + int ret; + unsigned long bogus_cmd =3D _IOC(_IOC_READ, 'J', 0xFF, 4); + + ret =3D ioctl(self->fd, bogus_cmd, NULL); + ASSERT_EQ(ret, -1); + ASSERT_EQ(errno, ENOTTY); +} + +/* + * Test 3: KEY_NEW with bad version field returns -EINVAL. + */ +TEST_F(cmh_mgmt, bad_version) +{ + struct cmh_ioctl_key_new req; + int ret; + + memset(&req, 0, sizeof(req)); + req.version =3D 0; /* invalid */ + req.ds_type =3D CMH_DS_AES_KEY; + req.len =3D 32; + req.flags =3D CMH_FLAG_PT; + req.cid =3D 0xDEAD; + + ret =3D ioctl(self->fd, CMH_IOCTL_KEY_NEW, &req); + ASSERT_EQ(ret, -1); + ASSERT_EQ(errno, EINVAL); +} + +/* + * Test 4: KEY_NEW creates a key, KEY_DELETE destroys it. + */ +TEST_F(cmh_mgmt, key_new_delete) +{ + struct cmh_ioctl_key_new new_req; + struct cmh_ioctl_key_grant del_req; + int ret; + + memset(&new_req, 0, sizeof(new_req)); + new_req.version =3D CMH_MGMT_V1; + new_req.ds_type =3D CMH_DS_AES_KEY; + new_req.len =3D 32; + new_req.flags =3D CMH_FLAG_PT; + new_req.cid =3D 0x5E1F7E57ULL; /* "SELFTEST" */ + + ret =3D ioctl(self->fd, CMH_IOCTL_KEY_NEW, &new_req); + ASSERT_EQ(ret, 0); + ASSERT_NE(new_req.ref, (uint64_t)0); + + /* Delete the key */ + memset(&del_req, 0, sizeof(del_req)); + del_req.version =3D CMH_MGMT_V1; + del_req.ref =3D new_req.ref; + + ret =3D ioctl(self->fd, CMH_IOCTL_KEY_DELETE, &del_req); + ASSERT_EQ(ret, 0); +} + +/* + * Test 5: KIC HKDF1 key derivation from hardware base key. + * Requires at least one KIC base key provisioned (KIC_KEY1). + */ +TEST_F(cmh_mgmt, kic_hkdf1) +{ + struct cmh_ioctl_kic_hkdf1 req; + static const char label[] =3D "kselftest-label"; + int ret; + + memset(&req, 0, sizeof(req)); + req.version =3D CMH_MGMT_V1; + req.key_len =3D 32; + req.base_key =3D CMH_KIC_KEY1; + req.cid =3D 0x4B534C46ULL; /* "KSLF" */ + req.label =3D (uint64_t)(uintptr_t)label; + req.label_len =3D sizeof(label) - 1; + req.flags =3D CMH_KIC_FLAG_TEMP; + + ret =3D ioctl(self->fd, CMH_IOCTL_KIC_HKDF1, &req); + if (ret < 0 && errno =3D=3D EIO) + SKIP(return, "KIC base key 1 not provisioned on this device"); + ASSERT_EQ(ret, 0); + ASSERT_NE(req.ref, (uint64_t)0); +} + +/* + * Test 6: ML-KEM-768 keygen using hardware RNG. + * Verifies the PQC keygen path end-to-end. + */ +TEST_F(cmh_mgmt, ml_kem_keygen) +{ + struct cmh_ioctl_ml_kem_keygen req; + /* ML-KEM-768: ek =3D 384*3+32 =3D 1184, dk =3D 768*3+96 =3D 2400 */ + uint8_t ek[1184]; + uint8_t dk[2400]; + int ret; + + memset(&req, 0, sizeof(req)); + req.version =3D CMH_MGMT_V1; + req.k =3D 3; /* ML-KEM-768 */ + req.flags =3D CMH_QSE_FLAG_HW_RNG; + req.seed =3D 0; /* HW RNG */ + req.z =3D 0; /* HW RNG */ + req.ek =3D (uint64_t)(uintptr_t)ek; + req.dk =3D (uint64_t)(uintptr_t)dk; + req.dk_cid =3D 0; + req.dk_ref =3D 0; + + memset(ek, 0, sizeof(ek)); + memset(dk, 0, sizeof(dk)); + + ret =3D ioctl(self->fd, CMH_IOCTL_ML_KEM_KEYGEN, &req); + if (ret < 0 && errno =3D=3D ENODEV) + SKIP(return, "QSE core not available on this hardware"); + ASSERT_EQ(ret, 0); + + /* Verify output is non-zero (extremely unlikely for random keys) */ + { + int i, nonzero =3D 0; + + for (i =3D 0; i < 64; i++) + nonzero +=3D (ek[i] !=3D 0); + ASSERT_GT(nonzero, 0); + } +} + +TEST_HARNESS_MAIN diff --git a/tools/testing/selftests/drivers/crypto/cmh/config b/tools/test= ing/selftests/drivers/crypto/cmh/config new file mode 100644 index 000000000000..063c1dd0e23b --- /dev/null +++ b/tools/testing/selftests/drivers/crypto/cmh/config @@ -0,0 +1 @@ +CONFIG_CRYPTO_DEV_CMH=3Dm --=20 2.43.7 From nobody Fri Sep 25 01:20:34 2026 Received: from SJ2PR03CU001.outbound.protection.outlook.com (mail-westusazon11022113.outbound.protection.outlook.com [52.101.43.113]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D1FEA4AFE32; Thu, 17 Sep 2026 22:59:40 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.43.113 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789685983; cv=fail; b=eHwAOsayoGhiKwhFj7AksgATBA1y2a8qz6E/hIFxQsGgn3GOpyJ1qXVbGnTUEd9mzAE79e7xSsfwFicYmyboU9USJtJQ45fVWMuJ6Z4sgyfU4oyZQaB5Xd3VOAvqs4VRcVf6quCY/QEXCgbn5qwjLSQlQL3jAthd28HJoaLTi5c= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789685983; c=relaxed/simple; bh=Tn8pRExeiR8E40eWcYWAXOOiT7BnACMD13Ttw0WsxJk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=FJ/73+mmmGygkJEQW7+pnOC9y6aEnkFavwuA1ZqZ/umviXz0wLGekdK6lmzI7CESJ+hwO7m0n4A/M/18n/le0+IB9BwZZpEttFMjMsPcW4lGZr78YgZ1Rch73h7h4l6G11dv7gCGckssO7CfjrOoNzJuDpBpTPKLQv3Anm8GNrA= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com; spf=fail smtp.mailfrom=rambus.com; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b=DgLhoNRA; arc=fail smtp.client-ip=52.101.43.113 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=rambus.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=rambus.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=rambus.com header.i=@rambus.com header.b="DgLhoNRA" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=tN7lqNyT+fSwOn3yvdhQ1MhzORs4dr7a27u8Mk0l0pNOe56qAFA/dMA2oiq7l/FrDSIblf3Zb51zbALKmWr43LerOlfxBFoZcuITXasqIHzI21sSXZlsV62aGS9NAgrTsKN4Rzvz077xCc2hm2wWC+sZdN1uRORnAMrvinnwhNhUaTl5GApK4H1eNs+uOolHAqwJxyNg8K2obJfZQYeAIw1jb4boFvEmNes3yEoZRj7AUTSisCelLmi+Q4DzI1+/WP/7zyCtNEaAgXX0JTyv1/BEJKseBDxVYIs8Jb/JQgJ3S4lwihQ8X1oXlQcVaCrV06aHwFL7dEsujFT8xrGWHw== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=qZtj1DDYvUW8oiuad6xKilGPwnsefmueMCv1X6ZHsLk=; b=omCukHw0omDdvc/FnRvqTbvyv0hLV4PHVzKrvS2SoI/VuPBbhdjfh2KEzhgjpMsmZJ+CDpFg5yTzgq4YXFpA8T3VSPJhXksNu3olkxqbJcaFC6BGrRDutxoy41lItpbgkl28kOKtQySDKZFa6u8KZZclFQwxZgioKPd/OYJEHwEKFZhqeIDIyKiHwEfg7mjNIFEcgtTF+UA7UuB4Gr1VpSAAVc30x+iY7kCotNteP42pK0I6C7NhCs+YvvwBNNAD+BO3WcxUykKWyoTABn94T8NdbI+UEFPnKuiqwh+JwbcbcjwJ+axBfC2yCPLggDGE/vdXpPQLf2Jwr5g//Jqp/g== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 192.86.86.210) smtp.rcpttodomain=cryptography.com smtp.mailfrom=rambus.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=rambus.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=rambus.com; s=selector1; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=qZtj1DDYvUW8oiuad6xKilGPwnsefmueMCv1X6ZHsLk=; b=DgLhoNRAjEYqCT1Yh6eottyjeyrPtcKB7dc5ztdI3f62vtl6HKCZneub3u7e+xvNV915lmytEBQoQENcCqhPxJmLWPxTf+xOGxNcyMVONR/7fZFFbvptEEcKTRlRJplahNKZU47l5RD/UPfNonDYg6+gWfthfpVoegENXCHzLul634Ag5bItV0nfeLGPkZ+Om8t/asyq0eO4xF7lcM1WmQY7BRElbyaJpEVjP8Oil4eYQA+tTrUFmQ6uOYXwcN4QkhPM9f7vjdHtMCuBCgDBBkHY0hfRQl1tgJIACivjIk8f0FDAxIqJ+PUf2SxNmdHYmyHeuaRQKNSAwO9du6ZSQg== Received: from CYZPR20CA0020.namprd20.prod.outlook.com (2603:10b6:930:a2::25) by LV2PR04MB994088.namprd04.prod.outlook.com (2603:10b6:408:426::5) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.428.13; Thu, 17 Sep 2026 22:59:37 +0000 Received: from BN7PEPF00000096.namprd03.prod.outlook.com (2603:10b6:930:a2:cafe::3a) by CYZPR20CA0020.outlook.office365.com (2603:10b6:930:a2::25) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.428.11 via Frontend Transport; Thu, 17 Sep 2026 22:59:36 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 192.86.86.210) smtp.mailfrom=rambus.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=rambus.com; Received-SPF: Pass (protection.outlook.com: domain of rambus.com designates 192.86.86.210 as permitted sender) receiver=protection.outlook.com; client-ip=192.86.86.210; helo=hqxsv-psmtppxy02.rambus.com; pr=C Received: from hqxsv-psmtppxy02.rambus.com (192.86.86.210) by BN7PEPF00000096.mail.protection.outlook.com (10.167.245.74) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.451.8 via Frontend Transport; Thu, 17 Sep 2026 22:59:35 +0000 Received: from hqxsv-cmdev3-aousherovitch.rambus.com (hqn-lb-int-float.rambus.com [10.12.20.20]) by hqxsv-psmtppxy02.rambus.com (Postfix) with ESMTP id AF665180174C; Thu, 17 Sep 2026 22:59:32 +0000 (UTC) From: Alex Ousherovitch To: Albert Ou , Alex Ousherovitch , Conor Dooley , "David S. Miller" , Herbert Xu , Jonathan Corbet , Krzysztof Kozlowski , Palmer Dabbelt , Paul Walmsley , Rob Herring , Saravanakrishnan Krishnamoorthy , Shuah Khan Cc: Alexandre Ghiti , devicetree@vger.kernel.org, Joel Wittenauer , linux-api@vger.kernel.org, linux-crypto@vger.kernel.org, linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-riscv@lists.infradead.org, Shuah Khan , Thi Nguyen Subject: [PATCH v5 19/19] MAINTAINERS: add Rambus CryptoManager Hub (CMH) Date: Thu, 17 Sep 2026 15:59:28 -0700 Message-ID: <20260917225929.2494111-20-aousherovitch@rambus.com> X-Mailer: git-send-email 2.43.7 In-Reply-To: <20260917225929.2494111-1-aousherovitch@rambus.com> References: <20260917225929.2494111-1-aousherovitch@rambus.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: BN7PEPF00000096:EE_|LV2PR04MB994088:EE_ X-MS-Office365-Filtering-Correlation-Id: 964510c7-5de6-4dc5-601d-08df150f5805 X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|1800799024|376014|7416014|82310400026|23010399003|36860700016|921020|22082099003|18002099003|10067099003|56012099006|11063799006|6133799003; X-Microsoft-Antispam-Message-Info: pUDXKUYtUCTFsYnP9goUQSr0v3JVyjqRFL99iRK+JkwrRCtgGjs92tMZDyEXigWkeIJ4Vt5xwxdSYgutuoE5hwNVHguenMof/enl2c+mS9y4Xkv3srydVJfxaOAwsuu0gMV+/+HBnodhmFGchZAIBwHgZCBYxaUJpFquoeGARRQpBP8kESof6TRG95vIGCRu/ymTHNcT7MB4miAUJCwxSoz4p5U8QsJfOk9C2NjTV7QK0w2y2MeLyMKJ0px1WeHK6gifibbdx7Y8UFS75QyOJ3Kx+36dQiRUMwZZ0Vh4aRwL5NKmMmo9EIJ1BfPE9uquDICbAUhCjnQI9CUs9cB9/xmYgw7ItJY6eDowi8iGPWBZfmZsvSaXSLUT6yGhOv2VCrcDWabUQFG3Y+DqPFGi7jyG7ilRvGE6QKQQ6KApuKghM0aaOe/1yZsilvNsMttnlQyf0I5Fo2SERMYUwlUKjTOmrMI7BpWXXJT8f5jTtm1HeRhW2RoZGuRDaPebMAbjQaTCbWdC0vzWsSCDGCV2F6PHFZo00vvWoUMOoknRnKHzpboXj1cEciNEJcSLcbWDUkOfLyQFJ9TlavRqvW/XSG00PAudNBegCTiXT2vBNwc2oTofsme8ImhpHEr2JvmCybru0XLHS+iAn1aAnx15faHWOqoE7V0cmuMHUq18yTW2y/5/9Q2zCZT0FXx99PGw5qjSgQ56kS0Wu+izatzxhIXgMqOsV1XzgU0OafHwfyCkm+I+75KrWIS/x5SyQWcI X-Forefront-Antispam-Report: CIP:192.86.86.210;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:hqxsv-psmtppxy02.rambus.com;PTR:ErrorRetry;CAT:NONE;SFS:(13230040)(1800799024)(376014)(7416014)(82310400026)(23010399003)(36860700016)(921020)(22082099003)(18002099003)(10067099003)(56012099006)(11063799006)(6133799003);DIR:OUT;SFP:1102; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: ZgEPzbraCTw6w66xnMD0x5AiIA/VwZPibwGcutr14/tn7C7f8fgdNjInAmu4SJzvyfqETlKrAgwSjmstkx7UfddGr2p+Iy8YnXPuxoWbOaEJiVK4X6/QPQukdTQZoZLH7NVjIjBiHjNlRdgO/Aa5IaKJM7K6cVs1IWufjthV4fLF5SHEwOtgYbbiEKDmy97zqhkQRwGNBjkm2v85Nb67odQHMjv2ldkK0If/Tu/VjNkz4UrXnv/GyKq/ORY1KwU5ZTrTfV00xGZ4npM21pfx9eTcAWrjbsBoaahqdEwfLDpaIoB9aH354zJU888FOUptXStMXyE52+3QQZl0nR7RaXU08qQg6MSmbV3TRbbhj79DpFhRjlHX94wg4vKl5nNGDgyBj7qAYepvJqk3EkTx4EdLWnGh+x72eR/Yfs8oL1xWLtc+BxPION3JaBKIyGrv X-OriginatorOrg: rambus.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 17 Sep 2026 22:59:35.3656 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: 964510c7-5de6-4dc5-601d-08df150f5805 X-MS-Exchange-CrossTenant-Id: bd0ba799-c2b9-413c-9c56-5d1731c4827c X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=bd0ba799-c2b9-413c-9c56-5d1731c4827c;Ip=[192.86.86.210];Helo=[hqxsv-psmtppxy02.rambus.com] X-MS-Exchange-CrossTenant-AuthSource: BN7PEPF00000096.namprd03.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: LV2PR04MB994088 Content-Type: text/plain; charset="utf-8" Add MAINTAINERS entry for the Rambus CryptoManager Hub (CMH) hardware crypto accelerator driver under drivers/crypto/cmh/. Signed-off-by: Alex Ousherovitch Co-developed-by: Saravanakrishnan Krishnamoorthy Signed-off-by: Saravanakrishnan Krishnamoorthy --- MAINTAINERS | 17 +++++++++++++++++ 1 file changed, 17 insertions(+) diff --git a/MAINTAINERS b/MAINTAINERS index c3bcdcaa9b92..9e7bcc4ba834 100644 --- a/MAINTAINERS +++ b/MAINTAINERS @@ -22760,6 +22760,23 @@ L: linux-wireless@vger.kernel.org S: Maintained F: drivers/net/wireless/ralink/ =20 +RAMBUS CRYPTOMANAGER HUB (CMH) HARDWARE CRYPTO ACCELERATOR +M: Alex Ousherovitch +M: Saravanakrishnan Krishnamoorthy +R: Joel Wittenauer +R: Thi Nguyen +L: linux-crypto@vger.kernel.org +S: Maintained +F: Documentation/ABI/testing/cmh-mgmt +F: Documentation/ABI/testing/debugfs-driver-cmh +F: Documentation/ABI/testing/sysfs-driver-cmh +F: Documentation/crypto/device_drivers/cmh.rst +F: Documentation/devicetree/bindings/crypto/rambus,cmh-v1030.yaml +F: Documentation/userspace-api/ioctl/cmh_mgmt.rst +F: drivers/crypto/cmh/ +F: include/uapi/linux/cmh_mgmt_ioctl.h +F: tools/testing/selftests/drivers/crypto/cmh/ + RAMDISK RAM BLOCK DEVICE DRIVER M: Jens Axboe S: Maintained --=20 2.43.7