From nobody Thu Sep 24 16:07:15 2026 Received: from mail-wm2-f13.google.com (mail-wm2-f13.google.com [74.125.225.141]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EE24B522F19 for ; Tue, 22 Sep 2026 08:01:31 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.225.141 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790064101; cv=none; b=H61j8v7Nm4O7ZACmL4DixRAwH3jsCMhQsI2+fokpj7X3ip0WqJEohe+QTeHIkX68oFKeU4hlfDlJbaLU/25D+uybetxGxZnphyYxqsNxGZsfLJadrD4ow+8YcicdJVY4o1gl3VgXiSS1bqqeHL6IrnMZlcFHCj2zIMM1e4FgKTo= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790064101; c=relaxed/simple; bh=0N9CYZPDiuUu44aUShjKzfgWHYRWpHfe6zVi7EjY5IU=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=aOYJ4c575iKg6SRl5JpjNhXVa4hE685fFpi6wgXnBsuMY+aDAFanRyB7n/bTXs7H9XQkcFbuQlljcOhH65E+VCe6ImGS2n1AM75qeUnFQERunoT8P3/uisAWyqr8nyN+w1DMl8ZEs7d3lmkZ8obcbWjdd25WQUcbunKA898yFXo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=LW/Q5g1/; arc=none smtp.client-ip=74.125.225.141 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="LW/Q5g1/" Received: by mail-wm2-f13.google.com with SMTP id 5b1f17b1804b1-49fbdc010d7so1073815e9.0 for ; Tue, 22 Sep 2026 01:01:30 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1790064088; x=1790668888; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=qmb5ZYsqks7WEIdeRWH6smgklDya3V2PeA5rJDQKGQU=; b=LW/Q5g1/musUc+wsQK4IusW0asOj6tj7c7eh51Rgj6l2IhYI8sQ3lqURkVbwZvykYh /AHx6mokkub/41wLr2HzABA5xH4IebKukxUyKBazaGLstabmgYqFj7H+vOxNEW/eX3tz x4koAMlgDU3rujnkWNkrCuHVqIJLB3zWsrcYyfNQwq0mLKHx1hfrDfq1tft6I6vfvS5c YUJZib79BdukfylpzuT2niwpsVmPcn3FJSz+OvQxy8xrFZr7eevJZvYqW81pDPXvUMA0 T0mCazA9tRBTtXPQ0Qi/f9gxRWZdslN952mm9xyiyIz/+opRhw6bNj4/kcK3Dm82hi/I 0jZQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790064088; x=1790668888; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=qmb5ZYsqks7WEIdeRWH6smgklDya3V2PeA5rJDQKGQU=; b=wkq/hUSKQ7zQ6zXd+2kZuv5Ha+4GTs987M0x5HVOjlAj2ooAQnCFmYCi72dlCjnDAR rrK9eXdPvv2il5fghB9gUL/w0LScluk26OVX6ITiPxsHeBvnw9MCpDOtcLfGpO+x13C+ +T3xwAQ1nB1GJgwAFbWz17Gpwri8KxZ/S7hkOLS4FdU/FYvNCPGBY/9qKYvzAmYkZ8W1 nnGYzxBPFTpf439T0dEzB5rjvjHVw19c+qYP38iT/9zzIG9r6gxRqKvGl17iGuOj5Vne ywmlPcozyUwZ7ygozqPN5VmghIstkZlXW19had6WMerrrcA6UxqkRgjT6o93VUmQmnTm +Y8Q== X-Forwarded-Encrypted: i=1; AKwUvBz/pQP0VdGICOu0ozg3AdBwUzWLoI9sFn5WwMoBSIqLuCM4HmbDDVhYtlrSg/qlhimTbnYrI8dVQa11NfE=@vger.kernel.org X-Gm-Message-State: AFuF++niPvB3qXcj7UlQf6EHrA0CoWzXX8sbQzBl6AapwfYcBWQB9YTn b39peLN6GXEe3UUASsflLJOo+/94tscSSbpG2sH/8Q3cE3I5d1ctD8f5 X-Gm-Gg: AYBFou0gh3fcInpIqxrBfuWdygr7as5J9hwS1E0rx23dW0hCxpNoZ6dbl3rE/b1iMfq zgwDjr6L1gzJueYjUM8w8ghJxqYCEDR7QxH53m5h6apwBmFE/vQpnmiC/nWK+7vgIhAXGO8qSLM rD7K57z2Lpfyup6U8a8Z3NLE0qPjsxWcZ/iegxF5hcA5E4n+D7yOXSmTCi4ekWs+eLjUrBXjV/O FeVMZ/Gq4coWldmEVavGzmQ7t8D7Y3NuJmDJG3rruAEo5jUk+xGz0BKxJmoDna7vRDKhoXb8fma sT3dQ8dIBnD3KQr2t7vEvwKuCHCyXByL1NdK5z/Vg9advnhPizLPjZSsnzHTF34fbxZyi/fpGWO iwo6/3iYIYTLoLIzXZJbJZYBC7irA0irWimForTM1WJyy1ebOnGzTLY+UUfCQpaFZM480QX/Dpr EeGQSIkV1xdMXyeeR4hjnOteybPnq8pdG7xv6yWtbe+ExOxuP4MlyWUzDrUv3ii4HkdnaPIM+FH lOr19fazaR1I8HG5CVIsgPjYqY9JaQMUPHmej/3f/XIZ2xp/IrWj2czcNlOp7xyq7r/tkrZnCOV uN8= X-Received: by 2002:a05:600c:354e:b0:49e:479b:c13b with SMTP id 5b1f17b1804b1-49fc7dff668mr165624765e9.1.1790064085766; Tue, 22 Sep 2026 01:01:25 -0700 (PDT) Received: from OrangePi5-Plus.BB-HOME (20014C4E1B80530056971C6280202175.dsl.pool.telekom.hu. [2001:4c4e:1b80:5300:5697:1c62:8020:2175]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-49fdaaf97e2sm18248625e9.2.2026.09.22.01.01.24 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 22 Sep 2026 01:01:25 -0700 (PDT) From: Igor Paunovic To: Tomeu Vizoso , Oded Gabbay , Heiko Stuebner Cc: Rob Herring , Krzysztof Kozlowski , Conor Dooley , Jeff Hugo , Robert Foss , Sidong Yang , Diederik de Haas , Sebastian Reichel , Jiaxing Hu , Nicolas Dufresne , Jonas Karlman , Guangshuo Li , =?UTF-8?q?H=C3=BCseyin=20BIYIK?= , dri-devel@lists.freedesktop.org, linux-rockchip@lists.infradead.org, linux-arm-kernel@lists.infradead.org, devicetree@vger.kernel.org, linux-kernel@vger.kernel.org, Igor Paunovic Subject: [PATCH v2 01/11] accel/rocket: search every core slot when a core is removed Date: Tue, 22 Sep 2026 10:01:04 +0200 Message-ID: <20260922080114.44662-2-royalnet026@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260922080114.44662-1-royalnet026@gmail.com> References: <20260922080114.44662-1-royalnet026@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" rocket_remove() decrements rdev->num_cores for each core it removes, while find_core_for_dev() searches slots 0 to num_cores - 1. Unbinding the cores in the order they were bound therefore loses the last one: by the time it is removed the search range has already shrunk past its slot, so find_core_for_dev() returns -1 and rocket_remove() gives up without doing anything. num_cores never reaches zero, rocket_device_fini() never runs, and the file-scoped rdev keeps pointing at a device that is going away. Binding the cores again starts from that stale count, because rocket_probe() takes rdev->num_cores as the slot to fill. On an RK3588, which describes three cores, the second round lands on slots 1, 2 and 3 while rdev->cores was allocated with room for three: rocket fdab0000.npu: Rockchip NPU core 1 version: 1179210309 rocket fdac0000.npu: Rockchip NPU core 2 version: 1179210309 rocket fdad0000.npu: Rockchip NPU core 3 version: 1179210309 The write to rdev->cores[3] is past the end of the array. Nothing in tree reads the core array often enough to notice, so the overrun is silent today. It turned up while testing a devfreq series on top of this, where a worker walks every core a few times a second, and UBSAN caught the first bool it read out of the overrun entry: UBSAN: invalid-load in drivers/accel/rocket/rocket_devfreq.c:47:10 load of value 5 is not a valid value for type '_Bool' Workqueue: devfreq_wq devfreq_monitor Record how many slots were allocated and search all of them. Every core is then found on removal, num_cores reaches zero, the device is torn down and a later bind starts from a clean rdev. This does not make unbinding a single core out of several work: probe still takes num_cores as the slot to fill, and rocket_open() still reaches for cores[0] whether or not anything is there. The fourth patch takes care of both. Found by unbinding and rebinding all three cores on an Orange Pi 5 Plus. With this applied the cores land in slots 0, 1 and 2 every time, whichever order they are bound in, and the shared supply goes back to a single user after each round. Fixes: ed98261b4168 ("accel/rocket: Add a new driver for Rockchip's NPU") Cc: stable@vger.kernel.org Assisted-by: LLM sparse checkpatch Signed-off-by: Igor Paunovic Tested-by: Jiaxing Hu # RK3576, two cores Tested-by: Sidong Yang # RK3588, 3 cores supersedes: the one on unbinding a single core now points at patch 4 --- Two paragraphs changed from the standalone posting, which this series supersedes: the one on unbinding a single core now points at patch 4 instead of saying what is left, and the round counts are dropped from the last one. Posting: https://lore.kernel.org/r/20260904125936.26234-1-royalnet026@gmail.com The two Tested-by tags are from that thread. The round counts were taken on a kernel that carried the use-after-free fixed in 8/11; they were re-measured with the whole series applied, see the notes on patch 4 and the cover letter. The single-core unbind and rocket_open() cases in that paragraph are 2/2 of the lifecycle series, now rebased into this one as patch 4: https://lore.kernel.org/r/20260731064933.12548-3-royalnet026@gmail.com drivers/accel/rocket/rocket_device.c | 2 ++ drivers/accel/rocket/rocket_device.h | 1 + drivers/accel/rocket/rocket_drv.c | 2 +- 3 files changed, 4 insertions(+), 1 deletion(-) diff --git a/drivers/accel/rocket/rocket_device.c b/drivers/accel/rocket/ro= cket_device.c index 46e6ee1e72c5f..efd004194c1af 100644 --- a/drivers/accel/rocket/rocket_device.c +++ b/drivers/accel/rocket/rocket_device.c @@ -31,6 +31,8 @@ struct rocket_device *rocket_device_init(struct platform_= device *pdev, if (of_device_is_available(core_node)) num_cores++; =20 + rdev->max_cores =3D num_cores; + rdev->cores =3D devm_kcalloc(dev, num_cores, sizeof(*rdev->cores), GFP_KE= RNEL); if (!rdev->cores) return ERR_PTR(-ENOMEM); diff --git a/drivers/accel/rocket/rocket_device.h b/drivers/accel/rocket/ro= cket_device.h index ce662abc01d3d..c62d567010696 100644 --- a/drivers/accel/rocket/rocket_device.h +++ b/drivers/accel/rocket/rocket_device.h @@ -19,6 +19,7 @@ struct rocket_device { =20 struct rocket_core *cores; unsigned int num_cores; + unsigned int max_cores; }; =20 struct rocket_device *rocket_device_init(struct platform_device *pdev, diff --git a/drivers/accel/rocket/rocket_drv.c b/drivers/accel/rocket/rocke= t_drv.c index 8bbbce594883e..2bcfe4ab3c68f 100644 --- a/drivers/accel/rocket/rocket_drv.c +++ b/drivers/accel/rocket/rocket_drv.c @@ -223,7 +223,7 @@ static int find_core_for_dev(struct device *dev) { struct rocket_device *rdev =3D dev_get_drvdata(dev); =20 - for (unsigned int core =3D 0; core < rdev->num_cores; core++) { + for (unsigned int core =3D 0; core < rdev->max_cores; core++) { if (dev =3D=3D rdev->cores[core].dev) return core; } --=20 2.43.0 From nobody Thu Sep 24 16:07:15 2026 Received: from mail-wm2-f12.google.com (mail-wm2-f12.google.com [74.125.225.140]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1E587525A69 for ; Tue, 22 Sep 2026 08:01:32 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.225.140 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790064102; cv=none; b=PhYgcjORZ+H6aZDm1kNI+NuDjqSKUHckmTuE+uCHZ4hanrYnUhkYpUMUpID0fXuYcwgl3tRYEHY5/nREweMBQmdofqBjEHuDVRTngCN/AIz6mhU6/BTkgmBjaBhtUgJNnZJDj3u41RjAA3i6oIGktxU1aExgB+nOU/yyBfqpig4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790064102; c=relaxed/simple; bh=CT/RCzxKrrCttvy9qEHMfrwnpwtAFxBYeP2FfoG304A=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=h5SU87OmumqbgY9YnjNMx4JYM2zJ4C24MJ+48g0wXXFO8l1O5cvxtlxYEqEYWfshGh9AsjVdFz+fpqWu2AB32VxnWf0+oLdWTupofxn2TdVCytkkDhAgTopkTKJ7ZtFz7uf/zE2OzR4R/luJ/FnPXEuR8i1YUwo+zeIxSiqoesg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=Xsa2yV6B; arc=none smtp.client-ip=74.125.225.140 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="Xsa2yV6B" Received: by mail-wm2-f12.google.com with SMTP id 5b1f17b1804b1-49d0726cdbcso2017835e9.0 for ; Tue, 22 Sep 2026 01:01:32 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1790064088; x=1790668888; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=VxEjmrMLU3ArphO/8yrozyx0jV5tnXGHISJN2k9QOtM=; b=Xsa2yV6BCJsZXDyz+BN1mAjcnQftUKoKmSLj735UQaL8pyETkyzkCl+sqDwWdG9jf+ 2FxB4DwD82G0B0LRB7zNL1tfk6zsFHrArnC5Gi0kfActdq2YKEaD/9i0AmTycaqKxBpr ljMyN+LpPFkt6UybFjOHyMGbofR4Kpu3DV+FpBf/seX1pFDR5/ZTik8M881jP3LyRVuh TQOY6b1bFhL30pbXTGmc+8C0gOXxJr8Xbo9U2MyIHSm4S9u0jDmIPaCYEoDVa9IndhmF xg8ZrYPGnGpJhBjl6yuDxCBHPrLLfS1ZWQ4uOYLxzKpZ1dYEsAVtporNWu5XO00mQQ9z 2gOA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790064088; x=1790668888; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=VxEjmrMLU3ArphO/8yrozyx0jV5tnXGHISJN2k9QOtM=; b=y24unNVJOrhgOkfeqosNhT+b8cuJMFgG/acOfMPvAN1JKrqP2pCeolGvNetlCzaRyg rTiIUQYLsSQP9113SZM8EnSoerLSaOGoRer0XMNEd0E2Vkhv8Xou8YQGBU7M9ef+dy/z 3IteAvP6jyhGrsdVbILm/QZUNylLZ4Zj3qLLsFseEXBR827v0DZxSCM5pm+TOPYar/Ry pzBZFpto+sIQqW1RjIbJl2ojmwQ5OXqL+FLL1+Ape/1iJwJdMOw/iPT16mYEj+HxRVxD Hoyd8ui33xFjxKNsTqaYYsCmTwMaBk/kvMMTP/Ny5NDdOW3+XNRnJDRnnAh9KgO3YpMJ Fscg== X-Forwarded-Encrypted: i=1; AKwUvBxc//Y36RPZkipfuv6BnT9P5vKXbqW5FUplWv0WU4VIiZq/6pdaXvMLU3AlrSPKEr7h/YXiBmOon1Izf5g=@vger.kernel.org X-Gm-Message-State: AFuF++mCeAS1nxq3FSnwjYLVCXy2o1U05h3fTGQ42I9ueHJY/hu2OkB2 LMRheaxMFe1RO7lQ3XRJPSRJ9YBdaOS7jmXT+u6u1L5WRNPeLNZo0LLk X-Gm-Gg: AYBFou2os6NgQVGAtGBWlOod7UicnRx75Sqa34m7wQ2sKB3T+A0jRWp2lHW/ERcXDSG ZQ3k3aHHgDSFtimzH/UkVg4xQ14FvEpACiOBCqqdy5SfGM5NmmAr3MudCbYvDKjrKuCC2dLaQKB JppqiR14Co53JeTS+qrQr8KxdZjUMF/sMK/6JY696liBkJmunY8eotcNZasXplXK6UF1JwqUT2M ESw2G75+0gN6eJTjpTONVlhE1Wl1GHG8tL8dK04g2BkZzD6nHaCYNzxuLxmtyd2CSfw6hs/u16u BJWzv21vCyewO2Sk+LpyhNzDAx+WtRBxwJjan3pmgneA5z4DjngYqbrurhYatkcLYfDz51TVrBf hMkQzjqjRwzGke1JMUKKqCWK5G+9CrP580VqT97lxX7F5ARLx00o1dGuBrvLgPh86Ry5rQs6wkJ KDQZGE1W/ZXVQcvl69KQeDIs+UD3JNCuX4tYrb/XaezOryFIFCSNxwtumOsmzBQdd91xscxCVjL ehCeDD6hgA0Yv5T4aa1BKcv7lgDsttcb0Nnim2LzdLH5HesKGmo88S+B4bPKUGdj/7Pp+OKBb71 TAs= X-Received: by 2002:a05:600c:1c13:b0:49e:7186:f36e with SMTP id 5b1f17b1804b1-49fc7dc1e75mr208969905e9.1.1790064088073; Tue, 22 Sep 2026 01:01:28 -0700 (PDT) Received: from OrangePi5-Plus.BB-HOME (20014C4E1B80530056971C6280202175.dsl.pool.telekom.hu. [2001:4c4e:1b80:5300:5697:1c62:8020:2175]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-49fdaaf97e2sm18248625e9.2.2026.09.22.01.01.25 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 22 Sep 2026 01:01:26 -0700 (PDT) From: Igor Paunovic To: Tomeu Vizoso , Oded Gabbay , Heiko Stuebner Cc: Rob Herring , Krzysztof Kozlowski , Conor Dooley , Jeff Hugo , Robert Foss , Sidong Yang , Diederik de Haas , Sebastian Reichel , Jiaxing Hu , Nicolas Dufresne , Jonas Karlman , Guangshuo Li , =?UTF-8?q?H=C3=BCseyin=20BIYIK?= , dri-devel@lists.freedesktop.org, linux-rockchip@lists.infradead.org, linux-arm-kernel@lists.infradead.org, devicetree@vger.kernel.org, linux-kernel@vger.kernel.org, Igor Paunovic Subject: [PATCH v2 02/11] accel/rocket: number the cores by devicetree position, not bind order Date: Tue, 22 Sep 2026 10:01:05 +0200 Message-ID: <20260922080114.44662-3-royalnet026@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260922080114.44662-1-royalnet026@gmail.com> References: <20260922080114.44662-1-royalnet026@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" rocket_job_hw_submit() programs the S_POINTER registers of a core with an extra bit derived from core->index, the way the vendor driver derives it from the hardware number of the core. rocket_probe() sets core->index to the slot the core takes in rdev->cores[], which is the order the cores bind in. The two agree only while the cores that bind are a prefix of the core nodes in the devicetree, in devicetree order. Unbind them and bind them back with a different core first, have one core's probe deferred behind a sibling's, or disable a core other than the last one, and every task submitted to a core whose slot is not its hardware number times out after 500 ms. The reset that follows does not help, and the inference finishes with wrong output. Observed on an Orange Pi 5 Plus with a KASAN build, over all six bind orders of the three cores: only the devicetree order ran clean. The other five produced 27 to 141 "NPU job timed out". In four of them no output tensor changed with the input, and the harness gave up before its first measured round; the fifth got through a six-second run with 27 timeouts and a wrong top-1 class. All but one of the timeouts land on the cores whose slot is not their hardware number, in proportion to the tasks the scheduler hands them, and in both directions of the mismatch. Number the cores by their position among the core nodes in the devicetree instead, which is what the hardware number is. The wrong value has been assigned since the driver was added, but it only reached the hardware once the extra bit was introduced, hence the Fixes tag below. Fixes: 0810d5ad88a1 ("accel/rocket: Add job submission IOCTL") Cc: stable@vger.kernel.org Assisted-by: LLM sparse checkpatch Signed-off-by: Igor Paunovic --- Supersedes the standalone posting: https://lore.kernel.org/r/20260905135612.7324-1-royalnet026@gmail.com Same diff. The message now says the numbers come from a KASAN build, corrects the timeout range to 27-141 against the raw log (it said 140), says what the four failed orders showed (the output did not change with the input; it said an oracle rejected the output), notes the one timeout that landed on a matching core, and drops the throughput figure, which was measured under KASAN. drivers/accel/rocket/rocket_core.h | 5 +++++ drivers/accel/rocket/rocket_drv.c | 31 +++++++++++++++++++++++++++++- 2 files changed, 35 insertions(+), 1 deletion(-) diff --git a/drivers/accel/rocket/rocket_core.h b/drivers/accel/rocket/rock= et_core.h index f6d7382854ca9..46ed8352a79d2 100644 --- a/drivers/accel/rocket/rocket_core.h +++ b/drivers/accel/rocket/rocket_core.h @@ -30,6 +30,11 @@ struct rocket_core { struct device *dev; struct rocket_device *rdev; + /* + * Hardware number of the core: its position among the core nodes in + * the devicetree. Not an index into rdev->cores[] - that slot is what + * find_core_for_dev() returns. + */ unsigned int index; =20 int irq; diff --git a/drivers/accel/rocket/rocket_drv.c b/drivers/accel/rocket/rocke= t_drv.c index 2bcfe4ab3c68f..7d927bb6b322d 100644 --- a/drivers/accel/rocket/rocket_drv.c +++ b/drivers/accel/rocket/rocket_drv.c @@ -157,10 +157,39 @@ static const struct drm_driver rocket_drm_driver =3D { .desc =3D "rocket DRM", }; =20 +/* + * The extra bit that rocket_job_hw_submit() sets in the S_POINTER registe= rs + * is the hardware number of the core, which is its position among the core + * nodes in the devicetree: a disabled core keeps its number. The slot a c= ore + * takes in rdev->cores[] is the order the cores happened to bind in, and = the + * two only agree while the cores that bind are a prefix of those nodes, in + * devicetree order. Every task submitted to a core whose slot is not its + * hardware number then times out. + */ +static int rocket_core_hw_index(struct device *dev) +{ + struct device_node *np; + int index =3D 0; + + for_each_matching_node(np, dev->driver->of_match_table) { + if (np =3D=3D dev->of_node) { + of_node_put(np); + return index; + } + index++; + } + + return -ENODEV; +} + static int rocket_probe(struct platform_device *pdev) { + int index =3D rocket_core_hw_index(&pdev->dev); int ret; =20 + if (index < 0) + return index; + if (rdev =3D=3D NULL) { /* First core probing, initialize DRM device. */ rdev =3D rocket_device_init(drm_dev, &rocket_drm_driver); @@ -176,7 +205,7 @@ static int rocket_probe(struct platform_device *pdev) =20 rdev->cores[core].rdev =3D rdev; rdev->cores[core].dev =3D &pdev->dev; - rdev->cores[core].index =3D core; + rdev->cores[core].index =3D index; =20 rdev->num_cores++; =20 --=20 2.43.0 From nobody Thu Sep 24 16:07:15 2026 Received: from mail-wr2-f12.google.com (mail-wr2-f12.google.com [74.125.225.76]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E906D525A90 for ; Tue, 22 Sep 2026 08:01:34 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.225.76 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790064098; cv=none; b=l8Mc+1FuAYDws84NMikd1Vuh2niiqvVA0gubyhgUtYpePqShbl5hmkY7+FggkrH99SduHnVluPuK/gMqK1t0fWS2YCXMfW/pJa5CI/qyRpsQ2v+gNaG/DuyRJ2CnSCaDadNFR0sIlYBMmQPQRT6JyVtx+PfiLTqQUpFP7uZ0Xh4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790064098; c=relaxed/simple; bh=R/0rtgRVcCRx4mq+bwcXP9RyWXmNvIpagAFBrXuOlAs=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=PDtPc0QF1ovm38Ob/o9HQbKOfEB2s0dqKQN9mPxOmx6M7IqSW9FE7skDmt6mqk9QyUmh8SBwBoAoJj/525Qr3fSHvu8m04J48EEe+jVwOH8i1x57f3n6sxIqsspgGrIkvyZdKGcNEYMgr/+qCjVh8iS6jpCoIni+xhH/g+7cwHU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=pSDIbHGM; arc=none smtp.client-ip=74.125.225.76 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="pSDIbHGM" Received: by mail-wr2-f12.google.com with SMTP id ffacd0b85a97d-48436686a40so356919f8f.3 for ; Tue, 22 Sep 2026 01:01:34 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1790064091; x=1790668891; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=uM6eJlIue+3ZnByR9U7PRXrgVvVnosyPWdGdGl87I7k=; b=pSDIbHGMCkMsnOvAmWiH+7BNamsOOkQmKnpQmUjfiSrtl/Qx7W2bP128l4zLWyvh7x N8J3TqVWkjw2RWE6Zahs2B5jPlaQPiUlAwYHbRguDL8TK6+lWSlQ0GArNmbiOgfZDZQC 5RN3xHLeSB3KeilmXbQHhVnP1ZKOeTxNF3/CQ6ADqar79a1xCKVq23cIgnCy/q0DVGq4 3cy/qop0K7sQhwhy3ncvdjMFzGezKhCADrO6UXs5S+EqfW2XbfBIwEjo0iemc7TZ5qjk sC3Yh2eaq5xWiedEG4Z6J9c3JjPzfUl8Gjut6r/7mBdxMUCTvhwub22Y1Uh2kRsz+gpR VMjA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790064091; x=1790668891; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=uM6eJlIue+3ZnByR9U7PRXrgVvVnosyPWdGdGl87I7k=; b=H4VcZ+iP6dhqDs8XR+t0ghlUv/4Nv+OWJqsBF4btWTSJdsT3REw3QrNalNKdksqmMP KJdUUHFBEnavpifBA9eSuHiw09Ko4qrpeyjvJxc9vPzNz+n0z51JMIcvOPUDk3rGRKMD ljiNRW8vqQw2QqRXKAf9rTVj3zfto8VcDxHQH1tnxxfMzDjTYfzraizFfCn7HXOOQTeo SORCYPJS41W4mCaYLuW0m8t+ehJDwqPTMlpu8l0Z+aoZlSJ+zWUExiB7NfKL9IJNf53H r9H78dYLlCn0VrQpAGKdNi6Vn2GdGOQvZIdv3RHGof7gCtzktMMplr2rZ6DnFKRJ0lRv OR0w== X-Forwarded-Encrypted: i=1; AKwUvBzZ2frp7BBEqodguKImPL7oe6ZSmoqdmPoW4kUSww0NkoV0DLJl0CsybTj/gcjsMCO3GcrweChVt8GChFg=@vger.kernel.org X-Gm-Message-State: AFuF++neUa09udIdxozIqM/qQCJsHUVf0CiFwmASwp+VnTNWk3mTos2N ZAL1jUrLO+B7qS+0AhRsp4ZrG2hzGesehg2ry7OP0K3uWIocn1Y3yzvp X-Gm-Gg: AYBFou1apa8vDFcFnQ3lml+PRa29GMQzIVRwMMT+y3ptrpWHA4Y82O65XZQsx9oJtvZ EhcobHsiDoPC1qFO8aMCOOWZqBd51ED1js7OmPQvGwNMV5/Hr9sVZaqoXXDFrJHy3AuwZf+n+TY pdHUks4CzbVAEO0qAvVyAc72FjZ6QQK71zmDgZkvmWkVHWVVuQgGPDpPbtNWfbXnruMWtXcocEs 1SNuxN039Sm5yBK7hklZRn5WzSUIz7XiZqPmZs5LBzmGpC/gRyr393VNBOmUHHnjQTeaLbWy/Yo e1MYhQ/oHRNR6x9ywaCJ5V9ySV8EybZ2xg3nWRf2xJWf18vryAq46l+DESyCwlMFhJJ0XVdutV0 4RZ9r/unCs9ptkCBltx/PBGcQuFwRzRLv8WDcoLFbrMiQYR316o/Z351X0toEGObAoPDL3TZE9R SdJiufzCHPycj6Z1ptyqcMA7gc8WO2ofsiAW7DWIMWNg6GzHLcGfYD5zqb9RTbJYh/LsS0y9A3W NbnMq+rdWKpbzwsM6V4axDsRvUaCpOFcRW7VCMrGx8z4Sw0P7rAsRj5QrJkJk6HrwTeiu/WcPO/ Mb55 X-Received: by 2002:a05:600c:1d1b:b0:49e:7cc6:ec88 with SMTP id 5b1f17b1804b1-49fc7e00555mr163235615e9.1.1790064089430; Tue, 22 Sep 2026 01:01:29 -0700 (PDT) Received: from OrangePi5-Plus.BB-HOME (20014C4E1B80530056971C6280202175.dsl.pool.telekom.hu. [2001:4c4e:1b80:5300:5697:1c62:8020:2175]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-49fdaaf97e2sm18248625e9.2.2026.09.22.01.01.28 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 22 Sep 2026 01:01:29 -0700 (PDT) From: Igor Paunovic To: Tomeu Vizoso , Oded Gabbay , Heiko Stuebner Cc: Rob Herring , Krzysztof Kozlowski , Conor Dooley , Jeff Hugo , Robert Foss , Sidong Yang , Diederik de Haas , Sebastian Reichel , Jiaxing Hu , Nicolas Dufresne , Jonas Karlman , Guangshuo Li , =?UTF-8?q?H=C3=BCseyin=20BIYIK?= , dri-devel@lists.freedesktop.org, linux-rockchip@lists.infradead.org, linux-arm-kernel@lists.infradead.org, devicetree@vger.kernel.org, linux-kernel@vger.kernel.org, Igor Paunovic Subject: [PATCH v2 03/11] accel/rocket: search every core slot when looking up a scheduler Date: Tue, 22 Sep 2026 10:01:06 +0200 Message-ID: <20260922080114.44662-4-royalnet026@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260922080114.44662-1-royalnet026@gmail.com> References: <20260922080114.44662-1-royalnet026@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" sched_to_core() walks rdev->cores[] up to rdev->num_cores, and rocket_remove() decrements num_cores for every core it removes. Unbind a core that is not the last one and the cores behind it fall outside the search, so sched_to_core() returns NULL for a core that is still bound and still running jobs. Neither caller checks the result: rocket_job_run(): rocket_fence_create(core), core->dev rocket_job_timedout(): dev_err(core->dev, "NPU job timed out") Unbinding the middle core of the three on an RK3588 while three clients are submitting to all of them faults twice, once from the surviving core's job queue and once from its reset work: KASAN: null-ptr-deref in range [0x0000000000000220-0x0000000000000227] Workqueue: fdad0000.npu drm_sched_run_job_work [gpu_sched] pc : rocket_job_run+0x234/0x838 [rocket] Call trace: rocket_job_run+0x234/0x838 [rocket] drm_sched_run_job_work+0x2cc/0xad8 [gpu_sched] process_one_work+0x640/0x14f0 KASAN: null-ptr-deref in range [0x0000000000000000-0x0000000000000007] Workqueue: rocket-reset-2 drm_sched_job_timedout [gpu_sched] pc : rocket_job_timedout+0xf0/0x1e0 [rocket] Call trace: rocket_job_timedout+0xf0/0x1e0 [rocket] drm_sched_job_timedout+0x188/0x6a0 [gpu_sched] Both are the third core: the workqueue names are its device and its core->index, and it was left at slot 2 while num_cores had dropped to 2. Search all the slots that were allocated, the way find_core_for_dev() now does. A core that is still bound is then found, and the two callers get the pointer they already assume they have. This does not make unbinding one core out of several safe. An open client keeps an entity pointing at the scheduler of the core that went away: drm_sched reports it as not ready for every job that lands on it, and the client waits in dma_fence_default_wait for a fence that will never signal. Stopping the NULL dereference is what belongs in a fix; the rest wants more thought. Reported-by: Sidong Yang Closes: https://lore.kernel.org/dri-devel/apwUewaRnoTNXHCt@rock-5b-plus/ Fixes: 0810d5ad88a1 ("accel/rocket: Add job submission IOCTL") Cc: stable@vger.kernel.org Assisted-by: LLM sparse checkpatch Signed-off-by: Igor Paunovic --- Unchanged from the standalone posting, which this series supersedes: https://lore.kernel.org/r/20260905150432.7477-1-royalnet026@gmail.com drivers/accel/rocket/rocket_job.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/drivers/accel/rocket/rocket_job.c b/drivers/accel/rocket/rocke= t_job.c index f404355058185..4bc4f9c8ee403 100644 --- a/drivers/accel/rocket/rocket_job.c +++ b/drivers/accel/rocket/rocket_job.c @@ -283,7 +283,7 @@ static struct rocket_core *sched_to_core(struct rocket_= device *rdev, { unsigned int core; =20 - for (core =3D 0; core < rdev->num_cores; core++) { + for (core =3D 0; core < rdev->max_cores; core++) { if (&rdev->cores[core].sched =3D=3D sched) return &rdev->cores[core]; } --=20 2.43.0 From nobody Thu Sep 24 16:07:15 2026 Received: from mail-wm2-f13.google.com (mail-wm2-f13.google.com [74.125.225.141]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1210B522696 for ; Tue, 22 Sep 2026 08:01:36 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.225.141 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790064109; cv=none; b=LZsHkDpx0YBGU5K4JGeEmIqZGMBn3WEGIuRbFBFoTTdeSBosRp+45qOC1kXRBKeMDDasVPAGmDV0e0vvQA1y3cfEAyrDkK7AIdP05ggJl1yFSy/lrRNPCeKkotOfZ2l+XLzFt9+ev4mND+saqBhBQp8W32pE9XA1CxrkTmvjE90= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790064109; c=relaxed/simple; bh=A510BS/niUth96kgK37SkXZKR9eu25CRfc7TjnqqkU4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=nlP7fJdZT4e4+WBN43Cjk8y8b/PMfniu6eVWP/7U5IZZe3C2V6Cej1KBzvriJriDG53sfpW9K2YI11IVddc3jTYwYocxCKM1d+ATqrbtgBXR+GUGBCnOyiKeaSlY1GhU70Yle5xKMTAc03FBlMVgSz4Jgw/JOQrMCH0BAsSjtjo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=qPihAMkw; arc=none smtp.client-ip=74.125.225.141 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="qPihAMkw" Received: by mail-wm2-f13.google.com with SMTP id 5b1f17b1804b1-49d0726cdbcso2017945e9.0 for ; Tue, 22 Sep 2026 01:01:36 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1790064092; x=1790668892; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=pp6BX0mvVT2CZt7W5iAnQc+T8SkoWRi5DiwXWjU82F8=; b=qPihAMkw0d9GewmanxpPEinT4zbUbedFAl5MhGKYboDmas5x4Qwvj5Wb2gziT+CiuW bwAOLwvKuNCnYCXxNqaEBDQomk8yu/0S5M4sy5qjXITuWCPUcUm33n8pVQ9q2M17W6S4 y1AWLKpmZzSev796eGgaac/sQxulXWWtAmNoXS30ADz6epliXmUJb1a+xWvyaa2E2Sz8 BGY+3ugvM++SXwPd2ysA/9pSmCAedrUwxWTWjJYCJ1RFrwfzQ5K/bK/jeMSQq5NO9MeO CzdBsYyCfDCdat/CANhxILvczd/+mYMfqDiPDrYxCWqJFaAhXbYhIMCAMMFlgHRY130Q 0yNQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790064092; x=1790668892; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=pp6BX0mvVT2CZt7W5iAnQc+T8SkoWRi5DiwXWjU82F8=; b=Kw8y2cgCEXqt4/xZVg0HJEFer3C4QnyXio86WV+OamiguEQp4vZ+Y0Ta7YfR3MOQ+a HUHuZGKfY2VYjBPgc/jOOnMA/sZHcU02VW7jMo474k231VbA1XkDYLgWwfevbEEBiNMm 0W7bSn1BlV+Tn0RDMEw96A5JuJ95a3CSbj/hPEQXW5vXyc2B3vJoQZLtsl125G37OVF5 KcZisFDKSx0g+U/GsbEVbj3SnrlveyrSghveq6JcFQaJOcwhM1GkCuDbOSpfiC4fWSES JLj5G+DOVTZMx1/xulW6HaqYcJD5cpzGK8n4JLh27jNVQMws9E3K5iiIh06eTTvqhVEE dJIQ== X-Forwarded-Encrypted: i=1; AKwUvBweApKFf5S9MOGmUSpNfcNcRBWsg0aU6Ikfzm0Pp9OSCmLa6DAwg8X4iKscwA2U6iEesEdLIF//IyaAn78=@vger.kernel.org X-Gm-Message-State: AFuF++lrgh4pc/xXPkaOcYXkVFOCLzUqnqLUZe8ira3IP2GB9+NMN6aC EzeDP0JKYJG4WfyTsBUBDkWgxLc+r0vXlnAvgKVbK8JhmC1TagOWxEEd X-Gm-Gg: AYBFou24mmnOiDpTbNS10TX/1+ns3NAcBOGxBQz89o/aZzgj7dObOAz8x9gG9aDF4lv 421xe3/4tJELLvZXnSMVJ92u1dUcM7XjrGzoC59xFpfonRpvmzXo06o2re/qrmaq+NAX4bfqQym o2Lju7HFUTSA9YYPN3o5GicnAxFU7a4UHyQQ5FH2Y3EKmhJSqp8sd/nrSJuPg6HXT3SqgnwdmY/ VtYO4JcyMgejxvkCoMm7tjx4Ag2ELhN6CUEbS7LxM/hTo1xRqLKvSuplcsse2SMp4luWOKgWo+o nFagH4lQSYzf20Ht9jP3svHqwlMb9ZdCWPKqNpaU4ijggMTzwpFbfjW3rcKr0jfdIJQ9q9nLlQN /g78WhQJ5mwhmuDZ0x70oxh5ckAg3XHAa20tgOslVrbL6okRpopKiKWLBMkqJdhBaGfU6ZfUAYK FySfkLoT8UomZt/VYYEzjYYihiLQ9AoUUY2BRCyp1RKAGKD2H1LSrtgKfU4/ruj1t1/HM+P0gEg Ff+OrgcRlx/hvsj+oGKcQNg5mf2/braQL8FEEZu6hKR/TSI4wPrbYi5S/JkR2A7OGcn96Itkek4 Nmm0MR5/RRKxeT4= X-Received: by 2002:a05:600c:6992:b0:49b:9241:7ff0 with SMTP id 5b1f17b1804b1-49fc7b7a756mr226927785e9.0.1790064090983; Tue, 22 Sep 2026 01:01:30 -0700 (PDT) Received: from OrangePi5-Plus.BB-HOME (20014C4E1B80530056971C6280202175.dsl.pool.telekom.hu. [2001:4c4e:1b80:5300:5697:1c62:8020:2175]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-49fdaaf97e2sm18248625e9.2.2026.09.22.01.01.29 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 22 Sep 2026 01:01:30 -0700 (PDT) From: Igor Paunovic To: Tomeu Vizoso , Oded Gabbay , Heiko Stuebner Cc: Rob Herring , Krzysztof Kozlowski , Conor Dooley , Jeff Hugo , Robert Foss , Sidong Yang , Diederik de Haas , Sebastian Reichel , Jiaxing Hu , Nicolas Dufresne , Jonas Karlman , Guangshuo Li , =?UTF-8?q?H=C3=BCseyin=20BIYIK?= , dri-devel@lists.freedesktop.org, linux-rockchip@lists.infradead.org, linux-arm-kernel@lists.infradead.org, devicetree@vger.kernel.org, linux-kernel@vger.kernel.org, Igor Paunovic Subject: [PATCH v2 04/11] accel/rocket: keep core slots stable across unbind and rebind Date: Tue, 22 Sep 2026 10:01:07 +0200 Message-ID: <20260922080114.44662-5-royalnet026@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260922080114.44662-1-royalnet026@gmail.com> References: <20260922080114.44662-1-royalnet026@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" rocket_probe() inserts a new core at slot num_cores, and rocket_remove() only decrements that counter without clearing the slot. That falls apart as soon as one core is unbound while its siblings stay bound: - the next bind reuses the slot of a still-live core and overwrites it while its IRQ handler (dev_id points into cores[]) and its DRM scheduler are still active; - rocket_open() unconditionally uses cores[0].dev, which after an unbind of core 0 is a stale pointer to an unbound device. On an RK3588 with three cores, unbinding the first one and binding it again puts it on top of the third: rocket fdab0000.npu: drm_sched_init: scheduler already initialized! One device now sits in two slots and the third core in none. The next unbind of the first core finds its stale slot, torn down already, and finishes the same scheduler a second time: Unable to handle kernel NULL pointer dereference at virtual address 00000= 00000000000 pc : drm_sched_fini+0x4c/0x1e0 [gpu_sched] Call trace: drm_sched_fini+0x4c/0x1e0 [gpu_sched] (P) rocket_job_fini+0x28/0x60 [rocket] rocket_core_fini+0x4c/0x78 [rocket] rocket_remove+0x78/0x110 [rocket] platform_remove+0x2c/0x68 device_remove+0x58/0xc0 device_release_driver_internal+0x214/0x2e0 device_driver_detach+0x24/0x50 unbind_store+0xd8/0xe8 Make .dev the slot-liveness marker: probe takes the first free slot and clears it again if core init fails, remove clears .dev after rocket_core_fini() and warns if the core cannot be found, lookups skip empty slots, and rocket_open() and rocket_job_open() use only live slots. num_cores keeps counting bound cores for the last-core teardown check. A missing core is no reason to refuse a new file: the device is one core short, not gone, and any bound core will do for the IOMMU domain, which is attached to the group of whichever core runs a job. With a single core live the scheduler list that drm_sched_entity_init() does not keep is freed at once, as rocket_job_close() only frees what the entity kept. Fixes: ed98261b4168 ("accel/rocket: Add a new driver for Rockchip's NPU") Cc: stable@vger.kernel.org Assisted-by: LLM sparse checkpatch Signed-off-by: Igor Paunovic --- v3, now in this series: - rebased onto the three fixes before it: "accel/rocket: search every core slot when a core is removed", which provides max_cores, "accel/rocket: number the cores by devicetree position, not bind order", so a slot no longer doubles as the hardware number, and "accel/rocket: search every core slot when looking up a scheduler" - the crash above: this is what the devfreq patches ran into when tested on a kernel without this one - free the scheduler list when a single core is live - Cc: stable v2: https://lore.kernel.org/r/20260731064933.12548-3-royalnet026@gmail.com - also clear the slot's .dev when rocket_core_init() fails (Jiaxing Hu) - check .dev in sched_to_core() (Jiaxing Hu) - document the synchronous-probe assumption at the slot scan v1: https://lore.kernel.org/dri-devel/20260730080355.177422-3-royalnet026@g= mail.com/ Unbinding a core that still has jobs in flight, or that an open file already holds the scheduler of, has further pre-existing issues that are out of scope for this bookkeeping fix. One of them, a use-after-free under KASAN, is described in the cover letter. The driver does not serialize probe and remove against open; this does not change that. For stable, this goes with the three patches before it. Verified on RK3588 (Orange Pi 5 Plus), 7.3.0-rc2 drm-misc-next plus this series, in-tree rocket, all three cores enabled: 25 rounds of unbinding and rebinding all three cores; 4 rounds of unbinding a single core (twice the devicetree-first core, twice the second one), each with an inference run while the core was absent and one more after all four; 5 rmmod/modprobe rounds; 3 unbind/rebind rounds and 1 rmmod with the clock raised to the 1 GHz OPP (CRU selector on the PVTPLL before each). No "scheduler already initialized" message and no oops; the regulator user count returns to its boot value after every round. Without this patch the same single-core round oopses in drm_sched_fini() on the second unbind, as shown above. The single-core, full, rmmod and raised-clock rounds were repeated (10 full rounds and 3 rmmod rounds this time) on a KASAN and PROVE_LOCKING build of the same tree: no report, and lockdep still enabled afterwards. drivers/accel/rocket/rocket_device.h | 2 + drivers/accel/rocket/rocket_drv.c | 55 +++++++++++++++++++++++----- drivers/accel/rocket/rocket_job.c | 38 ++++++++++++++----- 3 files changed, 77 insertions(+), 18 deletions(-) diff --git a/drivers/accel/rocket/rocket_device.h b/drivers/accel/rocket/ro= cket_device.h index c62d567010696..abb88a254e569 100644 --- a/drivers/accel/rocket/rocket_device.h +++ b/drivers/accel/rocket/rocket_device.h @@ -18,7 +18,9 @@ struct rocket_device { struct mutex sched_lock; =20 struct rocket_core *cores; + /* Number of currently bound cores. */ unsigned int num_cores; + /* Slot capacity (DT core count); slots with a NULL .dev are free. */ unsigned int max_cores; }; =20 diff --git a/drivers/accel/rocket/rocket_drv.c b/drivers/accel/rocket/rocke= t_drv.c index 7d927bb6b322d..b9b36c578db20 100644 --- a/drivers/accel/rocket/rocket_drv.c +++ b/drivers/accel/rocket/rocket_drv.c @@ -68,11 +68,21 @@ rocket_iommu_domain_put(struct rocket_iommu_domain *dom= ain) kref_put(&domain->kref, rocket_iommu_domain_destroy); } =20 +static struct rocket_core *rocket_first_live_core(struct rocket_device *rd= ev) +{ + for (unsigned int core =3D 0; core < rdev->max_cores; core++) + if (rdev->cores[core].dev) + return &rdev->cores[core]; + + return NULL; +} + static int rocket_open(struct drm_device *dev, struct drm_file *file) { struct rocket_device *rdev =3D to_rocket_device(dev); struct rocket_file_priv *rocket_priv; + struct rocket_core *core; u64 start, end; int ret; =20 @@ -85,8 +95,18 @@ rocket_open(struct drm_device *dev, struct drm_file *fil= e) goto err_put_mod; } =20 + /* + * Any bound core will do for the domain: it is attached to the group + * of whichever core runs a job, and the NPU IOMMUs are all the same. + */ + core =3D rocket_first_live_core(rdev); + if (!core) { + ret =3D -ENODEV; + goto err_free; + } + rocket_priv->rdev =3D rdev; - rocket_priv->domain =3D rocket_iommu_domain_create(rdev->cores[0].dev); + rocket_priv->domain =3D rocket_iommu_domain_create(core->dev); if (IS_ERR(rocket_priv->domain)) { ret =3D PTR_ERR(rocket_priv->domain); goto err_free; @@ -199,10 +219,21 @@ static int rocket_probe(struct platform_device *pdev) } } =20 - unsigned int core =3D rdev->num_cores; + unsigned int core; =20 dev_set_drvdata(&pdev->dev, rdev); =20 + /* + * Take the first free slot: cores can unbind and rebind in any + * order. The scan-then-claim relies on platform probes running + * sequentially; revisit if the driver ever enables async probe. + */ + for (core =3D 0; core < rdev->max_cores; core++) + if (!rdev->cores[core].dev) + break; + if (WARN_ON(core =3D=3D rdev->max_cores)) + return -ENXIO; + rdev->cores[core].rdev =3D rdev; rdev->cores[core].dev =3D &pdev->dev; rdev->cores[core].index =3D index; @@ -210,13 +241,18 @@ static int rocket_probe(struct platform_device *pdev) rdev->num_cores++; =20 ret =3D rocket_core_init(&rdev->cores[core]); - if (ret) { - rdev->num_cores--; + if (ret) + goto err_core; =20 - if (rdev->num_cores =3D=3D 0) { - rocket_device_fini(rdev); - rdev =3D NULL; - } + return 0; + +err_core: + rdev->cores[core].dev =3D NULL; + rdev->num_cores--; + + if (rdev->num_cores =3D=3D 0) { + rocket_device_fini(rdev); + rdev =3D NULL; } =20 return ret; @@ -229,10 +265,11 @@ static void rocket_remove(struct platform_device *pde= v) struct device *dev =3D &pdev->dev; int core =3D find_core_for_dev(dev); =20 - if (core < 0) + if (WARN_ON(core < 0)) return; =20 rocket_core_fini(&rdev->cores[core]); + rdev->cores[core].dev =3D NULL; rdev->num_cores--; =20 if (rdev->num_cores =3D=3D 0) { diff --git a/drivers/accel/rocket/rocket_job.c b/drivers/accel/rocket/rocke= t_job.c index 4bc4f9c8ee403..25ee4ab172a82 100644 --- a/drivers/accel/rocket/rocket_job.c +++ b/drivers/accel/rocket/rocket_job.c @@ -284,7 +284,7 @@ static struct rocket_core *sched_to_core(struct rocket_= device *rdev, unsigned int core; =20 for (core =3D 0; core < rdev->max_cores; core++) { - if (&rdev->cores[core].sched =3D=3D sched) + if (rdev->cores[core].dev && &rdev->cores[core].sched =3D=3D sched) return &rdev->cores[core]; } =20 @@ -511,22 +511,42 @@ void rocket_job_fini(struct rocket_core *core) int rocket_job_open(struct rocket_file_priv *rocket_priv) { struct rocket_device *rdev =3D rocket_priv->rdev; - struct drm_gpu_scheduler **scheds =3D kmalloc_objs(*scheds, - rdev->num_cores); - unsigned int core; + struct drm_gpu_scheduler **scheds; + unsigned int core, n =3D 0; int ret; =20 - for (core =3D 0; core < rdev->num_cores; core++) - scheds[core] =3D &rdev->cores[core].sched; + scheds =3D kmalloc_objs(*scheds, rdev->max_cores); + if (!scheds) + return -ENOMEM; + + /* Only the cores that are bound right now have a scheduler to offer. */ + for (core =3D 0; core < rdev->max_cores; core++) + if (rdev->cores[core].dev) + scheds[n++] =3D &rdev->cores[core].sched; + + if (!n) { + ret =3D -ENODEV; + goto err_free; + } =20 ret =3D drm_sched_entity_init(&rocket_priv->sched_entity, DRM_SCHED_PRIORITY_NORMAL, - scheds, - rdev->num_cores, NULL); + scheds, n, NULL); if (WARN_ON(ret)) - return ret; + goto err_free; + + /* + * drm_sched_entity_init() keeps the list only when it holds more + * than one scheduler, and rocket_job_close() frees what it kept. + */ + if (n < 2) + kfree(scheds); =20 return 0; + +err_free: + kfree(scheds); + return ret; } =20 void rocket_job_close(struct rocket_file_priv *rocket_priv) --=20 2.43.0 From nobody Thu Sep 24 16:07:15 2026 Received: from mail-wm2-f13.google.com (mail-wm2-f13.google.com [74.125.225.141]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 42633522F14 for ; Tue, 22 Sep 2026 08:01:35 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.225.141 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790064107; cv=none; b=ameb5hlss6JfHfomhV6CFr0QGSTzTBEi+BFprtCBjLu0sh7Hwp6a0IlVULv3XTgp01yoNwDKTrWyX6meiDyXYgJ77N3H6Cu+c3I9UqG6lge5yhAPKjG3jR7uW5r6KK8iMV2vETIM9roN1X3L7WPOuVLz9vesvD2Z6KHuFKXyWMU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790064107; c=relaxed/simple; bh=0xrVY0R026KpNfsviBE2/QZKTNvfEgT3VjQqd30Zu1k=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=kFLWroU5L35nV11vhgtX59rXn34YUliJhBLimW33G136w/aRmqxA+7Cd1JsEWer+Ps9CVyOuuN0/AogYvOvLM2ncyifg6CLNpKXi44yPb0FNC+R5m2qHNT0tq7aVC1yCzmIbvGlXJFHNZBN1QYGGRLqcKNf7GTakRC3C5PRAdjE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=Y6htxH0A; arc=none smtp.client-ip=74.125.225.141 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="Y6htxH0A" Received: by mail-wm2-f13.google.com with SMTP id 5b1f17b1804b1-49fbca514f8so2712655e9.2 for ; Tue, 22 Sep 2026 01:01:35 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1790064092; x=1790668892; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=hGF28/ongkg1pjDH+01IcEm6dgrPiae6qX55Ok4Xvzs=; b=Y6htxH0AHk6grdEK/nrW0GflTyjDSZnNsOLpl4zypHP8VM1mli2uXXKZTpp0nKEVah XjaTVcfSQ/xnGmk5RFdmvzxFmYbmJS1p/qZ9RO0lh0zRqVys5NuUbyKu15fAvC5+HkhL fpDCe2YUkyU5bG7/6JQVBXBA02iIiXEnahe3YUZzw3pq1paBE2xVmSbsgNwrikBNX17+ KvnOPd81nREC+264jRNDZw1LuWfMQgi8WII/1N37FIQ4WAZF0uofWDl970uNZyaqLcBI f2GmCuzyB6OdpM8ryAgX/uI77qEOIupCarLXmxvTKgTwOhFSj6pD7JZZv6IXQtA3Q73k crVQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790064092; x=1790668892; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=hGF28/ongkg1pjDH+01IcEm6dgrPiae6qX55Ok4Xvzs=; b=1J9LgqF5pFsxEJEUASiJ6J8H9Il+vP1QKGCeTs0PVsq9aA+U7RQKrRtIACDlFL/0XE KnQAjN67fFhGds4g7G22nsPsz1PBuN1bmqwf2FUBn4DUNkau5JLfrxuMWeNiL5gyOQ4Q DMYTfSUepvfiEN/47uE7idhIdZnRR4cDqk0VmFVbc2Pd9dYF+21rdf+FJ+3Mw6w5ucwX Qti90cq9MHUcLwwLM9GdydvzV87tK/6a/AemNceBnqyE0QgpU34lNM8dHRXdQybILaxf kah5JNDjiUtRFlqWuFl1t5LWtqPGHH6WIrdm57OdYVEL/Is9cmiqGMg7YvqQh6nKLSMC l1/A== X-Forwarded-Encrypted: i=1; AKwUvBxz/6ERDfeP/5dVPFIjztZ+lQ6BuGFDu9Bc5t54WaHPDtqwwwoXoOwfW4Tjq952lCZ75ubAfoqm+K22NSo=@vger.kernel.org X-Gm-Message-State: AFuF++nz7oVcj7qgVuYuDt6LbD1Brspk7M2OVO7pHmHwcAnsOsmrDpBZ ptOchXZKzeYiiEpc8wwz/4hliVKtg6QllAnokFTB142iaycUzi8jvS3LHJ6eCwvS X-Gm-Gg: AYBFou0p6mt9YxXi5oY4Of4wBCzsKEGNuIm+vHPavuGhyzZt3VxFkBuPSIdyVdwvhXu o99i3wapTP39+uYLlUU7YlqyCw27C+/wKtPnn7CeUqMSuI4XVODk1hthSZwI7NzQxZhbjjXdf/G 2x+IECxDoYnm2a0DrhtxMW83UaWKGozD0n3KTguJVOSLJMv46otvOGPRyzz5k9ctnTCvTTImdWc V1kT8dm/pq4ZavGB8uJeQBZo/Nxuv5TzX6gagKouW1j9Zsm48IkGwfzxJ5SPT/D4RK+Oqus/Qyq R6BnUyx6OimCPdQgmICim9v+YERayBZSesGBIKjG4spjeDaLu3YzSnOR/7ZRPsGTDnsM1FkkQ99 PPXLhySpYnKJzStfBpNgyz09QrGqjEgisLz1WlWX6rclmmMkW7DU8RsoR0tRg+zJWlziByE26rA 059qthneKo9Dqqcb67rxn5W9HUxaA/XehCx6T3s4wHgsDCbD4AamMme3YOCrbcgg7YLflVmRsil dgDkgHghIITkOhNBr5wbKzk32dx4MhS/uEWDm7ls6evQd2tpaY8W0q7r3CSA4JYgjwD+//dRyCg w0Y= X-Received: by 2002:a05:600c:3b99:b0:49e:6683:d227 with SMTP id 5b1f17b1804b1-49fc7bbfe4bmr199176535e9.0.1790064092385; Tue, 22 Sep 2026 01:01:32 -0700 (PDT) Received: from OrangePi5-Plus.BB-HOME (20014C4E1B80530056971C6280202175.dsl.pool.telekom.hu. [2001:4c4e:1b80:5300:5697:1c62:8020:2175]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-49fdaaf97e2sm18248625e9.2.2026.09.22.01.01.31 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 22 Sep 2026 01:01:31 -0700 (PDT) From: Igor Paunovic To: Tomeu Vizoso , Oded Gabbay , Heiko Stuebner Cc: Rob Herring , Krzysztof Kozlowski , Conor Dooley , Jeff Hugo , Robert Foss , Sidong Yang , Diederik de Haas , Sebastian Reichel , Jiaxing Hu , Nicolas Dufresne , Jonas Karlman , Guangshuo Li , =?UTF-8?q?H=C3=BCseyin=20BIYIK?= , dri-devel@lists.freedesktop.org, linux-rockchip@lists.infradead.org, linux-arm-kernel@lists.infradead.org, devicetree@vger.kernel.org, linux-kernel@vger.kernel.org, Igor Paunovic Subject: [PATCH v2 05/11] accel/rocket: request the core clocks by name Date: Tue, 22 Sep 2026 10:01:08 +0200 Message-ID: <20260922080114.44662-6-royalnet026@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260922080114.44662-1-royalnet026@gmail.com> References: <20260922080114.44662-1-royalnet026@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" rocket_core_init() hands core->clks to devm_clk_bulk_get() without ever setting the .id members. The rocket_core array is allocated with devm_kcalloc() in rocket_device_init(), and rocket_probe() only fills in .rdev, .dev and .index, so all four clk_bulk_data entries are requested with a NULL con_id (unlike core->resets, whose ids are set a few lines above). clk_get(dev, NULL) ends up in of_clk_get_hw(np, 0, NULL), and of_parse_clkspec() only consults "clock-names" when a name was passed, so the index stays 0 for all four entries. Every entry therefore ends up holding a handle to the *first* clock of the DT "clocks" property, i.e. ACLK_NPUn. Nothing fails: probe succeeds and the driver believes it owns four different clocks. The consequence is that rocket_device_runtime_resume() prepares and enables the AXI clock four times, while hclk, pclk and - most importantly - the NPU compute clock ("npu", SCMI_CLK_NPU on RK3588) are never prepared or enabled by this driver at all. The NPU still works only because the Rockchip power-domain driver sets GENPD_FLAG_PM_CLK and its attach_dev() callback walks the device node with of_clk_get() and adds every clock to the pm_clk list, so genpd happens to keep the remaining clocks running. The bug is therefore latent today, but it means the driver holds no reference to the clock that actually feeds the NPU, which stands in the way of any future frequency scaling (OPP/devfreq) work. Found on an Orange Pi 5 Plus (RK3588) by reading the live clock tree: /sys/kernel/debug/clk/clk_summary shows four "fdab0000.npu" consumer handles on aclk_npu0 (and likewise on aclk_npu1/aclk_npu2 for the other two cores), while hclk_npu0, pclk_npu_root and scmi_clk_npu have no "fdab0000.npu" consumer at all - their only consumers are the "npu@fdab0000" handles created by the power-domain driver via of_clk_get(). Set the ids explicitly, in the order mandated by the binding (Documentation/devicetree/bindings/npu/rockchip,rk3588-rknn-core.yaml): aclk, hclk, npu, pclk. After the change the driver holds one handle per distinct clock and clk_bulk_prepare_enable() covers all four. Note that this is a user-visible tightening for out-of-tree DTs: the old NULL-id requests resolved by index and succeeded no matter what "clock-names" contained, while the named requests fail probe with -ENOENT when one of the four names is missing. That is the right outcome for in-tree users - the binding requires exactly these four clock-names and rk3588-base.dtsi carries them on all three cores - but a DT that relied on the permissive lookup goes from silently running on the wrong clock handles to not probing at all, so record the change here where git log will find it. Fixes: ed98261b4168 ("accel/rocket: Add a new driver for Rockchip's NPU") Assisted-by: LLM sparse checkpatch Signed-off-by: Igor Paunovic Tested-by: Sidong Yang Tested-by: Diederik de Haas # NanoPC-T6 LTS, Nan= oPC-T6 Plus Reviewed-by: Sebastian Reichel Reviewed-by: Jiaxing Hu --- Same diff and the same message text (his copy re-wraps the lines) as 01/14 of Jiaxing Hu's RK3576 series, which carries this patch as well; whichever lands first, the other drops it. The trailers differ: Jiaxing asked me to drop his Signed-off-by from my own posting of it, as he did not pass this copy along, and Assisted-by is added because an LLM helped with the checks on this v2. drivers/accel/rocket/rocket_core.c | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/drivers/accel/rocket/rocket_core.c b/drivers/accel/rocket/rock= et_core.c index b3b2fa9ba645a..5dd260bacbff6 100644 --- a/drivers/accel/rocket/rocket_core.c +++ b/drivers/accel/rocket/rocket_core.c @@ -28,6 +28,10 @@ int rocket_core_init(struct rocket_core *core) if (err) return dev_err_probe(dev, err, "failed to get resets for core %d\n", cor= e->index); =20 + core->clks[0].id =3D "aclk"; + core->clks[1].id =3D "hclk"; + core->clks[2].id =3D "npu"; + core->clks[3].id =3D "pclk"; err =3D devm_clk_bulk_get(dev, ARRAY_SIZE(core->clks), core->clks); if (err) return dev_err_probe(dev, err, "failed to get clocks for core %d\n", cor= e->index); --=20 2.43.0 From nobody Thu Sep 24 16:07:15 2026 Received: from mail-wr2-f12.google.com (mail-wr2-f12.google.com [74.125.225.76]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5E595526ABA for ; Tue, 22 Sep 2026 08:01:38 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.225.76 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790064112; cv=none; b=AGH1zn4sjDOvzdSGIAWfR2ikkjYTPFxwbXYl2xGSPJKsTTKLaGQ+uOOOkwS9tPl+HAzuhQzBRatYqYG4+Xnp5h4EeLuqfzX1IZd8apAvnHYyH3XZOl1brzyJt7vZ7X5PoL0TpIXe+ehffDUzEVKAQvnEhfxtUd4a/8rAKdmiFKw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790064112; c=relaxed/simple; bh=4fAmXMpgpOU0t8UQuei0xUmAiikImQmBQYIZpvHFWWw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=eZg8t71FMqnVaVA15VpTcvj2bcShMsYMkQRJpgJbuCkLRq4carCmjHjy7CX4QfQXMj/fnfEyiPrEEvl5Qnt7JcADBlSm4Q3FUQpAkdFU0jIZPaM9TEVRmoLgM6ImwL8S7DkFy9YNqHxQ3qzHDsyMnOQi+t9M2Q0l1RrEiLs1rsM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=EfDMltdY; arc=none smtp.client-ip=74.125.225.76 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="EfDMltdY" Received: by mail-wr2-f12.google.com with SMTP id ffacd0b85a97d-482e1b30c94so340724f8f.1 for ; Tue, 22 Sep 2026 01:01:38 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1790064095; x=1790668895; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=gRVzD7/e6IROrFrMTSU32rtgAKkxa5C4zngEyHP7gSA=; b=EfDMltdYW1PUQlnAdKSDFTENYYpdIvKawKc8C9rs+DKMwvsCdWQIM9NgveLCdQ9D2b 2qUmsrkWVEAEAJPeoNx90PUOVkAG28NIOTcpnt2RBMUwCxnG0SRKQ32mJk0EyshQ+5w/ TOt4oEGRM/sV5bJWj3flKbOvrJ8DWL53IuvRAV8ECGSzOBeIrCyfHjCJ7YVrogzDbeHZ 5QAXfiACoTfaIpiAXn85aPyaQAq435plR2smwwL+wTNkU9cS9yOZ3zbAZAddroAKYsVr POWc/1JmzFAxmGTLYggl3VuCNU8te+IWKuSkoUGEicDezAZlnopKMNX3ah1wdR3n1Njo +2aw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790064095; x=1790668895; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=gRVzD7/e6IROrFrMTSU32rtgAKkxa5C4zngEyHP7gSA=; b=islXztIGKlgd6kpekAAg5XD+NPTruHupxzk3mLUvRVvge7H6cZ2zFRDN8jP5i8dPAO eZYE49rFcu/XuuhSCCNcd6EC12wxWLdn68Om0bsvflhja2M3b7HzxvjyXChBBHiT7w8w HMuQ459qy/esB4/paTuIDccnNt/EjTrybl7IwBSWonnaV2xuRqHxwaJqaO49p674i/QC wXB9Qumy3+qOMIrmW1/cH4UlswjN5vyPWLKW1adIo4bB/xnVY4LQGhm2ERe+n6cmBEwm Y2vm59Dafpj86QKRIWNN84B/1wWYTvKyYmgKcF1nFkCZ6YGiEmvqd23QyrPckXIa2xJk eyFA== X-Forwarded-Encrypted: i=1; AKwUvBzDTW2VsWr9CgLhc/7p+Ow684EnJf/ijftXcxsihZzvWihimBcqKf4N1YyQroG/qCha4hc7YNGQhKbdSwc=@vger.kernel.org X-Gm-Message-State: AFuF++mr4OXJ12gjiNkHoGwarg9ZZRlqGSo0UYzLPzjAIVaNpn36atfS wyLtxTe9ebrbK2e+CUGiqQ7+AtQmJLBVC4uU9JrB7Mu6S8UlInDnKV/h X-Gm-Gg: AYBFou0EiMn6Xt744RS4AXqtt0ZgZWJgFDYul6LKCvrQAHwzW4sRSTttxXiVgRlX5fu W4eQG5X4YQsKS6aX71ggiRj0rvOJ7p38nBEAQMynMGDNJNhaGpx3Wc+OmgrX72nx2yj58S36QuK aFK3etOrFxNSNut0SFkc/GxVcIXeVJlwVWp+0wsOies0ZTF9U+OPav8thFyT6bKurJHJ7hUJtDE EEf6iv/HTGlQamcEV1/uYiYOrkXl6bzg3kwlVlWneAlTXSujEhYZ9AMJa8d6hQmz+RoB8oPYgP1 4dpBNu9tUmXOQDp3GoYndLwy5G/7/QXnn7ou60xXpwO3f1VeD5fq+7IiEbOMkJAlCNWJ2osCCUq EPIB4poI1UXoS4ioRlRTgkaSV1sp22KWeuX4TCnOCJ6SVYk689GlwPEl8Ad5HsfZv+fgdggBvB/ LMg+7yghmLLcZPpvcp9YI4d9pl/Upek2zZhBcM89HMRehKfYrrOT0kNSDaArJhURv1fuPlZHVtv 0mLenlKvRCzChBSo5r4LE3jeY3dFJ96GgSv7SrcmPlZaS6cgbLQl8nxjfjNv6R2CCvAbOEX+kTf BKc= X-Received: by 2002:a05:600c:609b:b0:49c:f9b8:bae0 with SMTP id 5b1f17b1804b1-49fc7f2bdc8mr199260285e9.2.1790064093776; Tue, 22 Sep 2026 01:01:33 -0700 (PDT) Received: from OrangePi5-Plus.BB-HOME (20014C4E1B80530056971C6280202175.dsl.pool.telekom.hu. [2001:4c4e:1b80:5300:5697:1c62:8020:2175]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-49fdaaf97e2sm18248625e9.2.2026.09.22.01.01.32 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 22 Sep 2026 01:01:33 -0700 (PDT) From: Igor Paunovic To: Tomeu Vizoso , Oded Gabbay , Heiko Stuebner Cc: Rob Herring , Krzysztof Kozlowski , Conor Dooley , Jeff Hugo , Robert Foss , Sidong Yang , Diederik de Haas , Sebastian Reichel , Jiaxing Hu , Nicolas Dufresne , Jonas Karlman , Guangshuo Li , =?UTF-8?q?H=C3=BCseyin=20BIYIK?= , dri-devel@lists.freedesktop.org, linux-rockchip@lists.infradead.org, linux-arm-kernel@lists.infradead.org, devicetree@vger.kernel.org, linux-kernel@vger.kernel.org, Igor Paunovic Subject: [PATCH v2 06/11] dt-bindings: npu: rockchip: allow DVFS and thermal properties Date: Tue, 22 Sep 2026 10:01:09 +0200 Message-ID: <20260922080114.44662-7-royalnet026@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260922080114.44662-1-royalnet026@gmail.com> References: <20260922080114.44662-1-royalnet026@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" The three NPU cores on the RK3588 are fed by a single clock and a single supply, and the firmware accepts a fixed set of rates for that clock. Describing those rates as an operating-points-v2 table is what lets a driver scale the NPU instead of leaving it at whatever rate the bootloader set, so allow the property on the core node. Throttling the NPU from a thermal zone needs a core node to be usable as a cooling device, so allow #cooling-cells too. The OPP table belongs on every core, with opp-shared: the cores have no clock of their own, and one shared table for one shared clock is the same shape a CPU cluster uses. #cooling-cells goes on one core only, the one a thermal zone's cooling map names, because the cores cannot be throttled independently. Naming one representative node for a shared frequency domain is the established shape, as in "Cpufreq cooling device on CPU0" in Documentation/devicetree/bindings/thermal/thermal-cooling-devices.yaml. The schema cannot enforce which core carries #cooling-cells, because all three cores share a compatible string and a node name pattern, so that stays a devicetree convention. That is the same situation as for CPU cooling, where cpus.yaml does not restrict #cooling-cells to cpu@0 either. The example gains #cooling-cells; the operating-points-v2 property is exercised by the RK3588 devicetree later in this series. Assisted-by: LLM checkpatch dt_binding_check Signed-off-by: Igor Paunovic Acked-by: Conor Dooley --- v2: the text on where the properties go is rewritten for opp-shared on all three cores, following Nicolas Dufresne's review of v1 3/7. The only change to the schema file is the wording of the #cooling-cells description; the constraints and the example are unchanged. Conor, your Ack is kept on that basis; please say if it no longer holds. .../bindings/npu/rockchip,rk3588-rknn-core.yaml | 10 ++++++++++ 1 file changed, 10 insertions(+) diff --git a/Documentation/devicetree/bindings/npu/rockchip,rk3588-rknn-cor= e.yaml b/Documentation/devicetree/bindings/npu/rockchip,rk3588-rknn-core.ya= ml index caca2a4903cd1..beba1896156f5 100644 --- a/Documentation/devicetree/bindings/npu/rockchip,rk3588-rknn-core.yaml +++ b/Documentation/devicetree/bindings/npu/rockchip,rk3588-rknn-core.yaml @@ -42,6 +42,13 @@ properties: - const: npu - const: pclk =20 + "#cooling-cells": + description: + Present on one core only, the first, which stands for the shared NPU + clock as a cooling device. The other cores have no clock of their own + and cannot be throttled independently of it. + const: 2 + interrupts: maxItems: 1 =20 @@ -50,6 +57,8 @@ properties: =20 npu-supply: true =20 + operating-points-v2: true + power-domains: maxItems: 1 =20 @@ -100,6 +109,7 @@ examples: clocks =3D <&cru ACLK_NPU0>, <&cru HCLK_NPU0>, <&scmi_clk SCMI_CLK_NPU>, <&cru PCLK_NPU_ROOT>; clock-names =3D "aclk", "hclk", "npu", "pclk"; + #cooling-cells =3D <2>; interrupts =3D ; iommus =3D <&rknn_mmu_0>; npu-supply =3D <&vdd_npu_s0>; --=20 2.43.0 From nobody Thu Sep 24 16:07:15 2026 Received: from mail-wm2-f13.google.com (mail-wm2-f13.google.com [74.125.225.141]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8540951C05C for ; Tue, 22 Sep 2026 08:01:39 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.225.141 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790064122; cv=none; b=s3DYuFq+FvhfjWazEhbxy9mvOXhjoyFbDwxgJFIH/64cHaBoPpSe6QAZjGom9No9x4P7bPIGAU19HPQjsZcebeX4osDxNcN5fTxu2unWZJeF3ZXGah+PKYh+2EvcfgXmwTiMRFppOX81j4TulryCruliqp5FnqPm2ytUY0R8GNU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790064122; c=relaxed/simple; bh=bbyeP03gBPHvNx+YpSbids6Ib61D9Tlgwed4YbVP+I8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=P/s6cpH3t3gczFkSPfAOqtk8XQszknHL2315S4Qr/Mo5y/KKjh5O+phtWXXbMvV8MYTxrXxOQmw+1ZEKCcD2wmyHG204w/cPXA19Qb9el05Y4obAXnFVra54NgNh0mgWi4DvEI9h60ns3D+3kANv07b70K5aSxCbBqR1iC53DO8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=dhla/AT2; arc=none smtp.client-ip=74.125.225.141 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="dhla/AT2" Received: by mail-wm2-f13.google.com with SMTP id 5b1f17b1804b1-49fbca514f8so2712765e9.2 for ; Tue, 22 Sep 2026 01:01:38 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1790064095; x=1790668895; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=C/lMPPdyVDY8UDcv/G9hmnqUs4LxNskElZuOnfztMlQ=; b=dhla/AT2tMVXZWvnnkIAEbrVq8aMTS59v0WcUwEAWajKrjdMM0wIghac5M+bFqD1Bc OJ4+nMOmI3daeW4hV9YOATuNVdPNNCZTbJOPV8NANwxV2d9u1Lrx6ba/v8ntObHcLYRi YyvcVBalRB3e5SHp6og9VXCwUEZRbkpMClfVOCE+O7HeJHV9qTS1paa5Xv9VEqW6eILB RhHQeyzfVIcFPY42rwhvcqDI/UX4x9Jx70gn1kNFLQ1AWbNxXJ0QNuu38R7KXi+PgI06 SvWkcyP8aum0nNnqxp1fZuD+E9Uu5l2DPxKt1nP9pMwmOyuTdTkdtFzhcJcEE+KuWdCE oDQg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790064095; x=1790668895; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=C/lMPPdyVDY8UDcv/G9hmnqUs4LxNskElZuOnfztMlQ=; b=vqxScbJ/BvtHQJ83fyNCD6r1ezIvOTgTtoEfNEw39emgMJ9JDINrCanWGUeiyvmenj DC6CIyqgXlZvDXwX6hN6pMUP2K3WfCuqksRDtLvP1+MuO5Q/ns/fAlaagc1VbzTOUTiJ E9n3xf6UwWeEcTZZSFP0ECPaB6+Dc0VSncq4OK0Vnjc8uMRci9kQpw2kIo1niE2QVklx 4zuCZ1Y0RK+j4e0BXJYxpCM8V2LZ7iNd6WFNAHGDmiHc7v8BrhZB6VGnFgZQZAMF+DmK RRpkDf6zi+7354oSIcKRHMPlWKQH3wto8B8tFYStXKWNthVc/RSd6GUefXLDdUCETgbT wesA== X-Forwarded-Encrypted: i=1; AKwUvBxABhrr4Dl5uYivHKFZvUF7U/SGWm42qQ+3PRxBwTfcyg9Y7a5ahL8lsIplsanaikxGkAxmg2fUQbDCu+w=@vger.kernel.org X-Gm-Message-State: AFuF++llmDHHvhc1EziU8DwqoqGbBN73Foty5LI37EiYuXjYObiTqGvX ecY9m0ZcAzpVWrSrIPbSRxUVEhIxL1oLI2mOvY2PKkpW1egPzXMOo7dN X-Gm-Gg: AYBFou0h99SmSFNttps22clc+E/7C0nitWWtRX9ZHeTkvHfWIlJ9YIWJLOBi9tKaa79 rA1u3xsW+BQuHmAvg62ZuYPwCQGaKTmsowjyBbUHNISLxVGJ7bsebntOtBNWDwCKEEbihSEIDxC pr70iGQQPPcDO4o5sjvG2nTePTSgfABlC/dYOO6YdMaSd8JlhUelqQ6YUeFg4uGnIXJz4A/GYCo zNRpLPaBT63dNlpOg9/gVwBwmmPJ3lvRXnq4tZxAk6d21vSjeLJ6chUpnR3rT7Yok/nNUEgGe9J bj0MCzCJDSFA9RoJFECScgRZuVvKk5pYLsJQdeQKD2FL9A+XzDFBz5DR70OE5L1sS0mEvRYRQ4G ErK5pktsyPBUaE6zsFE1NE6gUR0TPJnLzmvduj0AMTT/tqUnt8vw1g72dcpQ+AkPLv0WDFZ9DbY FrGdJ7R0Yx69bWc0PHQQpXOLed3MZfQvkNcVl/s+tj2OwlKRzSfkz60zmlwi5HSTz0wt8fLT8aL W6lXfte6v+n5unjxBp6MX44/7zCJLBIDZnd6AVWrgmkN8lmHFxcElr+BKQIP5x6GJoDKMjFn0dB Y2uP4XmdWwM1Ww== X-Received: by 2002:a05:600c:354e:b0:49e:479b:c13b with SMTP id 5b1f17b1804b1-49fc7dff668mr165633455e9.1.1790064095154; Tue, 22 Sep 2026 01:01:35 -0700 (PDT) Received: from OrangePi5-Plus.BB-HOME (20014C4E1B80530056971C6280202175.dsl.pool.telekom.hu. [2001:4c4e:1b80:5300:5697:1c62:8020:2175]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-49fdaaf97e2sm18248625e9.2.2026.09.22.01.01.33 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 22 Sep 2026 01:01:34 -0700 (PDT) From: Igor Paunovic To: Tomeu Vizoso , Oded Gabbay , Heiko Stuebner Cc: Rob Herring , Krzysztof Kozlowski , Conor Dooley , Jeff Hugo , Robert Foss , Sidong Yang , Diederik de Haas , Sebastian Reichel , Jiaxing Hu , Nicolas Dufresne , Jonas Karlman , Guangshuo Li , =?UTF-8?q?H=C3=BCseyin=20BIYIK?= , dri-devel@lists.freedesktop.org, linux-rockchip@lists.infradead.org, linux-arm-kernel@lists.infradead.org, devicetree@vger.kernel.org, linux-kernel@vger.kernel.org, Igor Paunovic Subject: [PATCH v2 07/11] arm64: dts: rockchip: rk3588: add an OPP table for the NPU Date: Tue, 22 Sep 2026 10:01:10 +0200 Message-ID: <20260922080114.44662-8-royalnet026@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260922080114.44662-1-royalnet026@gmail.com> References: <20260922080114.44662-1-royalnet026@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" The NPU compute clock is driven by the firmware, which only accepts one of the rates in its own PVTPLL table: 300, 400, 500, 600, 700, 800, 900 and 1000 MHz through the PVTPLL, plus 200 MHz off GPLL. Anything else comes back as SCMI_INVALID_PARAMETERS, and that refusal never reaches the caller: the clock framework does not look at what the clock's set_rate returns, so clk_set_rate() reports success and the clock stays where it was. The table therefore has to name those rates exactly rather than describe a range. 200 MHz is included even though the vendor table stops at 300, because mainline pins the cores there with assigned-clock-rates, the firmware's table names 200 MHz exactly, on its GPLL path, and that is the rate the NPU boots and idles at. Its voltage is the same 700 mV the vendor uses for 300 MHz, so it is conservative. The voltages are the vendor's, and the upper half of the table matches the GPU table in this file step for step: 700 MHz at 700 mV, 800 at 750, 900 at 800, 1000 at 850. There is no PVTM or binning here, for the same reason the GPU table has none: mainline uses conservative worst-case voltages instead of per-chip nvmem data. The table is marked opp-shared and referenced from all three cores. They have one clock and one supply between them and cannot be scaled independently, and that is what opp-shared describes: one table for one clock, the way a CPU cluster shares its table. The full SoC range is described rather than a per-board subset, so that a board which cannot cool the upper rates drops them in its own .dts with a /delete-node/ on the OPP it does not want. A board may only delete OPPs that way, never invent intermediate ones: a rate that is not in the firmware's table is refused by the firmware, but the kernel never learns of it, so an invented OPP would be refused while the kernel went on reporting it as set. rk3588j.dtsi does not include this file; it carries its own derated tables for the CPU clusters and the GPU, and it gets no NPU table here. That is deliberate. The J part is rated lower than the rates in this table and none of it can be measured on the hardware this was written on, so inventing a derated NPU table would be guessing. Its NPU node stays disabled, so nothing binds and the cooling map added later in this series is simply never resolved. The same rates and voltages were arrived at independently by Nicolas Dufresne in a proof of concept that was never posted to the list; his version differs in that it marks 200 MHz as opp-suspend and drops the assigned-clock-rates pins. Link: https://gitlab.collabora.com/nicolas/linux/-/commits/rock5b-npu-poc-4 Assisted-by: LLM checkpatch dtbs_check Signed-off-by: Igor Paunovic --- v2: - opp-shared, and the table referenced from all three cores (Nicolas). - The opp-suspend paragraph is gone. In the v1 thread I said v2 would argue that the driver already puts the device back at its boot rate; that is again driver behaviour used as a devicetree argument, which is what Nicolas objected to, so I am not making it. Whether opp-suspend at 200 MHz describes the hardware is a question for the DT maintainers, in the cover letter. - "give a driver nowhere to return to" is gone for the same reason. - The paragraph on the table being inert until the driver patch is gone (Nicolas). - New: the firmware's refusal of a rate is not reported back through the clock framework. Found by reading clk_change_rate() in drivers/clk/clk.c after a test that requested a rate outside the table. arch/arm64/boot/dts/rockchip/rk3588-opp.dtsi | 54 ++++++++++++++++++++ 1 file changed, 54 insertions(+) diff --git a/arch/arm64/boot/dts/rockchip/rk3588-opp.dtsi b/arch/arm64/boot= /dts/rockchip/rk3588-opp.dtsi index b5d630d2c879f..59ecaef5101da 100644 --- a/arch/arm64/boot/dts/rockchip/rk3588-opp.dtsi +++ b/arch/arm64/boot/dts/rockchip/rk3588-opp.dtsi @@ -151,6 +151,48 @@ opp-1000000000 { opp-microvolt =3D <850000 850000 850000>; }; }; + + npu_opp_table: opp-table-npu { + compatible =3D "operating-points-v2"; + opp-shared; + + opp-200000000 { + opp-hz =3D /bits/ 64 <200000000>; + opp-microvolt =3D <700000 700000 850000>; + }; + opp-300000000 { + opp-hz =3D /bits/ 64 <300000000>; + opp-microvolt =3D <700000 700000 850000>; + }; + opp-400000000 { + opp-hz =3D /bits/ 64 <400000000>; + opp-microvolt =3D <700000 700000 850000>; + }; + opp-500000000 { + opp-hz =3D /bits/ 64 <500000000>; + opp-microvolt =3D <700000 700000 850000>; + }; + opp-600000000 { + opp-hz =3D /bits/ 64 <600000000>; + opp-microvolt =3D <700000 700000 850000>; + }; + opp-700000000 { + opp-hz =3D /bits/ 64 <700000000>; + opp-microvolt =3D <700000 700000 850000>; + }; + opp-800000000 { + opp-hz =3D /bits/ 64 <800000000>; + opp-microvolt =3D <750000 750000 850000>; + }; + opp-900000000 { + opp-hz =3D /bits/ 64 <900000000>; + opp-microvolt =3D <800000 800000 850000>; + }; + opp-1000000000 { + opp-hz =3D /bits/ 64 <1000000000>; + opp-microvolt =3D <850000 850000 850000>; + }; + }; }; =20 &cpu_b0 { @@ -188,3 +230,15 @@ &cpu_l3 { &gpu { operating-points-v2 =3D <&gpu_opp_table>; }; + +&rknn_core_0 { + operating-points-v2 =3D <&npu_opp_table>; +}; + +&rknn_core_1 { + operating-points-v2 =3D <&npu_opp_table>; +}; + +&rknn_core_2 { + operating-points-v2 =3D <&npu_opp_table>; +}; --=20 2.43.0 From nobody Thu Sep 24 16:07:15 2026 Received: from mail-wm2-f13.google.com (mail-wm2-f13.google.com [74.125.225.141]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CC3E84D598A for ; Tue, 22 Sep 2026 08:01:42 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.225.141 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790064117; cv=none; b=lGakGEhCeuKYNQ/253Ik2guR//IsXk6h48p1JkInOr1uQPA6RKGVsHa0JwenteL87qLSFmMpXac1Vz+KvL6pod56gGY8fJZ5ICoAYjv/odkvCC6GDukKtVb1QzkcNWcknqE45BhWNva8jSZBwbJCiQqWupr/n6Q0wfY+lq2LiWo= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790064117; c=relaxed/simple; bh=9ziRMOlF0fOTNkr1dL+Ne6qUpDk020zzCjQpb8OrhIs=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=dK97GVObORX5uTQbdnUdhyE5rVo2kHXdIfe3k8l8SBkeQsFzblUZ5snSaBbua+EG/dRiXQE161LU8qXKT2ASSNHR3FhrO0IJw4/XNKTePKnCzKR4j37fnKrl1KrXiHgFzi/ZUvXVtjgErpccNz51W8dviVYqaRUmJYu0PheyHG0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=OxyHqClv; arc=none smtp.client-ip=74.125.225.141 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="OxyHqClv" Received: by mail-wm2-f13.google.com with SMTP id 5b1f17b1804b1-49fbca514f8so2712845e9.2 for ; Tue, 22 Sep 2026 01:01:40 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1790064097; x=1790668897; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=YqiXccAHsUrIMDsLkLXR6P0dTgO5nQzYlI6kaXLWWHg=; b=OxyHqClvVDO/UmA1ObR7r4QdDC/7EQ830N3NMhjJrewQ0dIdYs9zq9gQMAGVOS6RW4 vUyhjDfInPLzpwI+xDzCXCzRqUeb7x6LJE/S9aDiCMLfPXZf2nCZZhi+HZxo9u5LgRNn 4PVJqiynscCNMuLDL0u31m6AAJU3x8dT+uTEhzHNVi2nJOdVTHLzPfvx6/rsaa34F4s0 92WE4RT+ALOMmFDIXhN3xXoR49A6SS5FWyJ7BJzL4D25Xd7q3MaygyyrK4ET+/ALdaTV dOLKTBFqG5LIPz1ZgOndwCSwy5lOdYJvS9QwqUmO1x8caizm1WFgEjTE6wmo8OyiVJS4 6Z9A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790064097; x=1790668897; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=YqiXccAHsUrIMDsLkLXR6P0dTgO5nQzYlI6kaXLWWHg=; b=lGzsfEOlzsm/rKyvhVjx1rVFlqKc52P2aykCRC3v1jrMFtD7jcXmQr0aOUcItkoJ+w DihIwcsq4PZqNRir4rrAUPinC8jTwUFzb9I2NQoUq7edr+PeoQTheSkzDtJ0VWENazoF v+Flzn5HZbZBZKqhUVS2McFvOJGhiUYD7w4DW+j8laWEnWMn9grjQvIdkyVrwblV7gE2 1oS0cBrgccC37QxhAKDQmvjBewb53roLFesKNEo3cKdeOOWH4+v+QjJf4u56+nwKWQ+2 4LFzsMrY1Ynhktebbz0qf0oaiYOUkblE1KN05YrW5kZzHhcA25y+gGsxwa3mPMG6Wpzl /smA== X-Forwarded-Encrypted: i=1; AKwUvBztvPDw12DRhIv9yJLXZ1ETvOUrBb51+Kw642rlU9eGLRdOjvIyJVaqRiadBkwd8yryWsBm0E9X3xPC3wM=@vger.kernel.org X-Gm-Message-State: AFuF++ldbLyF6pQ82gCv2V5ftv5eS/FJFUUQ6au+E0tGrGW3q44iejbM H7t2qgxjGtUMxI9sAaW8b6RIhmcih521QrhehawUsb0/LWsqqTUD9VL7 X-Gm-Gg: AYBFou3/zd6RLTAC6s82zRj4At62bevMa+ptEbritjiYvMNHLHXNPBIhNB40cgwV2wD maC0tiX8y+pVCqOuTo63D4vZDnne9pXYFR7hlf6xXe1mupiuoOUMK2nq+HSJptDrFM1Gu+R6GMQ wSWGMcg6jN1bUq3FLHWkraMI/OA4ix7akEKa1AjiqQeZWkwNLgGeEaEsQAZJZY6Nl/n+QX//bqc M7wzXISSYlNmnHLUJN2DSKvULwlW7SmJuGqxLvVrN9R6Q3LOM04vYEmPT2npsGneiXKwZ5LMJYU XgZy/N/YjCMA7Wi017WuIC7iz9/24TARpEAJaH7CzVs6CE5hLp+daaCPbfU1Dw3XXF95nauwc6A k0QfIQJzyOgDXjjsP6kVLBq2VeffWbOQqSKcjOymbIjcPeW9h/ft8wMLCO9xABeNxHPSPaJeuZo 1ansVpGQ7k1jwUUtYHfzj0ZEExp0ya2nGJg6c/5DUxv2cY3u3gXxEka9jj1pu/wt34HSvOpDnrK jz/jDqFGcbc+O3OP1RSZqdS46nFSyIn40VolqjGWXsMu+Etg7UgcNNmNI1WBf/n/bHsm2Huvhcc 7vE= X-Received: by 2002:a05:600c:3b99:b0:49e:6683:d227 with SMTP id 5b1f17b1804b1-49fc7bbfe4bmr199180665e9.0.1790064096621; Tue, 22 Sep 2026 01:01:36 -0700 (PDT) Received: from OrangePi5-Plus.BB-HOME (20014C4E1B80530056971C6280202175.dsl.pool.telekom.hu. [2001:4c4e:1b80:5300:5697:1c62:8020:2175]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-49fdaaf97e2sm18248625e9.2.2026.09.22.01.01.35 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 22 Sep 2026 01:01:36 -0700 (PDT) From: Igor Paunovic To: Tomeu Vizoso , Oded Gabbay , Heiko Stuebner Cc: Rob Herring , Krzysztof Kozlowski , Conor Dooley , Jeff Hugo , Robert Foss , Sidong Yang , Diederik de Haas , Sebastian Reichel , Jiaxing Hu , Nicolas Dufresne , Jonas Karlman , Guangshuo Li , =?UTF-8?q?H=C3=BCseyin=20BIYIK?= , dri-devel@lists.freedesktop.org, linux-rockchip@lists.infradead.org, linux-arm-kernel@lists.infradead.org, devicetree@vger.kernel.org, linux-kernel@vger.kernel.org, Igor Paunovic Subject: [PATCH v2 08/11] accel/rocket: restore the NPU clock boot rate before powering the cores down Date: Tue, 22 Sep 2026 10:01:11 +0200 Message-ID: <20260922080114.44662-9-royalnet026@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260922080114.44662-1-royalnet026@gmail.com> References: <20260922080114.44662-1-royalnet026@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" The compute clock is generated by a PVTPLL that lives inside the NPU power island. Powering an island up while that clock is above the rate the bootloader left it at does not work: the domain never acks the power-on, and the first register access into it afterwards takes an asynchronous SError. So the rate has to be back down before the last core goes away. Nothing in the driver raises the clock today, which makes this a no-op on its own, but it is the guard that has to be in the tree before anything does, and the next patches do. The .shutdown hook is the same guard for the handover: once devfreq is driving the clock, a kexec would otherwise pass the raised rate to the next kernel, which powers the islands up before it looks at it. What this cannot do is rescue a rate it did not set - the rate read at probe is taken as the boot rate whatever it is. The rate is read at probe rather than hardcoded. Mainline pins the RK3588 cores at 200 MHz with assigned-clock-rates, but that is a devicetree property, not a property of the hardware, and a SoC whose devicetree does not set it would be left running at a rate this driver had invented. All three cores share the clock, so only the last core to suspend may lower it; the others just drop the count. Lowering it is safe with the islands already down as long as the boot rate is one the firmware serves from GPLL, which on the RK3588 is the 200 MHz the devicetree pins: for that rate the firmware writes only CRU clock selectors, never a register inside the NPU. Assisted-by: LLM sparse checkpatch Signed-off-by: Igor Paunovic --- v2: v1 of this patch kept a struct clk handle in struct rocket_device, taken from the devres of the first core to probe, and used it from the runtime suspend of whichever core went down last. Unbinding the cores freed the handle underneath it: KASAN reported a slab-use-after-free in clk_set_rate() during a ten-round unbind/rebind test on this board after v1 was posted. No handle is kept any more; the callback uses the handle of the core it runs for, which is bound for as long as the call lasts. "A reboot" is dropped from the kexec sentence and the comment: whether the clock selectors survive the global reset the firmware does on reboot has not been checked. The argument that lowering the rate is safe is now limited to a boot rate the firmware serves from GPLL, which on the RK3588 is the 200 MHz the devicetree pins. drivers/accel/rocket/rocket_core.c | 10 ++++++ drivers/accel/rocket/rocket_device.h | 15 +++++++++ drivers/accel/rocket/rocket_drv.c | 47 ++++++++++++++++++++++++++++ 3 files changed, 72 insertions(+) diff --git a/drivers/accel/rocket/rocket_core.c b/drivers/accel/rocket/rock= et_core.c index 5dd260bacbff6..c736537cf28f6 100644 --- a/drivers/accel/rocket/rocket_core.c +++ b/drivers/accel/rocket/rocket_core.c @@ -12,6 +12,7 @@ #include =20 #include "rocket_core.h" +#include "rocket_device.h" #include "rocket_job.h" =20 int rocket_core_init(struct rocket_core *core) @@ -36,6 +37,15 @@ int rocket_core_init(struct rocket_core *core) if (err) return dev_err_probe(dev, err, "failed to get clocks for core %d\n", cor= e->index); =20 + /* + * Record what the compute clock was running at before anything here + * touched it, on the first core to probe. Reading it rather than + * hardcoding a rate keeps this working on a SoC whose devicetree does + * not pin the clock with assigned-clock-rates. + */ + if (!core->rdev->npu_boot_rate) + core->rdev->npu_boot_rate =3D clk_get_rate(core->clks[2].clk); + core->pc_iomem =3D devm_platform_ioremap_resource_byname(pdev, "pc"); if (IS_ERR(core->pc_iomem)) { dev_err(dev, "couldn't find PC registers %ld\n", PTR_ERR(core->pc_iomem)= ); diff --git a/drivers/accel/rocket/rocket_device.h b/drivers/accel/rocket/ro= cket_device.h index abb88a254e569..ba7c977cd6951 100644 --- a/drivers/accel/rocket/rocket_device.h +++ b/drivers/accel/rocket/rocket_device.h @@ -22,6 +22,21 @@ struct rocket_device { unsigned int num_cores; /* Slot capacity (DT core count); slots with a NULL .dev are free. */ unsigned int max_cores; + + /* + * The cores have no clock of their own: one clock feeds all of them, + * so any core's handle refers to the same thing. No handle is kept + * here: each one belongs to the devres of the core that asked for it + * and dies with that core's unbind, while this structure outlives any + * single core. Whoever needs the clock uses the handle of the core it + * was called for, which is bound for as long as the call lasts. + * + * npu_boot_rate is the rate the clock was left at before the driver + * touched it, and active_cores counts the cores that are runtime + * resumed right now. + */ + unsigned long npu_boot_rate; + atomic_t active_cores; }; =20 struct rocket_device *rocket_device_init(struct platform_device *pdev, diff --git a/drivers/accel/rocket/rocket_drv.c b/drivers/accel/rocket/rocke= t_drv.c index b9b36c578db20..8f03de1af488c 100644 --- a/drivers/accel/rocket/rocket_drv.c +++ b/drivers/accel/rocket/rocket_drv.c @@ -297,6 +297,30 @@ static int find_core_for_dev(struct device *dev) return -1; } =20 +/* + * Put the compute clock back where the bootloader had it. The cores share + * this clock, so this is only correct once none of them is running any mo= re. + * + * Lowering the rate is safe with the power islands down as long as the bo= ot + * rate is one the firmware serves from GPLL, which on the RK3588 is the + * 200 MHz the devicetree pins: for that rate the firmware touches only the + * CRU clock selectors, none of the NPU's own registers. + */ +static void rocket_npu_restore_boot_rate(struct rocket_core *core) +{ + struct rocket_device *rdev =3D core->rdev; + int err; + + if (!rdev->npu_boot_rate) + return; + + err =3D clk_set_rate(core->clks[2].clk, rdev->npu_boot_rate); + if (err) + dev_warn(core->dev, + "failed to restore the NPU boot rate of %lu Hz: %d\n", + rdev->npu_boot_rate, err); +} + static int rocket_device_runtime_resume(struct device *dev) { struct rocket_device *rdev =3D dev_get_drvdata(dev); @@ -312,6 +336,8 @@ static int rocket_device_runtime_resume(struct device *= dev) return err; } =20 + atomic_inc(&rdev->active_cores); + return 0; } =20 @@ -328,6 +354,9 @@ static int rocket_device_runtime_suspend(struct device = *dev) =20 clk_bulk_disable_unprepare(ARRAY_SIZE(rdev->cores[core].clks), rdev->core= s[core].clks); =20 + if (atomic_dec_and_test(&rdev->active_cores)) + rocket_npu_restore_boot_rate(&rdev->cores[core]); + return 0; } =20 @@ -336,9 +365,27 @@ EXPORT_GPL_DEV_PM_OPS(rocket_pm_ops) =3D { SYSTEM_SLEEP_PM_OPS(pm_runtime_force_suspend, pm_runtime_force_resume) }; =20 +/* + * A kexec hands the next kernel whatever rate is set here, and that kernel + * will power the islands up before it looks at the clock. + */ +static void rocket_shutdown(struct platform_device *pdev) +{ + struct rocket_device *rdev =3D dev_get_drvdata(&pdev->dev); + int core; + + if (!rdev) + return; + + core =3D find_core_for_dev(&pdev->dev); + if (core >=3D 0) + rocket_npu_restore_boot_rate(&rdev->cores[core]); +} + static struct platform_driver rocket_driver =3D { .probe =3D rocket_probe, .remove =3D rocket_remove, + .shutdown =3D rocket_shutdown, .driver =3D { .name =3D "rocket", .pm =3D pm_ptr(&rocket_pm_ops), --=20 2.43.0 From nobody Thu Sep 24 16:07:15 2026 Received: from mail-wm2-f12.google.com (mail-wm2-f12.google.com [74.125.225.140]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 204E0527599 for ; Tue, 22 Sep 2026 08:01:42 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.225.140 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790064129; cv=none; b=jJosED7kGWdODFU8hceaEkk5RvbKr6MAhd7zZgnlyCunKcOfFJNbnMpy7u43Vhn4Xsi1NhK0BrFywNr52XObKHFWAUrR63bnKsN36eL2o3U6WHb/g0KuCggJvKHC1la19YIbZD9zXkz3naszZtlJArTpmmiPVMwOtRqLFuudT70= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790064129; c=relaxed/simple; bh=4erbYgx+ZbBX8tKcANJBFfSKinrXJZH7FXiJj9SS9Ew=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=g48xJwHMotVdapIeLPqsZmsCO1aDy0vk4WKRgobpTMJ429obzcSc61EA7mMIm/Vtjtri3Qxr185RH/AGmkb5Gnk1WsCH5iIGM0vy0iH49dgenc+6DuGkDViUC30M/pW86NfJmggOBT7lMIxUyZfW2rhOHRTaZRGsoBJR73FlJCc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=NSvJ/UT4; arc=none smtp.client-ip=74.125.225.140 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="NSvJ/UT4" Received: by mail-wm2-f12.google.com with SMTP id 5b1f17b1804b1-49cfbdac7a1so1756475e9.1 for ; Tue, 22 Sep 2026 01:01:41 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1790064098; x=1790668898; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=I32qcJErlvaRW+Ft0guC5hzjNkhNg6mkhTsESIhHnNQ=; b=NSvJ/UT4eJ1An8Vfybf1psfXqwJGolij6/LXFWQGxWglmlgitpofi8R63EZgU2/CRr O6Zkdp2Ywe8EBSxWKVo3d29W6gow6cxftI0OSzgaWZGr8Wik0AU+hRSEfqKaGn3/hSLA NFsLgSkgmumHIiAlIRvJR5Bcwd6oj1XRtHihdXxRMZyPNomi3kFM3MsQSKs4sCLbWqBi bkyQiyt0gcovkS7czQMiZ7mrG4xEC+MPO80D0Tk3+mi2wySZV9Vt89bM1SA6u7AruqEL k4ckJZmZtqZjwtXYQnxBUS5jaXPPExih0ptX7M6KwAJDBeBCrVH4MC6xDZ6jOvh+wo4X MJqg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790064098; x=1790668898; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=I32qcJErlvaRW+Ft0guC5hzjNkhNg6mkhTsESIhHnNQ=; b=jAZQOJWP+HSbUVx2ExPaN6nyciLi12OpQsyA3s8eE8aMF/yCQckwreWlUwZiL1828i RxVhV0to0llAqvcJ3b+Q6KS03VMmX0JeedF1BRuo170ARFFD7WVldRz/AEMm+v+EOghS PPSS5Vqh3EmV/bNZojiW9OBMeuFYBji0M5/UPVYcGRlcp4/FSUlvRBPH5nPOFxu/GMg4 uIKKCcfpsbbh5THsuI3mxl8lT66co1gNpogTWucRJCgzwjP04EWgoCQtjG3ZruUf/EXX waWyp2aOPKWsaryP+kC6DVvVYOaBeaessiLnvUHQPFFXlUqE5fi/NJ6fC0gCZWj5hA6f je7g== X-Forwarded-Encrypted: i=1; AKwUvBz2x3/hxUujZdDcBw5QlILZ9YbbTDhh0sggebUbHfeb/1HEVNh40g7E/5w582mR1VcGDuQK8ctyVBuToCw=@vger.kernel.org X-Gm-Message-State: AFuF++lArv7bxBXv0j6PS9cdicsZSF9qG9bqkzUPzcXQIveM2Jv4dQ5z modyJ4U6i+ElQmpbSCHFVwWIy1W/BvT2DJRaqE3nPL9nsRSLzD8YPwsN X-Gm-Gg: AYBFou0m6Axf48HQ9KL+0Ve1FKgpiIkHx/5HTPCCXuK6bxtsJXgtT5fmnD31Z6XbUGG 959QySM/MbF0Yg/TEv4Ig4IBVbMUFW4GjQc1bGp6TtE8U6cEIBDJp1kBUmjVEl/TY8Bs7twub7Q wf+Prp6KtBSgXBz9CFdKwcOAeAqfrAGb+fchbIG4Ph4FnwSvFV0LKj/VYHLegbbASPiHyt2Hhmn z9ygZzRmyp0om1ZZCocccXHYdfmP2/1dqdCvr/qPzG6hsM1m+SJVIlIreXnrXo9RzmG896Udh20 P1s+Tfwc7+hbe7V2P8hqdon0Q/tEiIMiAa9faSga06K4sfLmiSlpOMqDhvm1pz8hbvR8+n1Aown biIgdesQCF/exGNfjF/vcBvmHw7cc5eJ1gbmhNrosMpmo3IpKNa6WAHbpZb4OoYQv6B/7fp+daB pkp97Wyxo3cFGhY5mjmPnWpo25A06FrDaEtCW8EXtNPpzTecMbspbtHqP/v7hzM7tD6yg0DDsTy NxVD5Hj/nsGcuGA7khSPRnxnU12GJYF52c8dJ3YjcBGUg0lJGKHM1nStOCrFTIA0Xnwqrphr+g/ Y8M= X-Received: by 2002:a05:600c:1c13:b0:49e:7186:f36e with SMTP id 5b1f17b1804b1-49fc7dc1e75mr208981335e9.1.1790064098083; Tue, 22 Sep 2026 01:01:38 -0700 (PDT) Received: from OrangePi5-Plus.BB-HOME (20014C4E1B80530056971C6280202175.dsl.pool.telekom.hu. [2001:4c4e:1b80:5300:5697:1c62:8020:2175]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-49fdaaf97e2sm18248625e9.2.2026.09.22.01.01.36 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 22 Sep 2026 01:01:37 -0700 (PDT) From: Igor Paunovic To: Tomeu Vizoso , Oded Gabbay , Heiko Stuebner Cc: Rob Herring , Krzysztof Kozlowski , Conor Dooley , Jeff Hugo , Robert Foss , Sidong Yang , Diederik de Haas , Sebastian Reichel , Jiaxing Hu , Nicolas Dufresne , Jonas Karlman , Guangshuo Li , =?UTF-8?q?H=C3=BCseyin=20BIYIK?= , dri-devel@lists.freedesktop.org, linux-rockchip@lists.infradead.org, linux-arm-kernel@lists.infradead.org, devicetree@vger.kernel.org, linux-kernel@vger.kernel.org, Igor Paunovic Subject: [PATCH v2 09/11] accel/rocket: add devfreq support Date: Tue, 22 Sep 2026 10:01:12 +0200 Message-ID: <20260922080114.44662-10-royalnet026@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260922080114.44662-1-royalnet026@gmail.com> References: <20260922080114.44662-1-royalnet026@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" The NPU has run at whatever rate the devicetree pinned it to since the driver was merged, which on the RK3588 is 200 MHz. The hardware reaches 1 GHz, and the firmware will change the rate on request, so let devfreq drive it from how busy the cores actually are. One devfreq device drives all of the cores, because they have one clock and one supply between them and cannot be scaled apart. It hangs off the first core in devicetree order that carries an OPP table, which on the RK3588, where all three cores reference the shared table, is rknn_core_0. The choice has to be fixed rather than "whichever core bound last": the devfreq device is named after that core, and a cooling map in the devicetree resolves against that core's node. core->index is the core's position among the core nodes, so the lowest index is the first node. The awkward part is that the clock is generated by a PVTPLL that sits inside the NPU power islands. An island powered up while the clock is above the rate the bootloader left never acknowledges the power-on, and the first register access into it afterwards takes an asynchronous SError. So before the rate goes up, every core is runtime resumed, and the references are held for as long as the clock stays raised: no island can transition while they are held, which is what makes a raised clock safe rather than merely unlikely to be caught out. The guard and the thing it guards arrive in the same commit, so no commit in the tree ever raises the rate without it. Unbinding any core takes the devfreq device down, and it comes back once every core is bound again: a rate change resumes all of them, so the device has to be whole for the guard to mean anything. Every walk of the core array here relies on that. The cost is that a boosted NPU does not power-gate individual cores. It is paid only above the boot rate; at the boot rate the references are dropped and runtime PM behaves as before. What that costs in milliwatts has not been measured on this board. A measurement, or a design that does not need the references at all, would both be welcome. Utilisation is aggregated as the maximum over the cores, not the sum: the rate has to satisfy the busiest core, and summing would report one saturated core out of three as a third of the load. Whether that is the right aggregation for a shared clock is a fair question for review. Unlike panfrost, panthor, lima and msm, the driver does not call devfreq_suspend_device() from runtime suspend. That call ends in cancel_delayed_work_sync() on the governor's worker, and the worker is what calls ->target(), which resumes every core: a runtime-suspend callback would wait for a worker that is waiting for that same callback. Instead ->target() returns early when the rate is unchanged, so a governor tick on an idle NPU costs one comparison and resumes nothing. System suspend is different: dpm_suspend() runs devfreq_suspend() over every registered devfreq before it walks the devices, from a path that holds none of this driver's locks. What the driver adds is lowering the rate and dropping the references from every core's ->suspend, not only the owning core's. pm_runtime_force_suspend() powers a core down whatever the usage count says, and the owning core need not be the first one suspended, so a suspend that aborts partway would otherwise leave the clock raised over islands that are already gated. The clock and the supply are claimed in one dev_pm_opp_set_config() call before the table is added. Naming the clock there is not optional: with no name the OPP core takes the first clock in the node, which is the bus clock. Configuration and table live on the owning core, which need not be the device being probed at the time, so none of it can be devres: the call hands back a token, and rocket_devfreq_fini() releases it, and the table, by hand. Left to devres, a later bind would fail on an OPP table the previous teardown never emptied. The rate is changed through dev_pm_opp_set_rate(), so that the supply moves with it. Nothing may set the rate behind the OPP core's back: it caches the OPP it last applied and skips a repeat request for it, so a raw clk_set_rate() would turn the next request for a raised rate into a silent no-op. The boot-rate restore added in the previous patch is converted accordingly. The rate the previous patch reads at probe is the firmware's own number and need not be in the OPP table: a board that does not pin the clock with assigned-clock-rates gets whatever the firmware's divider produces. So the boot rate is normalised to the table once, at init, to the lowest OPP that is not below it, and everything here compares against and returns to that OPP. Comparing cur_freq against a rate that is not in the table would leave the cores held after the first change, with no governor tick ever able to match it. ->get_cur_freq() and the status callback report the rate this driver last requested rather than asking the clock: devfreq queries them from sysfs with no runtime PM reference of its own, and the answer has to be a rate from the OPP table or devfreq_get_freq_level() does not find it and devfreq warns on every transition. "Requested" is the accurate word: a rate the firmware refuses does not come back as an error, because the clock framework ignores what the clock's set_rate returns. Keeping every request to the OPP table, which names only the rates the firmware accepts, keeps a request from being refused; what sysfs shows is still the request, not a measurement. A devicetree with no OPP table is not an error: the driver returns without a devfreq device and the NPU keeps its boot rate, as before this patch. The governor thresholds are a starting point taken from the other accelerators in tree, not a measurement. Inference workloads have not been profiled against them. Assisted-by: LLM sparse checkpatch Signed-off-by: Igor Paunovic --- v2: - The devfreq device goes on the first core in devicetree order with an OPP table, not on whichever core with the table bound first, as v1 would have done with the table on all three cores (Nicolas). - The boot rate is normalised to the OPP table at init. Found by reading the code: without assigned-clock-rates the raw rate is not in the table and the cores would have stayed held after the first change. Code read only so far; a devicetree without the property has not been booted. - Init holds every core while it programs the initial OPP, since rounding the boot rate up to the table may raise the clock. - "last programmed" is now "last requested", with the reason. - Comments and text updated for the shared table; a comment that counted ten governor ticks a second now counts twenty (50 ms polling). - The commit message now says that unbinding any core takes the devfreq device down until all are bound again; v1 did the same without saying so. - The utilisation counters of a core that was unbound and bound again start from zero. - A comment over hold_all() states the invariant the loops rely on. drivers/accel/rocket/Kconfig | 2 + drivers/accel/rocket/Makefile | 1 + drivers/accel/rocket/rocket_core.h | 11 + drivers/accel/rocket/rocket_devfreq.c | 500 ++++++++++++++++++++++++++ drivers/accel/rocket/rocket_devfreq.h | 65 ++++ drivers/accel/rocket/rocket_device.h | 3 + drivers/accel/rocket/rocket_drv.c | 59 ++- drivers/accel/rocket/rocket_job.c | 7 + 8 files changed, 646 insertions(+), 2 deletions(-) create mode 100644 drivers/accel/rocket/rocket_devfreq.c create mode 100644 drivers/accel/rocket/rocket_devfreq.h diff --git a/drivers/accel/rocket/Kconfig b/drivers/accel/rocket/Kconfig index 16465abe06607..00ee845c871fa 100644 --- a/drivers/accel/rocket/Kconfig +++ b/drivers/accel/rocket/Kconfig @@ -8,6 +8,8 @@ config DRM_ACCEL_ROCKET depends on MMU select DRM_SCHED select DRM_GEM_SHMEM_HELPER + select PM_DEVFREQ + select DEVFREQ_GOV_SIMPLE_ONDEMAND help Choose this option if you have a Rockchip SoC that contains a compatible Neural Processing Unit (NPU), such as the RK3588. Called by diff --git a/drivers/accel/rocket/Makefile b/drivers/accel/rocket/Makefile index 3713dfe223d6e..e0944b3e68121 100644 --- a/drivers/accel/rocket/Makefile +++ b/drivers/accel/rocket/Makefile @@ -5,6 +5,7 @@ obj-$(CONFIG_DRM_ACCEL_ROCKET) :=3D rocket.o rocket-y :=3D \ rocket_core.o \ rocket_device.o \ + rocket_devfreq.o \ rocket_drv.o \ rocket_gem.o \ rocket_job.o diff --git a/drivers/accel/rocket/rocket_core.h b/drivers/accel/rocket/rock= et_core.h index 46ed8352a79d2..c9995bd9e0553 100644 --- a/drivers/accel/rocket/rocket_core.h +++ b/drivers/accel/rocket/rocket_core.h @@ -7,6 +7,7 @@ #include #include #include +#include #include #include =20 @@ -60,6 +61,16 @@ struct rocket_core { struct drm_gpu_scheduler sched; u64 fence_context; u64 emit_seqno; + + /* + * Utilisation seen by devfreq, guarded by rdev->devfreq.busy_lock. A + * core runs one task at a time, so a flag is exact here and, unlike a + * counter, cannot be left skewed by a job the reset path tore down. + */ + bool busy; + ktime_t busy_time; + ktime_t idle_time; + ktime_t time_last_update; }; =20 int rocket_core_init(struct rocket_core *core); diff --git a/drivers/accel/rocket/rocket_devfreq.c b/drivers/accel/rocket/r= ocket_devfreq.c new file mode 100644 index 0000000000000..871fa370eb432 --- /dev/null +++ b/drivers/accel/rocket/rocket_devfreq.c @@ -0,0 +1,500 @@ +// SPDX-License-Identifier: GPL-2.0-only +/* Copyright 2025 Igor Paunovic */ + +#include +#include +#include +#include +#include +#include +#include + +#include "rocket_core.h" +#include "rocket_device.h" +#include "rocket_devfreq.h" + +/* + * One clock and one supply feed all of the NPU cores, so a single devfreq + * device drives them together. It hangs off the first core in devicetree + * order that carries an OPP table. + * + * The awkward part is that the clock is generated by a PVTPLL that sits + * inside the NPU power islands. An island that is powered up while the cl= ock + * is above the rate the bootloader left never acknowledges the power-on, = and + * the first register access into it afterwards takes an asynchronous SErr= or. + * + * Lowering the rate is fine at any time as long as the boot rate is one t= he + * firmware serves from GPLL, which on the RK3588 is the pinned 200 MHz: f= or + * that rate it writes only CRU clock selectors. Raising the rate is not, = so + * before it goes up every core is runtime resumed and the + * references are kept for as long as the clock stays raised. While they a= re + * held no core can suspend, so no island can transition at all, which is = the + * property that makes the raised clock safe rather than merely unlikely t= o be + * caught out. + * + * The cost is that a busy NPU does not power-gate individual cores. It is + * paid only above the boot rate: at the boot rate the references are drop= ped + * and runtime PM behaves exactly as it did before this file existed. The = cost + * in milliwatts has not been measured on this board, and a measurement or= a + * better idea would both be welcome. + */ + +static void rocket_devfreq_update_utilisation(struct rocket_core *core) +{ + ktime_t now =3D ktime_get(); + ktime_t elapsed =3D ktime_sub(now, core->time_last_update); + + if (core->busy) + core->busy_time =3D ktime_add(core->busy_time, elapsed); + else + core->idle_time =3D ktime_add(core->idle_time, elapsed); + + core->time_last_update =3D now; +} + +/* + * The devfreq device exists only while every slot in rdev->cores[] is + * filled: it goes up when num_cores reaches max_cores and comes down in + * rocket_remove() before the leaving core empties its slot. The loops + * over num_cores here and below rely on that. + */ +static int rocket_devfreq_hold_all(struct rocket_device *rdev) +{ + unsigned int i; + int ret; + + for (i =3D 0; i < rdev->num_cores; i++) { + ret =3D pm_runtime_resume_and_get(rdev->cores[i].dev); + if (ret < 0) { + while (i--) + pm_runtime_put_autosuspend(rdev->cores[i].dev); + + return ret; + } + } + + return 0; +} + +static void rocket_devfreq_release_all(struct rocket_device *rdev) +{ + unsigned int i; + + for (i =3D 0; i < rdev->num_cores; i++) + pm_runtime_put_autosuspend(rdev->cores[i].dev); +} + +/* Caller holds rdev->devfreq.lock. */ +static int rocket_devfreq_set_rate(struct rocket_device *rdev, unsigned lo= ng freq) +{ + struct rocket_devfreq *rdevfreq =3D &rdev->devfreq; + struct device *dev =3D rdevfreq->owner->dev; + int ret; + + ret =3D dev_pm_opp_set_rate(dev, freq); + if (ret) { + dev_err(dev, "failed to set the NPU rate to %lu Hz: %d\n", freq, ret); + return ret; + } + + WRITE_ONCE(rdevfreq->cur_freq, freq); + + return 0; +} + +static int rocket_devfreq_target(struct device *dev, unsigned long *freq, = u32 flags) +{ + struct rocket_device *rdev =3D dev_get_drvdata(dev); + struct rocket_devfreq *rdevfreq =3D &rdev->devfreq; + struct dev_pm_opp *opp; + int ret; + + opp =3D devfreq_recommended_opp(dev, freq, flags); + if (IS_ERR(opp)) + return PTR_ERR(opp); + dev_pm_opp_put(opp); + + guard(mutex)(&rdevfreq->lock); + + /* + * The governor calls this on every tick, including the ticks where it + * arrives at the rate the NPU is already running. Without this an idle + * NPU would resume all of its cores twenty times a second to set the + * rate they already have. + */ + if (*freq =3D=3D READ_ONCE(rdevfreq->cur_freq)) + return 0; + + if (!rdevfreq->cores_held) { + ret =3D rocket_devfreq_hold_all(rdev); + if (ret) { + /* + * Runtime PM is disabled on the way into system + * suspend, so this is the ordinary way for a governor + * tick that raced with it to end. + */ + dev_dbg(dev, "cannot resume the NPU cores to change rate: %d\n", ret); + return ret; + } + rdevfreq->cores_held =3D true; + } + + ret =3D rocket_devfreq_set_rate(rdev, *freq); + + /* At the boot rate the cores are free to suspend again. */ + if (READ_ONCE(rdevfreq->cur_freq) <=3D rdevfreq->boot_freq) { + rdevfreq->cores_held =3D false; + rocket_devfreq_release_all(rdev); + } + + return ret; +} + +static int rocket_devfreq_get_dev_status(struct device *dev, + struct devfreq_dev_status *status) +{ + struct rocket_device *rdev =3D dev_get_drvdata(dev); + ktime_t busy =3D 0, total =3D 0; + unsigned int i; + + scoped_guard(spinlock_irqsave, &rdev->devfreq.busy_lock) { + for (i =3D 0; i < rdev->num_cores; i++) { + struct rocket_core *core =3D &rdev->cores[i]; + + rocket_devfreq_update_utilisation(core); + + /* + * The cores share the clock, so what the rate has to + * satisfy is the busiest of them. Adding the cores up + * instead would report one saturated core out of three + * as a third of the load, and clock down underneath it. + */ + busy =3D max(busy, core->busy_time); + total =3D max(total, ktime_add(core->busy_time, core->idle_time)); + + core->busy_time =3D 0; + core->idle_time =3D 0; + } + } + + status->busy_time =3D ktime_to_ns(busy); + status->total_time =3D ktime_to_ns(total); + status->current_frequency =3D READ_ONCE(rdev->devfreq.cur_freq); + + dev_dbg(dev, "busy %lu total %lu %lu%% freq %lu MHz\n", + status->busy_time, status->total_time, + status->busy_time * 100 / max(status->total_time, 1UL), + status->current_frequency / 1000 / 1000); + + return 0; +} + +/* + * Report the rate this driver last requested, not what the firmware says. + * devfreq asks for the current frequency from sysfs as well, without a ru= ntime + * PM reference of its own, and the answer has to be one of the rates in t= he + * OPP table or devfreq_get_freq_level() will not find it and every transi= tion + * will be logged as unknown. + * + * "Requested" is the accurate word: a rate the firmware refuses does not = come + * back as an error, because the clock framework ignores what the clock's + * set_rate returns and the clock stays where it was. Keeping every reques= t to + * the OPP table, which names only the rates the firmware accepts, keeps a + * request from being refused; what is reported is still the request, not a + * measurement. + */ +static int rocket_devfreq_get_cur_freq(struct device *dev, unsigned long *= freq) +{ + struct rocket_device *rdev =3D dev_get_drvdata(dev); + + *freq =3D READ_ONCE(rdev->devfreq.cur_freq); + + return 0; +} + +static struct devfreq_dev_profile rocket_devfreq_profile =3D { + .timer =3D DEVFREQ_TIMER_DELAYED, + .polling_ms =3D 50, + .target =3D rocket_devfreq_target, + .get_dev_status =3D rocket_devfreq_get_dev_status, + .get_cur_freq =3D rocket_devfreq_get_cur_freq, +}; + +void rocket_devfreq_record_busy(struct rocket_core *core) +{ + struct rocket_devfreq *rdevfreq =3D &core->rdev->devfreq; + + if (!rdevfreq->devfreq) + return; + + scoped_guard(spinlock_irqsave, &rdevfreq->busy_lock) { + rocket_devfreq_update_utilisation(core); + core->busy =3D true; + } +} + +/* + * Idempotent on purpose: a job that times out is torn down by the reset p= ath, + * which cannot know whether the completion interrupt got there first. + */ +void rocket_devfreq_record_idle(struct rocket_core *core) +{ + struct rocket_devfreq *rdevfreq =3D &core->rdev->devfreq; + + if (!rdevfreq->devfreq) + return; + + scoped_guard(spinlock_irqsave, &rdevfreq->busy_lock) { + rocket_devfreq_update_utilisation(core); + core->busy =3D false; + } +} + +/* + * Put the clock back to the boot OPP through the OPP core. Called with no + * lock held, from the last core on its way down. + */ +int rocket_devfreq_set_boot_rate(struct rocket_device *rdev) +{ + struct rocket_devfreq *rdevfreq =3D &rdev->devfreq; + int ret; + + ret =3D dev_pm_opp_set_rate(rdevfreq->owner->dev, rdevfreq->boot_freq); + if (!ret) + WRITE_ONCE(rdevfreq->cur_freq, rdevfreq->boot_freq); + + return ret; +} + +/* + * System suspend, called from every core's ->suspend before it is forced = down. + * The first one to get here does the work and the rest are no-ops. + * + * It has to be every core and not just the one that owns the devfreq devi= ce. + * That core is suspended last, and a suspend that aborts partway - a pend= ing + * wakeup, or this driver's own -EBUSY on a core that is still busy - would + * never reach it, leaving the clock raised over already gated islands. + * + * No governor tick can be in flight here: dpm_suspend() calls devfreq_sus= pend() + * before it walks the devices, which stops the monitor on every registered + * devfreq. That is also where this driver does get devfreq_suspend_device= () - + * from the PM core, on a path that holds no runtime PM lock of ours. + */ +void rocket_devfreq_suspend(struct rocket_device *rdev) +{ + struct rocket_devfreq *rdevfreq =3D &rdev->devfreq; + + if (!rdevfreq->devfreq) + return; + + guard(mutex)(&rdevfreq->lock); + + if (!rdevfreq->cores_held) + return; + + rocket_devfreq_set_rate(rdev, rdevfreq->boot_freq); + rdevfreq->cores_held =3D false; + rocket_devfreq_release_all(rdev); +} + +int rocket_devfreq_init(struct rocket_device *rdev) +{ + static const char * const clk_names[] =3D { "npu", NULL }; + static const char * const supplies[] =3D { "npu", NULL }; + struct dev_pm_opp_config config =3D { + .clk_names =3D clk_names, + .regulator_names =3D supplies, + }; + struct rocket_devfreq *rdevfreq =3D &rdev->devfreq; + struct rocket_core *owner =3D NULL; + struct dev_pm_opp *opp; + struct device *dev; + unsigned long freq; + unsigned int i; + int ret; + + /* + * The devfreq device hangs off the first core in devicetree order that + * carries an OPP table: on the RK3588 every core references the shared + * table, so that is rknn_core_0. It has to be a fixed choice and not + * whichever core bound last, because the devfreq device is named after + * this core and a cooling map in the devicetree resolves against its + * node. core->index is the core's position among the core nodes, so + * the lowest index is the first node. + */ + for (i =3D 0; i < rdev->num_cores; i++) { + struct rocket_core *core =3D &rdev->cores[i]; + + if (!of_property_present(core->dev->of_node, "operating-points-v2")) + continue; + + if (!owner || core->index < owner->index) + owner =3D core; + } + + /* + * No OPP table is not an error. It asks for the NPU to stay at the + * rate it booted at, which is what this driver did before devfreq. + */ + if (!owner) + return 0; + + dev =3D owner->dev; + + /* + * None of this can be devres. It is attached to the owning core, while + * the device being probed right now is whichever core happened to bind + * last, so devres would outlive the teardown in rocket_devfreq_fini() + * and a later rebind would find the OPP table still populated. + * + * Naming the clock matters as much as claiming the supply: without it + * the OPP core takes the first clock in the node, which is the bus + * clock, and would scale that instead of the compute clock. + */ + ret =3D dev_pm_opp_set_config(dev, &config); + if (ret < 0) { + if (ret !=3D -ENODEV) + return dev_err_probe(dev, ret, + "failed to set the OPP clock and supply\n"); + + dev_info(dev, "no NPU supply described, leaving the clock alone\n"); + return 0; + } + rdevfreq->opp_token =3D ret; + + ret =3D dev_pm_opp_of_add_table(dev); + if (ret) { + if (ret !=3D -ENODEV) + dev_err_probe(dev, ret, "failed to add the OPP table\n"); + else + ret =3D 0; + + goto err_clear_config; + } + + mutex_init(&rdevfreq->lock); + spin_lock_init(&rdevfreq->busy_lock); + + for (i =3D 0; i < rdev->num_cores; i++) { + struct rocket_core *core =3D &rdev->cores[i]; + + /* A core that was unbound and bound again starts over. */ + core->busy =3D false; + core->busy_time =3D 0; + core->idle_time =3D 0; + core->time_last_update =3D ktime_get(); + } + + freq =3D rdev->npu_boot_rate; + opp =3D devfreq_recommended_opp(dev, &freq, 0); + if (IS_ERR(opp)) { + ret =3D dev_err_probe(dev, PTR_ERR(opp), + "no OPP covers the %lu Hz boot rate\n", + rdev->npu_boot_rate); + goto err_remove_table; + } + + /* + * From here on the boot rate is the OPP it maps to. The raw rate is + * the firmware's number and need not be in the table at all. + */ + rdevfreq->boot_freq =3D freq; + + /* + * Program the supply for the rate the NPU is already running, so that + * the regulator is not switched off underneath it by + * regulator_late_cleanup(). That OPP is the boot rate rounded up to + * the table, so this may raise the clock, and a raised clock is only + * ever programmed with every core held: the same rule as ->target(). + */ + ret =3D rocket_devfreq_hold_all(rdev); + if (ret) { + dev_pm_opp_put(opp); + dev_err_probe(dev, ret, + "cannot resume the NPU cores to set the initial OPP\n"); + goto err_remove_table; + } + ret =3D dev_pm_opp_set_opp(dev, opp); + rocket_devfreq_release_all(rdev); + dev_pm_opp_put(opp); + if (ret) { + dev_err_probe(dev, ret, "failed to set the initial OPP\n"); + goto err_remove_table; + } + + rdevfreq->cur_freq =3D freq; + rdevfreq->owner =3D owner; + rocket_devfreq_profile.initial_freq =3D freq; + + /* + * A starting point taken from the other accelerators in tree, not a + * measurement: inference workloads have not been profiled against + * these thresholds. + */ + rdevfreq->gov_data.upthreshold =3D 50; + rdevfreq->gov_data.downdifferential =3D 10; + + rdevfreq->devfreq =3D devfreq_add_device(dev, &rocket_devfreq_profile, + DEVFREQ_GOV_SIMPLE_ONDEMAND, + &rdevfreq->gov_data); + if (IS_ERR(rdevfreq->devfreq)) { + ret =3D PTR_ERR(rdevfreq->devfreq); + rdevfreq->devfreq =3D NULL; + rdevfreq->owner =3D NULL; + + dev_err_probe(dev, ret, "failed to add the devfreq device\n"); + goto err_remove_table; + } + + return 0; + +err_remove_table: + mutex_destroy(&rdevfreq->lock); + dev_pm_opp_of_remove_table(dev); +err_clear_config: + dev_pm_opp_clear_config(rdevfreq->opp_token); + rdevfreq->opp_token =3D 0; + + return ret; +} + +void rocket_devfreq_fini(struct rocket_device *rdev) +{ + struct rocket_devfreq *rdevfreq =3D &rdev->devfreq; + struct device *dev; + + if (!rdevfreq->devfreq) + return; + + dev =3D rdevfreq->owner->dev; + + devfreq_remove_device(rdevfreq->devfreq); + rdevfreq->devfreq =3D NULL; + + /* + * Lower the clock before letting go of the cores, not after: a core + * that suspends while the clock is still raised would be unable to + * come back. + */ + scoped_guard(mutex, &rdevfreq->lock) { + if (rdevfreq->cores_held) { + rocket_devfreq_set_rate(rdev, rdevfreq->boot_freq); + rdevfreq->cores_held =3D false; + rocket_devfreq_release_all(rdev); + } + } + + /* + * Undo the OPP setup by hand, in the reverse order. Everything above + * lives on the owning core rather than on the device that was probing + * when it was set up, so nothing here is released by devres; leaving + * the table populated would make the next bind fail with the OPP core + * complaining that it is not empty. After this the owner is gone, so + * anything that still wants the boot rate asks the clock directly. + */ + rdevfreq->owner =3D NULL; + mutex_destroy(&rdevfreq->lock); + dev_pm_opp_of_remove_table(dev); + dev_pm_opp_clear_config(rdevfreq->opp_token); + rdevfreq->opp_token =3D 0; +} diff --git a/drivers/accel/rocket/rocket_devfreq.h b/drivers/accel/rocket/r= ocket_devfreq.h new file mode 100644 index 0000000000000..65a9de6d37389 --- /dev/null +++ b/drivers/accel/rocket/rocket_devfreq.h @@ -0,0 +1,65 @@ +/* SPDX-License-Identifier: GPL-2.0-only */ +/* Copyright 2025 Igor Paunovic */ + +#ifndef __ROCKET_DEVFREQ_H__ +#define __ROCKET_DEVFREQ_H__ + +#include +#include +#include + +struct rocket_core; +struct rocket_device; + +struct rocket_devfreq { + struct devfreq *devfreq; + struct devfreq_simple_ondemand_data gov_data; + + /* + * The core the devfreq device hangs off: the first in devicetree order + * to carry an OPP table. NULL when no core does. + */ + struct rocket_core *owner; + + /* + * The OPP clock and supply configuration is attached to the owning + * core, which is not the device this driver is probing when it is set + * up, so it cannot be devres. Zero means nothing is attached. + */ + int opp_token; + + /* + * Serialises rate changes against each other and against the set of + * runtime PM references taken below. + * + * This is never taken from a runtime PM callback. A rate change + * resumes every core while holding it, so a core that took it on its + * way down would wait for a rate change that is waiting for that same + * core to finish suspending. + */ + struct mutex lock; + unsigned long cur_freq; + bool cores_held; + + /* + * The boot rate as an OPP: the lowest rate in the table that is not + * below rdev->npu_boot_rate. The raw boot rate is whatever the firmware + * reported at probe and need not be in the table, and comparing + * cur_freq against a rate that is not in the table would leave the + * cores held for good after the first change. Everything here compares + * against and returns to this instead. + */ + unsigned long boot_freq; + + /* Guards the utilisation fields of every core. */ + spinlock_t busy_lock; +}; + +int rocket_devfreq_init(struct rocket_device *rdev); +void rocket_devfreq_fini(struct rocket_device *rdev); +void rocket_devfreq_suspend(struct rocket_device *rdev); +int rocket_devfreq_set_boot_rate(struct rocket_device *rdev); +void rocket_devfreq_record_busy(struct rocket_core *core); +void rocket_devfreq_record_idle(struct rocket_core *core); + +#endif /* __ROCKET_DEVFREQ_H__ */ diff --git a/drivers/accel/rocket/rocket_device.h b/drivers/accel/rocket/ro= cket_device.h index ba7c977cd6951..a91d5a9b09ed3 100644 --- a/drivers/accel/rocket/rocket_device.h +++ b/drivers/accel/rocket/rocket_device.h @@ -11,6 +11,7 @@ #include =20 #include "rocket_core.h" +#include "rocket_devfreq.h" =20 struct rocket_device { struct drm_device ddev; @@ -37,6 +38,8 @@ struct rocket_device { */ unsigned long npu_boot_rate; atomic_t active_cores; + + struct rocket_devfreq devfreq; }; =20 struct rocket_device *rocket_device_init(struct platform_device *pdev, diff --git a/drivers/accel/rocket/rocket_drv.c b/drivers/accel/rocket/rocke= t_drv.c index 8f03de1af488c..c6eab2239b6a9 100644 --- a/drivers/accel/rocket/rocket_drv.c +++ b/drivers/accel/rocket/rocket_drv.c @@ -14,6 +14,7 @@ #include =20 #include "rocket_device.h" +#include "rocket_devfreq.h" #include "rocket_drv.h" #include "rocket_gem.h" #include "rocket_job.h" @@ -244,6 +245,20 @@ static int rocket_probe(struct platform_device *pdev) if (ret) goto err_core; =20 + /* + * Every core described in the devicetree has to be bound before the + * devfreq device goes up. A rate change resumes all of them and keeps + * them resumed, and a core that had not probed yet would come up later + * underneath a raised clock. + */ + if (rdev->num_cores =3D=3D rdev->max_cores) { + ret =3D rocket_devfreq_init(rdev); + if (ret) { + rocket_core_fini(&rdev->cores[core]); + goto err_core; + } + } + return 0; =20 err_core: @@ -268,6 +283,9 @@ static void rocket_remove(struct platform_device *pdev) if (WARN_ON(core < 0)) return; =20 + /* The devfreq device drives every core, so it goes before any of them. */ + rocket_devfreq_fini(rdev); + rocket_core_fini(&rdev->cores[core]); rdev->cores[core].dev =3D NULL; rdev->num_cores--; @@ -314,7 +332,18 @@ static void rocket_npu_restore_boot_rate(struct rocket= _core *core) if (!rdev->npu_boot_rate) return; =20 - err =3D clk_set_rate(core->clks[2].clk, rdev->npu_boot_rate); + /* + * Go through the OPP core once there is a table, never behind its + * back: it caches the OPP it last applied and skips a request for that + * same OPP, so a raw clk_set_rate() here would make the next request + * for the raised rate a silent no-op, with sysfs reporting a rate the + * hardware was not running. + */ + if (rdev->devfreq.owner) + err =3D rocket_devfreq_set_boot_rate(rdev); + else + err =3D clk_set_rate(core->clks[2].clk, rdev->npu_boot_rate); + if (err) dev_warn(core->dev, "failed to restore the NPU boot rate of %lu Hz: %d\n", @@ -360,14 +389,38 @@ static int rocket_device_runtime_suspend(struct devic= e *dev) return 0; } =20 +static int rocket_device_suspend(struct device *dev) +{ + struct rocket_device *rdev =3D dev_get_drvdata(dev); + + /* + * pm_runtime_force_suspend() below powers this core down whatever the + * runtime PM usage count says, so the references taken while the clock + * is raised do not hold it off. Put the rate back first; the call is a + * no-op on every core after the first. + */ + rocket_devfreq_suspend(rdev); + + return pm_runtime_force_suspend(dev); +} + EXPORT_GPL_DEV_PM_OPS(rocket_pm_ops) =3D { RUNTIME_PM_OPS(rocket_device_runtime_suspend, rocket_device_runtime_resum= e, NULL) - SYSTEM_SLEEP_PM_OPS(pm_runtime_force_suspend, pm_runtime_force_resume) + SYSTEM_SLEEP_PM_OPS(rocket_device_suspend, pm_runtime_force_resume) }; =20 /* * A kexec hands the next kernel whatever rate is set here, and that kernel * will power the islands up before it looks at the clock. + * + * Take the devfreq device down before restoring the rate rather than afte= r. + * Nothing freezes workqueues on this path, so a governor tick that landed + * after the restore would raise the clock straight back up and hand on ex= actly + * what this is here to prevent. + * + * The hook runs once per core. After the first call the devfreq device is + * gone, and each later restore asks the clock for the rate it already has, + * which the clock framework drops before it reaches the firmware. */ static void rocket_shutdown(struct platform_device *pdev) { @@ -377,6 +430,8 @@ static void rocket_shutdown(struct platform_device *pde= v) if (!rdev) return; =20 + rocket_devfreq_fini(rdev); + core =3D find_core_for_dev(&pdev->dev); if (core >=3D 0) rocket_npu_restore_boot_rate(&rdev->cores[core]); diff --git a/drivers/accel/rocket/rocket_job.c b/drivers/accel/rocket/rocke= t_job.c index 25ee4ab172a82..7846b627b00f4 100644 --- a/drivers/accel/rocket/rocket_job.c +++ b/drivers/accel/rocket/rocket_job.c @@ -15,6 +15,7 @@ =20 #include "rocket_core.h" #include "rocket_device.h" +#include "rocket_devfreq.h" #include "rocket_drv.h" #include "rocket_job.h" #include "rocket_registers.h" @@ -149,6 +150,8 @@ static void rocket_job_hw_submit(struct rocket_core *co= re, struct rocket_job *jo =20 rocket_pc_writel(core, TASK_DMA_BASE_ADDR, PC_TASK_DMA_BASE_ADDR_DMA_BASE= _ADDR(0x0)); =20 + rocket_devfreq_record_busy(core); + rocket_pc_writel(core, OPERATION_ENABLE, PC_OPERATION_ENABLE_OP_EN(1)); =20 dev_dbg(core->dev, "Submitted regcmd at 0x%llx to core %d", task->regcmd,= core->index); @@ -348,6 +351,8 @@ static void rocket_job_handle_irq(struct rocket_core *c= ore) rocket_pc_writel(core, OPERATION_ENABLE, 0x0); rocket_pc_writel(core, INTERRUPT_CLEAR, 0x1ffff); =20 + rocket_devfreq_record_idle(core); + scoped_guard(mutex, &core->job_lock) if (core->in_flight_job) { if (core->in_flight_job->next_task_idx < core->in_flight_job->task_coun= t) { @@ -379,6 +384,8 @@ rocket_reset(struct rocket_core *core, struct drm_sched= _job *bad) if (core->in_flight_job) pm_runtime_put_noidle(core->dev); =20 + rocket_devfreq_record_idle(core); + iommu_detach_group(NULL, core->iommu_group); =20 core->in_flight_job =3D NULL; --=20 2.43.0 From nobody Thu Sep 24 16:07:16 2026 Received: from mail-wm2-f13.google.com (mail-wm2-f13.google.com [74.125.225.141]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 16421525A92 for ; Tue, 22 Sep 2026 08:01:42 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.225.141 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790064124; cv=none; b=X+/L9D5dCAR/SFZ2N/2hGZysPVm+VsDXXl6RinWoYq+fVpjRQvnuFA55AF5ZrpII7GdVNB75/Q8r4DikxHsPRY6xF8EjnY4DyBY9v4jllZJ7ShMDmYDIIZ2V5g4kztEBaG/TsWLuz6ZcrlL6U3lw7qjw5Ni6gW4F3g275r/JzF0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790064124; c=relaxed/simple; bh=2W7J4uciibwZjfedWf0nIKQ9arloV2QAqcu52OEo04Q=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=TN2DfFS1D1xKrJvTNZrOQf1ZM7qbwtU8spLs4lVMLVVVIYeom3/G01U1cXdo7w/5+hZF9nWVjKF+1ZULnc0z35UscH7QlyvQpvwirA2vhnqn0J233d+Y55nPmnzB0O01MnlSeYezH0WGRJbsxFaBiXDG3poYLh/aDu3OOf71wmE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=mQL9e2fF; arc=none smtp.client-ip=74.125.225.141 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="mQL9e2fF" Received: by mail-wm2-f13.google.com with SMTP id 5b1f17b1804b1-49c5a927a1fso2575385e9.2 for ; Tue, 22 Sep 2026 01:01:42 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1790064099; x=1790668899; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=cUwmXn/jLJ9RJT7pjA9WmiefPpGzIu17sEQJkrs8eXE=; b=mQL9e2fFJlcQaP2tWvPb5kX3i+wpKhtOghrA46WaLa6ivFGLYNN++cw/uNYWHvMBaB BHDMyhMWJS33O0Dh81XRhtzn1njUWH+lb1S8pftaQ8lPnvR8zAAu29VNIHadYI61lbh1 Upgu859BD2nQuRnzcwP6cXJDxz8b1OhIhYCi/SbTij9O/zLL2Hs81uTt6+SMnWnRR5C/ kG/SRoHBVkk5Br/KWuasLGfIV9xckgFFkc3TSm8gVFGtck8Maq+G0hGBdTXOR/CEp03p mKZ3iMqXx1bt7RZ/pTrycxWcQO+7IhKm7KXx3gYIxzMtk+qpDk0xPvHSqRF5ykQEbT32 qtcw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790064099; x=1790668899; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=cUwmXn/jLJ9RJT7pjA9WmiefPpGzIu17sEQJkrs8eXE=; b=eKhnh3PADltOjcFYgHEvrCE0ip8GKD7ooSiNsplBt35a45Ffl8tnY1Uuw8wPWAXg92 7B6zPATDeGPdHOUSG+3nEpI0o/d1NnxYkXjBHgOp1OpxJAESOUgLF2hJ/o8S4K88a3vD TFv9DVpZclBgTmfJiO05J+ZjwTT1zIJTY7q2j4QyCDHdkYxovU5mOg/PD/5vglcecpLj z+ShZWBnnCUreJATzoM1BGCSjY7vu5iSQordcdV11HdEWn4wMPOmXqsq/NloiZROGdn1 KAXXGV3FZs7qUGia5YiEc/WNttENjVocLiVBiaHLk6kN6SRWLx2BXMDC5fpYOT5ZMtKc Uxng== X-Forwarded-Encrypted: i=1; AKwUvBzHQV3DAF6ihgypTgjnmh7pZ9OA0khDg3IRkQ0UVxmP8ekfGsN2U3xyN8+YKn5Qr01YovEmtTfHiWsMg0g=@vger.kernel.org X-Gm-Message-State: AFuF++nGjrH4PxJRfAnVZqvTQvm7C+4aReqsFCzm3bzGdnLBVqtaudF9 3lvdRbO+aOaG4Z8nAUH/WCosxGi9OoGtvR1Rppp50TxQl3f25yTyLsEn X-Gm-Gg: AYBFou3Hf2F0Xi5BLispKqVdD8hK2udR3IQWep6bX61LZ/msooZI0nRpws1Imt/skpm uaCqiHFcBWGYN7nxinFV8fluYZg7ScrzigvI/l2hya1hzGvg4PEK7qWmz59q7QIjDxhM2X3BnQX DUBBEegU2MR+zWtTGtNlLf/iZW71AUwE8OpFu/RCnTRBSQJAuswwQ0xxCmDnx00e+FSrL7k7xh6 IkgMMOKeGP4CO87TE9h1eKWGy42irHKHv+v9N21BKCNJN8SGfBezTt6tealELXgVWW8n4Lg8GZB EdqTzAKy5BG0WQ/gYCwSiKvqiZGf25GUmdUBnkldbdlAQCsBv68XeMEhc7/BAmhCJtc688KCoxx JVzU8z4bTliw/Sq/dA+az6y/7fyTZQa/S6skF6JPEzh5OBcfDxcb4abCNFfc2xtF6zqzQoeyuKf ZV8UX1O5s5pugZdNodTqewNJ4QZFx+1JH30CyqUzMRYzFyV7grZ4QM95Ffub3AUkJSjqnzidfc8 im+Dk8GHcz5lcywjpxpAJk5X6IAz20EWBPq6lFEgoIqx+r1ltpQTVraZfu9vMBsrfREsQ4PuseR Q4w= X-Received: by 2002:a05:600c:a00a:b0:499:d95a:41f with SMTP id 5b1f17b1804b1-49fc7b83237mr163730805e9.0.1790064099450; Tue, 22 Sep 2026 01:01:39 -0700 (PDT) Received: from OrangePi5-Plus.BB-HOME (20014C4E1B80530056971C6280202175.dsl.pool.telekom.hu. [2001:4c4e:1b80:5300:5697:1c62:8020:2175]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-49fdaaf97e2sm18248625e9.2.2026.09.22.01.01.38 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 22 Sep 2026 01:01:39 -0700 (PDT) From: Igor Paunovic To: Tomeu Vizoso , Oded Gabbay , Heiko Stuebner Cc: Rob Herring , Krzysztof Kozlowski , Conor Dooley , Jeff Hugo , Robert Foss , Sidong Yang , Diederik de Haas , Sebastian Reichel , Jiaxing Hu , Nicolas Dufresne , Jonas Karlman , Guangshuo Li , =?UTF-8?q?H=C3=BCseyin=20BIYIK?= , dri-devel@lists.freedesktop.org, linux-rockchip@lists.infradead.org, linux-arm-kernel@lists.infradead.org, devicetree@vger.kernel.org, linux-kernel@vger.kernel.org, Igor Paunovic Subject: [PATCH v2 10/11] accel/rocket: register a devfreq cooling device Date: Tue, 22 Sep 2026 10:01:13 +0200 Message-ID: <20260922080114.44662-11-royalnet026@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260922080114.44662-1-royalnet026@gmail.com> References: <20260922080114.44662-1-royalnet026@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" With devfreq driving the NPU clock, a thermal zone can now throttle the NPU by capping that clock. Register the cooling device so a devicetree can bind it to a zone. The _em variant is used, not because there is an energy model today but so that there will be one the day a power coefficient for this NPU is measured. There is none now: the NPU core nodes carry no dynamic-power-coefficient, Rockchip does not publish one, and a made-up number would be worse than no number. devfreq_cooling_em_register() logs the missing model at debug level and registers the cooling device anyway, so what this gets today is step-wise throttling with no power model for the IPA governor to use. Measuring the coefficient is follow-up work. Registration is allowed to fail. A kernel built without DEVFREQ_THERMAL gets a stub that returns an error, and losing throttling is not a reason to refuse to drive the NPU at all, so the failure is logged and probe carries on. The cooling device is unregistered by hand before the devfreq device it is attached to goes away. Assisted-by: LLM sparse checkpatch Signed-off-by: Igor Paunovic --- v2: wording only, for the shared table. drivers/accel/rocket/rocket_devfreq.c | 25 +++++++++++++++++++++++++ drivers/accel/rocket/rocket_devfreq.h | 2 ++ 2 files changed, 27 insertions(+) diff --git a/drivers/accel/rocket/rocket_devfreq.c b/drivers/accel/rocket/r= ocket_devfreq.c index 871fa370eb432..c81887734de77 100644 --- a/drivers/accel/rocket/rocket_devfreq.c +++ b/drivers/accel/rocket/rocket_devfreq.c @@ -3,6 +3,7 @@ =20 #include #include +#include #include #include #include @@ -446,6 +447,25 @@ int rocket_devfreq_init(struct rocket_device *rdev) goto err_remove_table; } =20 + /* + * Thermal throttling is optional, so a kernel built without + * DEVFREQ_THERMAL keeps a working NPU rather than a failed probe. + * + * The _em variant is used so that the driver is ready for an energy + * model the day a power coefficient for this NPU is measured. There is + * none today: the NPU core nodes have no dynamic-power-coefficient, + * the vendor does not publish one, and inventing a number would be + * worse than having none. Without it the EM registration inside is + * skipped and throttling is step-wise, with no power model for IPA to + * use. + */ + rdevfreq->cooling =3D devfreq_cooling_em_register(rdevfreq->devfreq, NULL= ); + if (IS_ERR(rdevfreq->cooling)) { + dev_info(dev, "no devfreq cooling device (%pe), NPU will not be throttle= d\n", + rdevfreq->cooling); + rdevfreq->cooling =3D NULL; + } + return 0; =20 err_remove_table: @@ -468,6 +488,11 @@ void rocket_devfreq_fini(struct rocket_device *rdev) =20 dev =3D rdevfreq->owner->dev; =20 + if (rdevfreq->cooling) { + devfreq_cooling_unregister(rdevfreq->cooling); + rdevfreq->cooling =3D NULL; + } + devfreq_remove_device(rdevfreq->devfreq); rdevfreq->devfreq =3D NULL; =20 diff --git a/drivers/accel/rocket/rocket_devfreq.h b/drivers/accel/rocket/r= ocket_devfreq.h index 65a9de6d37389..26a3749078b9f 100644 --- a/drivers/accel/rocket/rocket_devfreq.h +++ b/drivers/accel/rocket/rocket_devfreq.h @@ -10,9 +10,11 @@ =20 struct rocket_core; struct rocket_device; +struct thermal_cooling_device; =20 struct rocket_devfreq { struct devfreq *devfreq; + struct thermal_cooling_device *cooling; struct devfreq_simple_ondemand_data gov_data; =20 /* --=20 2.43.0 From nobody Thu Sep 24 16:07:16 2026 Received: from mail-wm2-f12.google.com (mail-wm2-f12.google.com [74.125.225.140]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1D5E552842D for ; Tue, 22 Sep 2026 08:01:43 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.225.140 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790064116; cv=none; b=pqxb/DbF8AJL5P0BpUC9CNP73fqd7xaSybgtdfXg0EKf0YLyGZ6c81azWlcvGOzcqXMzPIwwTdetuz4ZX/mJ1YSSgjfTN13WY74BndHiLSvOCIfOhaL2XhbZ3thLaGqkSZQz3I/zWcCmTiw7yn0DjJ/g79GnXlE+zkl4qnrj02U= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790064116; c=relaxed/simple; bh=8caMejK8hgruIPCiytRhmYLDcurgxrGWdTVCVpsRw0U=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=o9GblOOWOq2XBtYaJhxzoHTROWUlQ5NJwyvxLJ5N3/AtRy20zB1dazYeEHGyn7OVeoS7WNpDiI94UvWlyyTUlcOBy5DFF0rkxRwYpInqLNVL7Jv1D9yDGma/yhQPh9gspAyYBg3Em2cqkqLdiMSVxkf1rtUt69qI6o/ICcqrnlY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=qf+nY3bJ; arc=none smtp.client-ip=74.125.225.140 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="qf+nY3bJ" Received: by mail-wm2-f12.google.com with SMTP id 5b1f17b1804b1-49d0726cdbcso2018165e9.0 for ; Tue, 22 Sep 2026 01:01:43 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1790064101; x=1790668901; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=k8SSEreitnBFuEXQj3F3dce2S1yQFmhptrw7zyYrACU=; b=qf+nY3bJ0l4kr1VWLZG7cczDhzi0I0pbqnIGx0g2Ca16CelTZIS3wgyAlHU7Q3ZF+T l3GyoPioQCGr4o+Ymi5fe+rwTJUIcf0pxEbt0dm3UrHYNiQ2Efu4KEEtsfpadkOj5WKu PD0UrdqVCAcrqyYHOLqoODqlg6q1UUw1ysOVm0QCwDu2eqtaUEhU3a8OYqlKEmlf2p4l pLZf/ZLXsvoGHP+OHJKWCI5v60PbFo/olg49RqIdQcQaezuKKnii/GtCGtgEK850S1Iu 4gZ9nvv+RcYzjfPYilp3D6iyRufxOlOSvym9fKuTyrAF3r5f7CYQMePOTWFYjnYTLam9 vOlw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790064101; x=1790668901; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=k8SSEreitnBFuEXQj3F3dce2S1yQFmhptrw7zyYrACU=; b=TlwDTwrF87g29tB73HRP8BC0h0fd0SxCdw2b2ezAVBTNQqXIShjRNOcgOGrOIEjyzM MOKxo7p6t6RT0Q5lLDIiRgvU3OL3HBX4Rihu81PQt70tg2rD1Vn5/fVLRH8l7xqHTwxm am52ykrUOddy7mNZl37xJ68WHiik29VSPAyKeBk3yYK65D0vFrgd/G6m6qX0VnGjl4Kd T70WUqx1UFSOH8uyKHchtmK24LamlFpE7/7MN6rb+kQrGO7EE9zKAlN2ArtBgFTLqG/F PhWyhS4kVNBUqLHT/iMF+qYsuXEyvCgHCEpfV2CGpSlRdKWTqvoT6tkt2+jJskbReYU8 N+BA== X-Forwarded-Encrypted: i=1; AKwUvByzDwtJAB9DJ6np+aFb2C2bdAu935xw3uvKqxv3tT4o9cVrOUw4iC/9T1sUktJBXNL3vQUmii7S5FoFIRQ=@vger.kernel.org X-Gm-Message-State: AFuF++ndbba/2qEmk0LCwzuwg+0ayrEZmP+LcrgudfozBWVs5DAd1F2Q VfhrpAEMz2EqWKelu+atx/y/AeZjKkDm/2DrqCYYFNX5xqrKvHNX2iZu X-Gm-Gg: AYBFou3fRfoNKEqvNtsbIkanH9/vmX2m9/2VfpCkBYkDTI1xIp2fmGq2tzT7FgbJTSy g1Zrzj2PKhf1kYd14e4NlZAtQaJrcHxSsEa4A6NRk8Qp+apw3HNK9oQVBoCpGUcsGQvBTc2ZbcK xFDBrIxXQ4yHDefajVVl2GvBVZQZqvtu5Cz94roOqWww5VWGgSSwdTAtJV5Rt0yFq8WJplmu4A6 8RLoaUH0jmZhda0q2Avfq5TdV+K4hIxpSRtFvm2YbCpGp/zm1t56uZCiEwNZ7ikom7MXg3+LY4z 3ranushpWU50YM6bF4d6QF6GWEd5BbA3DzF4mFAjZtqibqqccp7/waY3GoPjy7BzuoBmEJDkNn/ YE20XkM6MVBFvG4H7smpKrE/sIUGj6jHcdGCtdEdo4iWVF+gW9ucCKLzdLpquYbeQ5ED4v/KsEZ Pzx+b3X/zev/MAIHNMwLSYST2zQyKxlgUbwWYYPZxSTAIuwPaijLYl3gNo+uL4pxtT8b2uMQjq9 ru58fRWtFn1NwXggKWZZzxmti8BpvrXAysApq4FR5ZmnFBqsxtppdxCX+b//+dnVoxPJIVcSKJ7 3Hc= X-Received: by 2002:a05:600c:6992:b0:49b:9241:7ff0 with SMTP id 5b1f17b1804b1-49fc7b7a756mr226941795e9.0.1790064100738; Tue, 22 Sep 2026 01:01:40 -0700 (PDT) Received: from OrangePi5-Plus.BB-HOME (20014C4E1B80530056971C6280202175.dsl.pool.telekom.hu. [2001:4c4e:1b80:5300:5697:1c62:8020:2175]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-49fdaaf97e2sm18248625e9.2.2026.09.22.01.01.39 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 22 Sep 2026 01:01:40 -0700 (PDT) From: Igor Paunovic To: Tomeu Vizoso , Oded Gabbay , Heiko Stuebner Cc: Rob Herring , Krzysztof Kozlowski , Conor Dooley , Jeff Hugo , Robert Foss , Sidong Yang , Diederik de Haas , Sebastian Reichel , Jiaxing Hu , Nicolas Dufresne , Jonas Karlman , Guangshuo Li , =?UTF-8?q?H=C3=BCseyin=20BIYIK?= , dri-devel@lists.freedesktop.org, linux-rockchip@lists.infradead.org, linux-arm-kernel@lists.infradead.org, devicetree@vger.kernel.org, linux-kernel@vger.kernel.org, Igor Paunovic Subject: [PATCH v2 11/11] arm64: dts: rockchip: rk3588: add passive cooling to the NPU thermal zone Date: Tue, 22 Sep 2026 10:01:14 +0200 Message-ID: <20260922080114.44662-12-royalnet026@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260922080114.44662-1-royalnet026@gmail.com> References: <20260922080114.44662-1-royalnet026@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" The NPU zone has had only a critical trip at 115 degrees, which is a shutdown and not a cooling policy. Now that the NPU can be throttled by capping its clock, give the zone a passive trip and a cooling map, in the same shape and at the same temperatures as the GPU zone right above it: 85 degrees with 2 degrees of hysteresis, and a 100 ms passive polling delay. The #cooling-cells property goes on rknn_core_0, the first of the three core nodes, as the one node that stands for the shared clock as a cooling device. The other two cores have no clock of their own and cannot be throttled independently of it. The cooling map is inert until the driver registers a cooling device: thermal_of_should_bind() only resolves a map entry once a matching cdev appears, so this patch on its own changes nothing but the trip point. On a board that leaves the NPU disabled the passive trip has no cooling device to act on and only changes the polling rate above 85 degrees, as the GPU zone's trip already does on a board without the GPU. The thermal path itself has not been exercised on the board this was written on. Reaching 85 degrees on an NPU workload with the fan curve here has not been possible, so what is verified is that the zone parses and binds, not that throttling engages at temperature. Assisted-by: LLM checkpatch dtbs_check Signed-off-by: Igor Paunovic --- v2: commit message only: wording for the shared table, and a sentence on boards that leave the NPU disabled. The diff is unchanged. arch/arm64/boot/dts/rockchip/rk3588-base.dtsi | 17 ++++++++++++++++- 1 file changed, 16 insertions(+), 1 deletion(-) diff --git a/arch/arm64/boot/dts/rockchip/rk3588-base.dtsi b/arch/arm64/boo= t/dts/rockchip/rk3588-base.dtsi index 376ad04e07869..c3d22b08f415b 100644 --- a/arch/arm64/boot/dts/rockchip/rk3588-base.dtsi +++ b/arch/arm64/boot/dts/rockchip/rk3588-base.dtsi @@ -1162,6 +1162,7 @@ rknn_core_0: npu@fdab0000 { clock-names =3D "aclk", "hclk", "npu", "pclk"; assigned-clocks =3D <&scmi_clk SCMI_CLK_NPU>; assigned-clock-rates =3D <200000000>; + #cooling-cells =3D <2>; resets =3D <&cru SRST_A_RKNN0>, <&cru SRST_H_RKNN0>; reset-names =3D "srst_a", "srst_h"; power-domains =3D <&power RK3588_PD_NPUTOP>; @@ -3212,17 +3213,31 @@ map0 { }; =20 npu_thermal: npu-thermal { - polling-delay-passive =3D <0>; + polling-delay-passive =3D <100>; polling-delay =3D <0>; thermal-sensors =3D <&tsadc 6>; =20 trips { + npu_alert: npu-alert { + temperature =3D <85000>; + hysteresis =3D <2000>; + type =3D "passive"; + }; + npu_crit: npu-crit { temperature =3D <115000>; hysteresis =3D <0>; type =3D "critical"; }; }; + + cooling-maps { + map0 { + trip =3D <&npu_alert>; + cooling-device =3D + <&rknn_core_0 THERMAL_NO_LIMIT THERMAL_NO_LIMIT>; + }; + }; }; }; =20 --=20 2.43.0