From nobody Fri Sep 25 19:20:46 2026 Received: from fout-a7-smtp.messagingengine.com (fout-a7-smtp.messagingengine.com [103.168.172.150]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B9DB04A4832; Fri, 25 Sep 2026 14:15:42 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.150 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790345746; cv=none; b=gNNeaFjfc5dNo4dUkT+evxxP6FViDaGZog61gQcNESU/HxUv4jicHR4Yy2J1jgIu/3T4ULjmlpcrgq0UMwJhf8p8ZjLzjeSGR74oQL6KiMRSodlsJkuKQXiJomyQ4Gj7SDNGwg53p0+BiZj24qYN4LaoXbARnpUzVI037NRCMj8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790345746; c=relaxed/simple; bh=rjAyW73kANZYejHA2P8FKv5mgGBCwXIlpRSNuuQa0mo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Al+HyJMBkxF2RmCrH2/0FS6b35x2BmMvzORJla9bBrsiz3ejBGj87XVHeEp+bN0va3xiJt5J3PXwu27GRi8e25pYqu6uHRMZRX3xlkIakhhXeOQ3Ixejedlt0Z9cIlk0H0OipucJMZO+EX1Mo7JVSSvPC9XhZ6Z+FMkNBFw3WAY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=UyZwljfN; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=BzeEOn+G; arc=none smtp.client-ip=103.168.172.150 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="UyZwljfN"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="BzeEOn+G" Received: from phl-compute-12.internal (phl-compute-12.internal [10.202.2.52]) by mailfout.phl.internal (Postfix) with ESMTP id 7D1ADEC0248; Fri, 25 Sep 2026 10:15:41 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-12.internal (MEProxy); Fri, 25 Sep 2026 10:15:41 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm3; t=1790345741; x= 1790432141; bh=1ktQH64MdNmCJGeuXk6bCJk+uHJRmtrdtVuZp+kcp9Y=; b=U yZwljfNmmNsdB8n+hOXtnnChlP1wJjmQWJrYH7xUgC4pjMHpJvRmKtM1YznFptm4 /5aU3qOkNDBCUR5Ekrzthjpk68pAFk8LRbNuMLsk4ja3L025TfI7xTKfEDkUEiuA uf7lbK2VeEQ5jQ2jZdviXCjfxwYPvEflGnzk/qsj4be9Nnaqm6D8SGVf7pfMe590 ynWadFnE3O+hULnb1codqCIvhSNBVQLuhbcaz5uFKD1DKLWQAHEUmSJ+fuCfC+jM 1jCpxfH65m40/QefFcQizi9HSAV476sWppyOus3YpG2sWd0u9NiE5C6pQj1/33Dk yRlQPdHT+Lhk9sINfEHjQ== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm1; t=1790345741; x=1790432141; bh=1 ktQH64MdNmCJGeuXk6bCJk+uHJRmtrdtVuZp+kcp9Y=; b=BzeEOn+Gst4do58kA TSFumWtjcXH7cLsUUprgvYmUTTTsvTXLymG2lFyx6Ygda5lR+UG2nzT5w5/hYP3K 7aWzUYt1hAGAqJRrz8sKJIK/bNVmeDgJlZW8R/HO5l+K1OQY5Sgj5AVCMKvnLMBs WLfNE0vFN47BUU4KjGrEtNjgum0csh8cJOz1DzbbNG/EaDYeE/F4WUggBOdEWgNw OwK/Mozij9v4eF87IMY4Wo61jKtvm8275mM1hy/VnUO2B9nlHqKU7fsRYWaoi6NO DcHqq8xCkSPLeUVhTvBJS9XqKJSH5z/y6R1mpAVp+YJBGXoEpG8fcp23IlKlJufN MsWWA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGKDwwNiF627+asCW7qlOfbc3YO4/BWQQAIa1dD7j4y1/LpRsbneXz52sgooVW2NB 3cVSkvlr4fQi4SwiDMxRaBoBRNO6zHRHeGXMGxJYTRHb1PBeh27/Q3VWbCGBZujoWHPoE6 R+Qc/6k1bBYz+Haf5IzKmyGQxC9qTL+WrK7JzMZqDmDHuiASYj35os3WnrFE1SWLZ6a+qY bCW2glHxjWyrurmiVGB62MJKctiau11N+au9WzHHk3KgIfj2lfSeRdK+MffwyEB0gNSuDN eSmAis7xZg1iofNwJ76iNU49c75bqPHfinmsayTQbADoZ6F/gk68t7NA/yonOE6sVDH8F4 GrbXyWpBbmuCShau1sGcJPX2ea641v6zI8ks6qcI5j3hSLjFADE7BuGhe0vuVlNoqzfY8t WVuJko3s4/EaAbKXVDlSiK/vwhT4qPpLZ/+FqCHRvsC759cyF0XW71AAukG71aEi2jZFUX pwpx0bqJYALapEWolycem0U3kP3vmzPOXIOj1Hub0si31e+psklr2msjxYj1zN/UVXnpR3 l54NK6TZA0prmxnuN6kEtDxGXCwoSx6sH45l+gmQXXPvGs0pxRk1QybcpcNWwpDPhsWJ/l jG5nBYXAQTrGqcrJUAaYSwUdM5xAAh3hPP1A5kgnfim8niTKYziJK65pCqcg X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 25 Sep 2026 10:15:40 -0400 (EDT) From: Kiryl Shutsemau To: Will Deacon , Robin Murphy , Joerg Roedel , Thierry Reding , Jonathan Hunter , Jason Gunthorpe , Nicolin Chen , Breno Leitao Cc: "Kiryl Shutsemau (Meta)" , Krishna Reddy , Pranjal Shrivastava , Mostafa Saleh , Ashish Mhetre , Shameer Kolothum , Yuanhe Shu , Kyle McMartin , Usama Arif , kernel-team@meta.com, linux-arm-kernel@lists.infradead.org, iommu@lists.linux.dev, linux-tegra@vger.kernel.org, linux-kernel@vger.kernel.org Subject: [PATCH v7 1/2] iommu/arm-smmu-v3: Add a cmdq_max_n_shift module parameter Date: Fri, 25 Sep 2026 15:15:29 +0100 Message-ID: <20260925141532.1274962-2-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260925141532.1274962-1-kirill@shutemov.name> References: <20260925141532.1274962-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The command queue depth comes straight from the maximum the hardware advertises in IDR1, which reaches megabytes of coherent DMA per queue. A system with several SMMUv3 instances pays that per instance, and the Tegra241 CMDQV pays it again for every VCMDQ it preallocates. Queue depth only bounds how many commands may be in flight before a sync. A machine driving a handful of devices, or one with a tight memory budget, has no use for the maximum, and no way to say so. Add cmdq_max_n_shift, a cap on the depth given as the log2 of the entry count, the form the hardware itself takes in the LOG2SIZE field of CMDQ_BASE. It defaults to the largest depth the driver allocates, so an unset parameter changes nothing. Decide the depth in arm_smmu_cmdq_max_n_shift(), which applies the parameter to the IDR1 value, so the queue is allocated at the requested size. The Tegra241 CMDQV sizes its VCMDQs from IDR1 itself, so route that through the same helper. Floor the request at one page worth of entries. Without the floor, a small request trips the CMDQ_BATCH_ENTRIES check in arm_smmu_device_hw_probe() and the SMMU fails to probe. The floor also costs nothing: coherent DMA is page granular, so a shallower queue occupies the same memory as one that fills the page. Assisted-by: LLM Reviewed-by: Breno Leitao Reviewed-by: Jason Gunthorpe Signed-off-by: Kiryl Shutsemau (Meta) Reviewed-by: Nicolin Chen Tested-by: Nicolin Chen --- drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c | 40 ++++++++++++++++++- drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h | 1 + .../iommu/arm/arm-smmu-v3/tegra241-cmdqv.c | 2 +- 3 files changed, 40 insertions(+), 3 deletions(-) diff --git a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c b/drivers/iommu/ar= m/arm-smmu-v3/arm-smmu-v3.c index 5732f3ba0122..810d3ce75089 100644 --- a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c +++ b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c @@ -40,6 +40,11 @@ module_param(disable_msipolling, bool, 0444); MODULE_PARM_DESC(disable_msipolling, "Disable MSI-based polling for CMD_SYNC completion."); =20 +static unsigned int cmdq_max_n_shift =3D CMDQ_MAX_SZ_SHIFT; +module_param(cmdq_max_n_shift, uint, 0444); +MODULE_PARM_DESC(cmdq_max_n_shift, + "Cap on the command queue depth, as log2 of the number of entries. Defaul= ts to the hardware maximum; the queue never shrinks below one page."); + static const struct iommu_ops arm_smmu_ops; static struct iommu_dirty_ops arm_smmu_dirty_ops; =20 @@ -4412,6 +4417,37 @@ static struct iommu_dirty_ops arm_smmu_dirty_ops =3D= { }; =20 /* Probing and initialisation functions */ + +/** + * arm_smmu_queue_max_n_shift() - pick the log2 depth of a queue + * @hw_max_n_shift: log2 depth the hardware advertises in IDR1 + * @ent_sz_shift: log2 of the queue entry size in bytes + * @limit_n_shift: log2 depth to cap the queue at + * + * @limit_n_shift is floored at one page, because coherent DMA is page + * granular: a shallower queue occupies the same memory as one that fills = the + * page, and arm_smmu_init_one_queue() stops shrinking at a page too. + */ +static u32 arm_smmu_queue_max_n_shift(u32 hw_max_n_shift, u32 ent_sz_shift, + u32 limit_n_shift) +{ + u32 floor_n_shift =3D PAGE_SHIFT - ent_sz_shift; + + limit_n_shift =3D max(limit_n_shift, floor_n_shift); + return min(hw_max_n_shift, limit_n_shift); +} + +/* + * Command queues are also allocated by the Tegra241 CMDQV for its VCMDQs,= which + * need the same depth decision. + */ +u32 arm_smmu_cmdq_max_n_shift(u32 hw_max_n_shift) +{ + /* Capped to ensure natural alignment */ + return arm_smmu_queue_max_n_shift(hw_max_n_shift, CMDQ_ENT_SZ_SHIFT, + min(CMDQ_MAX_SZ_SHIFT, cmdq_max_n_shift)); +} + int arm_smmu_init_one_queue(struct arm_smmu_device *smmu, struct arm_smmu_queue *q, void __iomem *page, unsigned long prod_off, unsigned long cons_off, @@ -5156,8 +5192,8 @@ static int arm_smmu_device_hw_probe(struct arm_smmu_d= evice *smmu) smmu->features |=3D ARM_SMMU_FEAT_ATTR_TYPES_OVR; =20 /* Queue sizes, capped to ensure natural alignment */ - smmu->cmdq.q.llq.max_n_shift =3D min_t(u32, CMDQ_MAX_SZ_SHIFT, - FIELD_GET(IDR1_CMDQS, reg)); + smmu->cmdq.q.llq.max_n_shift =3D + arm_smmu_cmdq_max_n_shift(FIELD_GET(IDR1_CMDQS, reg)); if (smmu->cmdq.q.llq.max_n_shift <=3D ilog2(CMDQ_BATCH_ENTRIES)) { /* * We don't support splitting up batches, so one batch of diff --git a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h b/drivers/iommu/ar= m/arm-smmu-v3/arm-smmu-v3.h index 50f8321e979c..ea4c87bbe253 100644 --- a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h +++ b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.h @@ -1165,6 +1165,7 @@ static inline void arm_smmu_domain_inv(struct arm_smm= u_domain *smmu_domain) =20 void __arm_smmu_cmdq_skip_err(struct arm_smmu_device *smmu, struct arm_smmu_cmdq *cmdq); +u32 arm_smmu_cmdq_max_n_shift(u32 hw_max_n_shift); int arm_smmu_init_one_queue(struct arm_smmu_device *smmu, struct arm_smmu_queue *q, void __iomem *page, unsigned long prod_off, unsigned long cons_off, diff --git a/drivers/iommu/arm/arm-smmu-v3/tegra241-cmdqv.c b/drivers/iommu= /arm/arm-smmu-v3/tegra241-cmdqv.c index 6644075c1431..710a4c694b94 100644 --- a/drivers/iommu/arm/arm-smmu-v3/tegra241-cmdqv.c +++ b/drivers/iommu/arm/arm-smmu-v3/tegra241-cmdqv.c @@ -663,7 +663,7 @@ static int tegra241_vcmdq_alloc_smmu_cmdq(struct tegra2= 41_vcmdq *vcmdq) /* Cap queue size to SMMU's IDR1.CMDQS and ensure natural alignment */ regval =3D readl_relaxed(smmu->base + ARM_SMMU_IDR1); q->llq.max_n_shift =3D - min_t(u32, CMDQ_MAX_SZ_SHIFT, FIELD_GET(IDR1_CMDQS, regval)); + arm_smmu_cmdq_max_n_shift(FIELD_GET(IDR1_CMDQS, regval)); =20 /* Use the common helper to init the VCMDQ, and then... */ ret =3D arm_smmu_init_one_queue(smmu, q, vcmdq->page0, --=20 2.54.0 From nobody Fri Sep 25 19:20:46 2026 Received: from fhigh-a5-smtp.messagingengine.com (fhigh-a5-smtp.messagingengine.com [103.168.172.156]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DE21F4A9D56; Fri, 25 Sep 2026 14:15:44 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.156 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790345746; cv=none; b=FRJWDAKUVL5MaRrWI7LECvIXb3r7GIOAs3oeBzSyr6P/J9ed2y3brGtzehb9/r0VecFFuQr1DKhpz/8ByAYhuOYmmw6yrMXlIl7Avk4UIuNSwX53FmUtJv0IkbdKILY4QM5/O/h+ezausNVTzem5ghsmr3EK3xxGDo71BApUsD0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790345746; c=relaxed/simple; bh=uL62P1C5xKl+vpUBDxHCmDZ5EXwnuj5Sjo+mo9Gevgo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=NPivLdK0J9BWZcRcgo0B3raTxtBKvEBJToMjvrmuCuFweaJQ8lVEfduaGT+U+tfXmTBS/Hk9BsfJOx+YPkHvt5uzG/2+YxQ9MZiAjBjHRxxUt/xdaF/fO35QJL27WUIU9wBHTEnzqtV/dcZWZHeVm4uLXuw9w7cVULRfkJU7aQ8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=cFdcGhfG; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=bNGQDokD; arc=none smtp.client-ip=103.168.172.156 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="cFdcGhfG"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="bNGQDokD" Received: from phl-compute-03.internal (phl-compute-03.internal [10.202.2.43]) by mailfhigh.phl.internal (Postfix) with ESMTP id 45CD11400168; Fri, 25 Sep 2026 10:15:43 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-03.internal (MEProxy); Fri, 25 Sep 2026 10:15:43 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm3; t=1790345743; x= 1790432143; bh=CSvNW4KC0jILNq7/I9jV1syURG3TP6r5qeySe9d+kW4=; b=c FdcGhfGwM72PqbT32460E1BwZRiHfv+1as0xeCO1B/Fy/N7LzMTdM+cVxx9J45Eg jEuqzVEvOY4wPtxGp0IIICGUHaSUCeM0LJ/vfo+1tdpefCYiH+OPYJqFzW270vqW fKyA8fqpBoDEVCBkltQWYDKDdzbX5aWloaxDl5k0nn+acivm2VU+4O85Xddno9AY 7SzdwXLzggobGf9KgvoJM8P5142wvO0WQO6L+4gYqcnLG3AWXLr9mKlwK53KmLB5 3+h1pQbrnY5tNm/N4QpFfkqy/3SnR7WU2Y/fBuyGnNaK/juEgzGWV50diijgpIPs 8mEWBmt1/k3Kwx2sAh8kA== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm1; t=1790345743; x=1790432143; bh=C SvNW4KC0jILNq7/I9jV1syURG3TP6r5qeySe9d+kW4=; b=bNGQDokDhYo6Wnzd8 uepseBC5jEaPxcmj8Ao0GMXeC2B6v/V91MiewK1WkdgI7RxWtZiroqFd9sVPS81H 9f2PkqvovHto+Ot/NdGfeKGePORDxB4uregYOb8eqmOS+NvuGYQdsd+CublMc/89 zBqbk4fOXJCVIhXdN+c+VyVeDQJK4qYiwQlnGu+Xb0ZZlzf2eAarTmcmEwlSGbEH X8MxqGhmCr/tWXpLC9ILdsarPA7NYV4ic+rc/WMjd9XIkE22Km8wUth46A88j0/F kh7MjEUos0EthaZ4LnHACAWwFGZTKcszwxN3VH+n0e8ledtZfVJ8KdZgZZ16zSKp fw3xA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGKDwwNiF627+asCW7qlOfbc3YO4/BWQQAIa1dD7j4y1/LpRsbneXz52sgooVW2NB 3cVSkvlr4fQi4SwiDMxRaBoBRNO6zHRHeGXMGxJYTRHb1PBeh27/Q3VWbCGBZujoWHPoE6 R+Qc/6k1bBYz+Haf5IzKmyGQxC9qTL+WrK7JzMZqDmDHuiASYj35os3WnrFE1SWLZ6a+qY bCW2glHxjWyrurmiVGB62MJKctiau11N+au9WzHHk3KgIfj2lfSeRdK+MffwyEB0gNSuDN eSmAis7xZg1iofNwJ76iNU49c75bqPHfinmsayTQbADoZ6F/gk68t7NA/yonOE6sVDH8HE D3ftdBE7WnFsPxGcCrWYEQmSAbK7DmTVXeozDOm33abuyGIyUWUYmk42rWMIOSulW5MN5U r+bQ8LHYnJXbH1N1JBlBmKseL7dxUcK49tdAO5EHONR0qoxm8SRl9vmTzDDoKINrV6q1NO v+9uVOC68bBiJKHaZIqSsEIkN9+lFZz8a0ANE3olguIulErakGj3YNqcofh4GyHDCbHfeb bzNkJfo1dosM7+p857j3d16Oey7PxmHreAx04gx6kNEDsjaz2YX3jobUGkB8uhdIxfxWZx OIlwu1cuf1GpDQxmkBIhKPcoHqF/wRcEGrRzImSbCqjKAsFHjqnJeqbC0f2w X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 25 Sep 2026 10:15:42 -0400 (EDT) From: Kiryl Shutsemau To: Will Deacon , Robin Murphy , Joerg Roedel , Thierry Reding , Jonathan Hunter , Jason Gunthorpe , Nicolin Chen , Breno Leitao Cc: "Kiryl Shutsemau (Meta)" , Krishna Reddy , Pranjal Shrivastava , Mostafa Saleh , Ashish Mhetre , Shameer Kolothum , Yuanhe Shu , Kyle McMartin , Usama Arif , kernel-team@meta.com, linux-arm-kernel@lists.infradead.org, iommu@lists.linux.dev, linux-tegra@vger.kernel.org, linux-kernel@vger.kernel.org Subject: [PATCH v7 2/2] iommu/arm-smmu-v3: Default queue depths to one page in a kdump kernel Date: Fri, 25 Sep 2026 15:15:30 +0100 Message-ID: <20260925141532.1274962-3-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260925141532.1274962-1-kirill@shutemov.name> References: <20260925141532.1274962-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" All three queues are sized from the maxima the hardware advertises in IDR1 and allocated at probe, up to 4 MB each on a 4K-page kernel. The capture kernel already disables two of them: arm_smmu_device_reset() drops CR0_EVTQEN and CR0_PRIQEN. It still allocates both at full size. A kdump capture kernel runs from a small crashkernel reservation, and every SMMUv3 instance pays that cost again, up to 12 MB apiece. It goes to queues that either serve the handful of devices used to save the dump or are switched off outright, and it is memory the dump itself needs. Size all three queues at one page worth of entries when is_kdump_kernel(). The queues carry commands and fault records rather than DMA data, so dump throughput is unaffected. A shallower command queue only bounds how many commands may be in flight before a sync, which does not matter for the few devices that save the dump. The cmdq_max_n_shift parameter therefore does not apply in a capture kernel. Suggested-by: Kyle McMartin Assisted-by: LLM Reviewed-by: Jason Gunthorpe Tested-by: Yuanhe Shu Signed-off-by: Kiryl Shutsemau (Meta) Reviewed-by: Breno Leitao Reviewed-by: Nicolin Chen Tested-by: Nicolin Chen --- drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c | 30 ++++++++++++++++----- 1 file changed, 24 insertions(+), 6 deletions(-) diff --git a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c b/drivers/iommu/ar= m/arm-smmu-v3/arm-smmu-v3.c index 810d3ce75089..a5ad57432dfd 100644 --- a/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c +++ b/drivers/iommu/arm/arm-smmu-v3/arm-smmu-v3.c @@ -43,7 +43,7 @@ MODULE_PARM_DESC(disable_msipolling, static unsigned int cmdq_max_n_shift =3D CMDQ_MAX_SZ_SHIFT; module_param(cmdq_max_n_shift, uint, 0444); MODULE_PARM_DESC(cmdq_max_n_shift, - "Cap on the command queue depth, as log2 of the number of entries. Defaul= ts to the hardware maximum; the queue never shrinks below one page."); + "Cap on the command queue depth, as log2 of the number of entries. Defaul= ts to the hardware maximum; the queue never shrinks below one page. A kdump= kernel always uses one page."); =20 static const struct iommu_ops arm_smmu_ops; static struct iommu_dirty_ops arm_smmu_dirty_ops; @@ -4424,6 +4424,8 @@ static struct iommu_dirty_ops arm_smmu_dirty_ops =3D { * @ent_sz_shift: log2 of the queue entry size in bytes * @limit_n_shift: log2 depth to cap the queue at * + * A kdump capture kernel gets one page worth of entries whatever the limi= t. + * * @limit_n_shift is floored at one page, because coherent DMA is page * granular: a shallower queue occupies the same memory as one that fills = the * page, and arm_smmu_init_one_queue() stops shrinking at a page too. @@ -4433,10 +4435,27 @@ static u32 arm_smmu_queue_max_n_shift(u32 hw_max_n_= shift, u32 ent_sz_shift, { u32 floor_n_shift =3D PAGE_SHIFT - ent_sz_shift; =20 + if (is_kdump_kernel()) + return min(hw_max_n_shift, floor_n_shift); + limit_n_shift =3D max(limit_n_shift, floor_n_shift); return min(hw_max_n_shift, limit_n_shift); } =20 +static inline u32 arm_smmu_evtq_max_n_shift(u32 hw_max_n_shift) +{ + /* Capped to ensure natural alignment */ + return arm_smmu_queue_max_n_shift(hw_max_n_shift, EVTQ_ENT_SZ_SHIFT, + EVTQ_MAX_SZ_SHIFT); +} + +static inline u32 arm_smmu_priq_max_n_shift(u32 hw_max_n_shift) +{ + /* Capped to ensure natural alignment */ + return arm_smmu_queue_max_n_shift(hw_max_n_shift, PRIQ_ENT_SZ_SHIFT, + PRIQ_MAX_SZ_SHIFT); +} + /* * Command queues are also allocated by the Tegra241 CMDQV for its VCMDQs,= which * need the same depth decision. @@ -5191,7 +5210,6 @@ static int arm_smmu_device_hw_probe(struct arm_smmu_d= evice *smmu) if (reg & IDR1_ATTR_TYPES_OVR) smmu->features |=3D ARM_SMMU_FEAT_ATTR_TYPES_OVR; =20 - /* Queue sizes, capped to ensure natural alignment */ smmu->cmdq.q.llq.max_n_shift =3D arm_smmu_cmdq_max_n_shift(FIELD_GET(IDR1_CMDQS, reg)); if (smmu->cmdq.q.llq.max_n_shift <=3D ilog2(CMDQ_BATCH_ENTRIES)) { @@ -5206,10 +5224,10 @@ static int arm_smmu_device_hw_probe(struct arm_smmu= _device *smmu) return -ENXIO; } =20 - smmu->evtq.q.llq.max_n_shift =3D min_t(u32, EVTQ_MAX_SZ_SHIFT, - FIELD_GET(IDR1_EVTQS, reg)); - smmu->priq.q.llq.max_n_shift =3D min_t(u32, PRIQ_MAX_SZ_SHIFT, - FIELD_GET(IDR1_PRIQS, reg)); + smmu->evtq.q.llq.max_n_shift =3D + arm_smmu_evtq_max_n_shift(FIELD_GET(IDR1_EVTQS, reg)); + smmu->priq.q.llq.max_n_shift =3D + arm_smmu_priq_max_n_shift(FIELD_GET(IDR1_PRIQS, reg)); =20 /* SID/SSID sizes */ smmu->ssid_bits =3D FIELD_GET(IDR1_SSIDSIZE, reg); --=20 2.54.0