From nobody Tue Sep 29 08:22:34 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 47EB63F9F22; Mon, 10 Aug 2026 16:17:48 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786378668; cv=none; b=jsT5u1+497mmZxcjWRNyKEu/oOAS/8iA9RJveKXBame2UcRjQO2rPVGFg3ZQUeJ2AHTnm0P6R5yc6bo+LtF9p6o5p4vVuBaK89rophESEuwcvIWvsshBH6q/ES404XM6gpCL5czErjHurkHelaYYwK2kddBs/XwzpiySy8or1FM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786378668; c=relaxed/simple; bh=Q+qdx6oRKdXlA4l6g9eyHfPrIoqSy/Lo6Jb4tRgnhO0=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=lJNFiHjWeW/yCNWVBZKfi4X10Xzap9SKqOTSKvY1vU3AjWAU5E7KoXOKd2PkeJrCHTjwUd1wEvBIafUmvTYuwZkOmWXfREjb/aO+pb1KM8lJ+cBU1bvFA0rBk2QMWv860N53b630+4Pxx53EPemz7MofXtdP14VquEgIRHSMNEQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=I76pR+Gt; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="I76pR+Gt" Received: by smtp.kernel.org (Postfix) with ESMTPS id CA463C19425; Mon, 10 Aug 2026 16:17:47 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1786378667; bh=Q+qdx6oRKdXlA4l6g9eyHfPrIoqSy/Lo6Jb4tRgnhO0=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=I76pR+GtsYmIxx7++WdnINkB0HjzufkoAYoEPtGL5tIiBsXHrK8EzdHKnKDfyK1VX 6EIoWO/5fn7z38/9XY/mAQ2/Dgjcsei8O/dQLFp3grFiU+TwIyKnCiOEvItOUG/KCO L2kIeCTlU51WKgL6HLUETgNJHxjRUSDqmtMgshlZFjQ4hVlDjaa7xAPYXqskgrCAAe +MQa+fRZYP8akGvJn/FrT2iLwdnlhKPOeKQSlvAdUBfhc7KgSlOS3e/j8bBy7CGKW9 LFzaP9qUrK/zXTlAsxUzNtVHMj2ern3W2bS0akAC/c2Vh1fjaHM/2JvxOPIRTyP1mQ uC14XmIEeiU1Q== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id AB4E7C5AD55; Mon, 10 Aug 2026 16:17:47 +0000 (UTC) From: Junrui Luo via B4 Relay Date: Tue, 11 Aug 2026 00:13:10 +0800 Subject: [PATCH 1/5] drm/amdgpu: free prt_va on the open_kms error path Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260811-amdgpu-fixes-v1-1-4954a417b8ff@outlook.com> References: <20260811-amdgpu-fixes-v1-0-4954a417b8ff@outlook.com> In-Reply-To: <20260811-amdgpu-fixes-v1-0-4954a417b8ff@outlook.com> To: Alex Deucher , =?utf-8?q?Christian_K=C3=B6nig?= , David Airlie , Simona Vetter , Sumit Semwal , Junwei Zhang , =?utf-8?q?Nicolai_H=C3=A4hnle?= , Prike Liang , Arvind Yadav , Shashank Sharma , Leo Liu , Felix Kuehling Cc: amd-gfx@lists.freedesktop.org, dri-devel@lists.freedesktop.org, linux-kernel@vger.kernel.org, linux-media@vger.kernel.org, linaro-mm-sig@lists.linaro.org, Junrui Luo , Yuhao Jiang X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=openpgp-sha256; l=1692; i=moonafterrain@outlook.com; h=from:subject:message-id; bh=q+sGDSGgjTe169BOxCoszY1QMkkcItYVcawWxv7vRl8=; b=owJ4nJvAy8zAJVb4wiKgu++DA+NptSSGrMqfS21uv6tOUDx47X7r4QuP/FZ8ClS7bTFpRZEoX +RZK0cG9yMdpSwMYlwMsmKKLMcLLn2z8N2iu8VnSzLMHFYmkCEMXJwCMBGWTob/xa1RtyoWnp8X 01Y5487D+wGaU7rzbSqCD6xP/M5yOVx8FsN/zx72XsaG69aXLBa+bud7evlC06pi4xcHHzEwysW EbN3DCAAXk07f X-Developer-Key: i=moonafterrain@outlook.com; a=openpgp; fpr=C770D2F6384DB42DB44CB46371E838508B8EF040 X-Endpoint-Received: by B4 Relay for moonafterrain@outlook.com/default with auth_id=909 X-Original-From: Junrui Luo Reply-To: moonafterrain@outlook.com From: Junrui Luo amdgpu_driver_open_kms() creates fpriv->prt_va with amdgpu_vm_bo_add() before mapping the CSA and the seq64 buffer. If either mapping fails the function jumps to error_vm, which only calls amdgpu_vm_fini() and then frees fpriv. amdgpu_vm_fini() releases the amdgpu_bo_va_mapping objects reachable from vm->freed and the vm->va rbtree, but it never frees a struct amdgpu_bo_va, so the bo_va allocated for prt_va and the dma_fence stub reference it holds are both lost. The success path does get this right: amdgpu_driver_postclose_kms() reserves the root PD and calls amdgpu_vm_bo_del(adev, fpriv->prt_va) before amdgpu_vm_fini(). Only the open() unwind is missing it. Drop the bo_va on the error path as well, reserving the root PD as amdgpu_vm_bo_del() requires. Fixes: b85891bd6d1b ("drm/amdgpu: IOCTL interface for PRT support v4") Reported-by: Yuhao Jiang Assisted-by: Claude:claude-opus-5 Signed-off-by: Junrui Luo --- drivers/gpu/drm/amd/amdgpu/amdgpu_kms.c | 5 +++++ 1 file changed, 5 insertions(+) diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_kms.c b/drivers/gpu/drm/amd/= amdgpu/amdgpu_kms.c index 242c48e85912..7ef1c1dcc207 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_kms.c +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_kms.c @@ -1553,6 +1553,11 @@ int amdgpu_driver_open_kms(struct drm_device *dev, s= truct drm_file *file_priv) pasid =3D 0; } =20 + if (fpriv->prt_va && + !WARN_ON(amdgpu_bo_reserve(fpriv->vm.root.bo, true))) { + amdgpu_vm_bo_del(adev, fpriv->prt_va); + amdgpu_bo_unreserve(fpriv->vm.root.bo); + } amdgpu_vm_fini(adev, &fpriv->vm); =20 error_pasid: --=20 2.51.2 From nobody Tue Sep 29 08:22:34 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 47BCB3F86F0; Mon, 10 Aug 2026 16:17:48 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786378668; cv=none; b=esvyE6XUTgFDirRehbfPx2rwusNeS9UvEJ9lNpSHDhrIlgdi3cHp8TMvG7b9SzPMjua6PiZBsGxzIIHxapCaVqMGrZ7oxxz0HPKGo68iq8031zt5GSYSG7IB/8gy8mj/gw1pbH0NddgFPFAiTKMXWaQvAvj9wPmhtap8uowaeVM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786378668; c=relaxed/simple; bh=TtwXuOo7f/6QzG98rp5LXPITn3SYcjaZqv2NU1Zo/2M=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=boYCPxNkPSEi532xKlbfeTOq8d1yVrCD9w15VfG1ggGGRXlrfi9JoSUu/SDH9jf0YNZcOxuIr9UjOckLABt6wZrfG720sq5+dZGk0BBR48Oq0/bub9dsHPVb75jNYp6Gt8AKZtJAgsFveO0SGyTc7HMHkZBtORATehahJ5UBqaY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=cASpjYhd; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="cASpjYhd" Received: by smtp.kernel.org (Postfix) with ESMTPS id DB0C6C2BCF7; Mon, 10 Aug 2026 16:17:47 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1786378667; bh=TtwXuOo7f/6QzG98rp5LXPITn3SYcjaZqv2NU1Zo/2M=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=cASpjYhdIVf0owsIGF8yCYwZQ3IeDnTbp0qffmKhq+RzZRsJ8gwg9MMdMibzLnlJV uGpq8UcU7DTtT6CZyY8tOHHQEVQwLel8jWb1CT1tZwF2x9FsOxuwCm6VmAO04fdxFR 09hUZEUWe9FpOArXNKFPX4gvyoMuvJArv17DmWFvf87OaTOY7i/TnkaHX95vLyNg94 sQWXxmb0ly5a5xFoVEoVNiq8CCLZHfmoXm+xYOAgZNf+2kdcPAf7TuwN3Nmk+BI8Ke 3dRzdzYgSkxW77ceUN9jHx3xirPu8egARiWeBs5hTxwyPc2Wjva9wPkqm4RbuUxIEy T/baJexq4Wh+w== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id BCCEAC5B56D; Mon, 10 Aug 2026 16:17:47 +0000 (UTC) From: Junrui Luo via B4 Relay Date: Tue, 11 Aug 2026 00:13:11 +0800 Subject: [PATCH 2/5] drm/amdgpu: reject PRT mappings as user queue buffer VAs Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260811-amdgpu-fixes-v1-2-4954a417b8ff@outlook.com> References: <20260811-amdgpu-fixes-v1-0-4954a417b8ff@outlook.com> In-Reply-To: <20260811-amdgpu-fixes-v1-0-4954a417b8ff@outlook.com> To: Alex Deucher , =?utf-8?q?Christian_K=C3=B6nig?= , David Airlie , Simona Vetter , Sumit Semwal , Junwei Zhang , =?utf-8?q?Nicolai_H=C3=A4hnle?= , Prike Liang , Arvind Yadav , Shashank Sharma , Leo Liu , Felix Kuehling Cc: amd-gfx@lists.freedesktop.org, dri-devel@lists.freedesktop.org, linux-kernel@vger.kernel.org, linux-media@vger.kernel.org, linaro-mm-sig@lists.linaro.org, Junrui Luo , Yuhao Jiang , stable@vger.kernel.org X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=openpgp-sha256; l=2330; i=moonafterrain@outlook.com; h=from:subject:message-id; bh=lfRKX90eM1yWKlq0c1LdgUl6q5HeCXaJonPRnO3TU20=; b=owJ4nJvAy8zAJVb4wiKgu++DA+NptSSGrMqfS8MzkkW8LcUEFkwJaFi1V09slfqyG/p2N5wes iw49vLaimkdpSwMYlwMsmKKLMcLLn2z8N2iu8VnSzLMHFYmkCEMXJwCMBFRB4b/SdMkBBe5RRtO Kl8g8cNstk7TrfXJ4Q4LbT0f2EaHsxRdYGS4rlpeoZampVt1JDRg451N3yeHZ+1T8PzAf9bk7c5 PjdbcAPu3SQM= X-Developer-Key: i=moonafterrain@outlook.com; a=openpgp; fpr=C770D2F6384DB42DB44CB46371E838508B8EF040 X-Endpoint-Received: by B4 Relay for moonafterrain@outlook.com/default with auth_id=909 X-Original-From: Junrui Luo Reply-To: moonafterrain@outlook.com From: Junrui Luo amdgpu_userq_input_va_validate() resolves a user-supplied queue_va, rptr_va or wptr_va to a VM mapping and latches userq_va_mapped on the owning bo_va. It only checks that a mapping exists and that the requested span is contained in it, never that the mapping has a backing BO. PRT mappings do not: amdgpu_gem_va_ioctl() routes every AMDGPU_VM_PAGE_PRT map through fpriv->prt_va, created via amdgpu_vm_bo_add(adev, vm, NULL), so base.bo stays NULL while amdgpu_vm_bo_insert_map() still sets mapping->bo_va. A VA inside such a mapping therefore passes validation and marks fpriv->prt_va as userq mapped. The flag is never cleared. On the next unmap of any PRT mapping in that VM, amdgpu_vm_bo_unmap() sees userq_va_mapped and calls amdgpu_userq_gem_va_unmap_validate(), which reads bo_va->base.bo->tbo.base.resv before its ip_mask guard, leading to a NULL pointer dereference. Fix by rejecting a mapping without a backing BO in the validation helper, so the invariant amdgpu_userq_gem_va_unmap_validate() relies on holds by construction. A sparse mapping has no memory behind it and cannot serve as a ring, rptr or wptr buffer. Fixes: 2e7ceac0ea41 ("drm/amdgpu: validate userq va for GEM unmap") Reported-by: Yuhao Jiang Assisted-by: Claude:claude-opus-5 Cc: stable@vger.kernel.org Signed-off-by: Junrui Luo --- drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c | 8 ++++++++ 1 file changed, 8 insertions(+) diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c b/drivers/gpu/drm/am= d/amdgpu/amdgpu_userq.c index 6d3ed55e9ab4..bec107216811 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c @@ -259,6 +259,14 @@ int amdgpu_userq_input_va_validate(struct amdgpu_devic= e *adev, if (!va_map) return -EINVAL; =20 + /* + * A PRT mapping has no backing BO and so can't carry the eviction + * fence which amdgpu_userq_gem_va_unmap_validate() waits on. Reject it + * here, otherwise that helper dereferences a NULL bo on GEM unmap. + */ + if (!va_map->bo_va->base.bo) + return -EINVAL; + /* Lookup guarantees start_page is mapped; ensure full span is covered. */ if ((end_addr >> AMDGPU_GPU_PAGE_SHIFT) <=3D va_map->last) { va_map->bo_va->userq_va_mapped =3D true; --=20 2.51.2 From nobody Tue Sep 29 08:22:34 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 47CED3F8886; Mon, 10 Aug 2026 16:17:48 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786378668; cv=none; b=M1qzLYx9VmtDRQgIPdKFVjQM2GZIK6ttX2TqUdAaCYd8dKvokrGW5M6qPJxY+k5tZ0KvtKCbUQ6doP54fvA4RCYhpvPU/L5WBRO+gcox/QdjpGBKiU8U0JBU9er/AUNAREYBQW2eM+tgzjy+zNWP48X/dv2Pq8DSHwznTPMhamk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786378668; c=relaxed/simple; bh=HkNyNZT7BLjrADpqU1n8ztglAus49jq8Yw+tgDgVUvU=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=Vkqg7M8QtjXwHwnT9pEp1HFd80Agg3JWTBhypL9cAI+UC+vl7INK9Y0e0X4SX8zdHXSgfUrPQqwtCvhug28vQ4gZx0sqcyXnL/OQT4m/gCLSh1bvRkYIuOrBYkJjgqeBnyVgITo1wITC9Uevda7NC7wq6WJBI1/9u+KuAF+T+10= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=i2jpmx4Q; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="i2jpmx4Q" Received: by smtp.kernel.org (Postfix) with ESMTPS id E5EC8C2BCFA; Mon, 10 Aug 2026 16:17:47 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1786378667; bh=HkNyNZT7BLjrADpqU1n8ztglAus49jq8Yw+tgDgVUvU=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=i2jpmx4Q4ebk8BEoBUwdjKOUO0z5Z0VxWoyglwAW20nUSmG3OlfLeLxyl2fvFt7jP xYKFvpw3+4j8uai+CcIpp4SnhfuiCrjHa6RZnX5AaZ53Io/9NqCNNlQrgTDf4BDBDx W6QsqdfVUOqzjSts31CxOLyhrBupV+76pNRR7I6cLP+PJwuGd8vcbc6pXwhBG60qNI gGs7Y78WN/YX7pA/1i64DESaTUDqM1zzggxPUa/wVc5EJSNJ3fl6S5cvZbLUcmOyY3 11Y3ApOKz/NfiWAhlosKzjkJW5ai4cEpw8HHVZ4c12ft0hKIeGaXLZLfUXOVsuovs1 5jSNe7FiF0ZxQ== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id CDDD9C5AD7B; Mon, 10 Aug 2026 16:17:47 +0000 (UTC) From: Junrui Luo via B4 Relay Date: Tue, 11 Aug 2026 00:13:12 +0800 Subject: [PATCH 3/5] drm/amdgpu/userq: bound the eviction fence rearm retry loop Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260811-amdgpu-fixes-v1-3-4954a417b8ff@outlook.com> References: <20260811-amdgpu-fixes-v1-0-4954a417b8ff@outlook.com> In-Reply-To: <20260811-amdgpu-fixes-v1-0-4954a417b8ff@outlook.com> To: Alex Deucher , =?utf-8?q?Christian_K=C3=B6nig?= , David Airlie , Simona Vetter , Sumit Semwal , Junwei Zhang , =?utf-8?q?Nicolai_H=C3=A4hnle?= , Prike Liang , Arvind Yadav , Shashank Sharma , Leo Liu , Felix Kuehling Cc: amd-gfx@lists.freedesktop.org, dri-devel@lists.freedesktop.org, linux-kernel@vger.kernel.org, linux-media@vger.kernel.org, linaro-mm-sig@lists.linaro.org, Junrui Luo , Yuhao Jiang , stable@vger.kernel.org X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=openpgp-sha256; l=5540; i=moonafterrain@outlook.com; h=from:subject:message-id; bh=pc/SsaipZ7NjkbbEcRk7SSITnHVNauzCNXg3L1I8WBA=; b=owJ4nJvAy8zAJVb4wiKgu++DA+NptSSGrMqfSz92XTihnLXcXp3xqavki9yqH1L2sVO2HFpjM +XilKVyuskdpSwMYlwMsmKKLMcLLn2z8N2iu8VnSzLMHFYmkCEMXJwCMBGBmwz/ww/FVdWI7Mpu +JisIs0z7010snVrz2THoHUyi+O2u/rlMDI07BfkelX35s67M4qbmO7UcvAH1GuneZfsmPoxYfl jO1Z+AAhUSms= X-Developer-Key: i=moonafterrain@outlook.com; a=openpgp; fpr=C770D2F6384DB42DB44CB46371E838508B8EF040 X-Endpoint-Received: by B4 Relay for moonafterrain@outlook.com/default with auth_id=909 X-Original-From: Junrui Luo Reply-To: moonafterrain@outlook.com From: Junrui Luo amdgpu_userq_ensure_ev_fence() loops until the eviction fence is both present and unsignaled. The only producer of such a fence is amdgpu_evf_mgr_rearm(), which runs as the very last step of amdgpu_userq_vm_validate(). Every failure point ahead of it - the kzalloc() in the rearm itself, amdgpu_hmm_range_alloc(), the ttm_bo_validate() calls, the GART binding of the wptr BOs - makes amdgpu_userq_restore_worker() give up with only a drm_file_err(). Nothing propagates that back, so the waiting thread reschedules the worker and flushes it again, forever. Both flush_delayed_work() and mutex_lock() sleep in TASK_UNINTERRUPTIBLE, so the looping task cannot be killed and the OOM killer cannot reclaim it. An unprivileged render node client reaches this from both AMDGPU_USERQ and AMDGPU_USERQ_SIGNAL. The eviction fence sequence number is already bumped by every successful rearm, so use it as the loop's progress condition: if a completed flush of the restore worker did not move it then no rearm happened and retrying cannot help. Return -ENOMEM in that case and let both callers report it to userspace. Fixes: a242a3e4b5be ("drm/amdgpu: simplify eviction fence suspend/resume") Reported-by: Yuhao Jiang Assisted-by: Claude:claude-opus-5 Cc: stable@vger.kernel.org Signed-off-by: Junrui Luo --- drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c | 21 +++++++++++++++++++-- drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h | 4 ++-- drivers/gpu/drm/amd/amdgpu/amdgpu_userq_fence.c | 10 +++++++++- 3 files changed, 30 insertions(+), 5 deletions(-) diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c b/drivers/gpu/drm/am= d/amdgpu/amdgpu_userq.c index bec107216811..208b53ae5bd1 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c @@ -448,12 +448,16 @@ static void amdgpu_userq_cleanup(struct amdgpu_usermo= de_queue *queue) * Ensures that a valid and not yet signaled eviction fence is attached to= the * usermode queue before any queue operations proceed. If it is signalled,= then * rearm a new eviction fence. + * + * Returns 0 with @uq_mgr->userq_mutex held, or -ENOMEM with the mutex rel= eased + * when the restore worker could not rearm the fence. */ -void +int amdgpu_userq_ensure_ev_fence(struct amdgpu_userq_mgr *uq_mgr, struct amdgpu_eviction_fence_mgr *evf_mgr) { struct dma_fence *ev_fence; + int seq, prev_seq =3D -1; =20 retry: /* Flush any pending resume work to create ev_fence */ @@ -463,7 +467,16 @@ amdgpu_userq_ensure_ev_fence(struct amdgpu_userq_mgr *= uq_mgr, ev_fence =3D amdgpu_evf_mgr_get_fence(evf_mgr); if (dma_fence_is_signaled(ev_fence)) { dma_fence_put(ev_fence); + seq =3D atomic_read(&evf_mgr->ev_fence_seq); mutex_unlock(&uq_mgr->userq_mutex); + /* + * The sequence number is only bumped by a successful rearm, so + * if the flush above ran the worker without moving it then the + * restore failed and looping again would never terminate. + */ + if (seq =3D=3D prev_seq) + return -ENOMEM; + prev_seq =3D seq; /* * Looks like there was no pending resume work, * add one now to create a valid eviction fence @@ -472,6 +485,8 @@ amdgpu_userq_ensure_ev_fence(struct amdgpu_userq_mgr *u= q_mgr, goto retry; } dma_fence_put(ev_fence); + + return 0; } =20 =20 @@ -747,7 +762,9 @@ amdgpu_userq_create(struct drm_file *filp, union drm_am= dgpu_userq *args) if (r) goto clean_mqd; =20 - amdgpu_userq_ensure_ev_fence(&fpriv->userq_mgr, &fpriv->evf_mgr); + r =3D amdgpu_userq_ensure_ev_fence(&fpriv->userq_mgr, &fpriv->evf_mgr); + if (r) + goto erase_doorbell; =20 /* don't map the queue if scheduling is halted */ if (!adev->userq_halt_for_enforce_isolation || diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h b/drivers/gpu/drm/am= d/amdgpu/amdgpu_userq.h index 6412a7f7b6ef..c35909bf7ceb 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h @@ -164,8 +164,8 @@ void amdgpu_userq_mgr_fini(struct amdgpu_userq_mgr *use= rq_mgr); =20 void amdgpu_userq_evict(struct amdgpu_userq_mgr *uq_mgr); =20 -void amdgpu_userq_ensure_ev_fence(struct amdgpu_userq_mgr *userq_mgr, - struct amdgpu_eviction_fence_mgr *evf_mgr); +int amdgpu_userq_ensure_ev_fence(struct amdgpu_userq_mgr *userq_mgr, + struct amdgpu_eviction_fence_mgr *evf_mgr); =20 u32 amdgpu_userq_get_supported_ip_mask(struct amdgpu_device *adev); bool amdgpu_userq_enabled(struct drm_device *dev); diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq_fence.c b/drivers/gpu/= drm/amd/amdgpu/amdgpu_userq_fence.c index 7e80442ec3e5..1c287ce59736 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq_fence.c +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq_fence.c @@ -523,7 +523,15 @@ int amdgpu_userq_signal_ioctl(struct drm_device *dev, = void *data, goto put_queue; =20 /* We are here means UQ is active, make sure the eviction fence is valid = */ - amdgpu_userq_ensure_ev_fence(&fpriv->userq_mgr, &fpriv->evf_mgr); + r =3D amdgpu_userq_ensure_ev_fence(&fpriv->userq_mgr, &fpriv->evf_mgr); + if (r) { + /* The fence is not initialized yet, so unwind it by hand */ + amdgpu_userq_fence_put_fence_drv_array(fence); + amdgpu_userq_fence_driver_put(fence->fence_drv); + kvfree(fence->fence_drv_array); + kfree(fence); + goto put_queue; + } =20 /* Create the new fence */ amdgpu_userq_fence_init(queue, fence, wptr); --=20 2.51.2 From nobody Tue Sep 29 08:22:34 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 47B043F660F; Mon, 10 Aug 2026 16:17:48 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786378668; cv=none; b=Glj9dOvewQYxBAqu4b0gU6KRC5doM2AL6/QyGd7jwO/gOaEiwh9h2xGled5RwyFL0lq+f7SqTmj9HrQrX802Htid4VwZwsUZoFaoDFk+qd+EtX7cmRz6hUf1qoUV+35S6j88pQyx+OueKXuGCSrUJufd1ckN+mZ5/D2kTUz1P3c= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786378668; c=relaxed/simple; bh=qrWihAGCrqFkMcNdLvUZTQMgUT+Pgkggyran7rTHm9M=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=SvOTmGJX3LRNufdTuavcsRLtJnMyqrxEu8VbkQyggZ2l7LMJM4WNPx/ynUIZkmNePAU8keZts7+ScQ6pUbp1YRarjeHlDxaM4r3ONZetfxDyd9jtMLWyJpnv9ExJv/M063reivDVWz2cLU2+HqxK2u++FHlLJ9GjYy2n5923Jz0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=Rxn/c43A; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="Rxn/c43A" Received: by smtp.kernel.org (Postfix) with ESMTPS id 0459BC2BCFC; Mon, 10 Aug 2026 16:17:48 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1786378668; bh=qrWihAGCrqFkMcNdLvUZTQMgUT+Pgkggyran7rTHm9M=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=Rxn/c43A0RClcJ3wW3eopMFE4AD5cbJ1MV+B6K+ukP/gjaBjPM+V/EkXOHI1/E3K/ +u3t0cMH1k2fU1kg8s79UhC4mm0udFDIb5Y+X6iJZIO/8PwaznQAgfIi3Udu9WNzw0 mHXCWhgk2PIKxozXWm79EcCRb9QOMj6YphU3aMRi0socbjfFGRLywBgbTIOYpt+xz/ Km7O88SSXCQqY1APsRbO1/1WrePt2P2hlwLfK8TZEOfAdpCdaO9FrcJNyEkEWg8W6L KOMhjUAEN69wqLxvmogNjSdUOAGMBXFv2VSKOi6OAQggkAVP1YGwXltHWVDlTQaj/p vvoR3vl3h1m7g== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id DDD8AC5B572; Mon, 10 Aug 2026 16:17:47 +0000 (UTC) From: Junrui Luo via B4 Relay Date: Tue, 11 Aug 2026 00:13:13 +0800 Subject: [PATCH 4/5] drm/amdgpu: enforce UVD handle ownership on destroy Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260811-amdgpu-fixes-v1-4-4954a417b8ff@outlook.com> References: <20260811-amdgpu-fixes-v1-0-4954a417b8ff@outlook.com> In-Reply-To: <20260811-amdgpu-fixes-v1-0-4954a417b8ff@outlook.com> To: Alex Deucher , =?utf-8?q?Christian_K=C3=B6nig?= , David Airlie , Simona Vetter , Sumit Semwal , Junwei Zhang , =?utf-8?q?Nicolai_H=C3=A4hnle?= , Prike Liang , Arvind Yadav , Shashank Sharma , Leo Liu , Felix Kuehling Cc: amd-gfx@lists.freedesktop.org, dri-devel@lists.freedesktop.org, linux-kernel@vger.kernel.org, linux-media@vger.kernel.org, linaro-mm-sig@lists.linaro.org, Junrui Luo , Yuhao Jiang , stable@vger.kernel.org X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=openpgp-sha256; l=2299; i=moonafterrain@outlook.com; h=from:subject:message-id; bh=UtGNeM13NSreEA2vgkQPwwzgGSG/EJ+HK2TwqY3ru6I=; b=owJ4nJvAy8zAJVb4wiKgu++DA+NptSSGrMqfS40VjhmWnd2/6hD3na7ZrbHHfV8FbmefWj17e lyFd6J4TWdHKQuDGBeDrJgiy/GCS98sfLfobvHZkgwzh5UJZAgDF6cATMQrieGffnHOVano9rq9 HMXn1TZyyjWGrQ8Ji//dz/WvaKP8bcHdjAynyuLigxq3Zwcmm+R/+VGSxGH59HiN/utz+bN4Pwe s+MQEAO2LS68= X-Developer-Key: i=moonafterrain@outlook.com; a=openpgp; fpr=C770D2F6384DB42DB44CB46371E838508B8EF040 X-Endpoint-Received: by B4 Relay for moonafterrain@outlook.com/default with auth_id=909 X-Original-From: Junrui Luo Reply-To: moonafterrain@outlook.com From: Junrui Luo amdgpu_uvd_cs_msg() validates that a decode message references a handle owned by the submitting client, rejecting a mismatch between adev->uvd.filp[i] and ctx->parser->filp. The handles[] and filp[] tables are per-device and shared by every drm_file that opens the render node. The destroy message performs no such check: it walks the whole table and clears every slot matching the handle taken from the command stream buffer. A client can therefore destroy a handle owned by another client, clearing the victim's slot and tearing down its session in UVD firmware, so subsequent decode submissions fail with -ENOENT. Since amdgpu_uvd_free_handles() only reaps slots whose handle is non-zero, the cleared slot also retains a stale filp pointer until reused. Apply the decode arm's ownership test to the destroy arm. The kunmap is hoisted above the loop, matching the create and decode arms, so the new error return cannot leak the amdgpu_bo_kmap() reference. Kernel-initiated teardown goes through amdgpu_uvd_send_msg() and never runs the parser. Fixes: 5146419e6feb ("drm/amdgpu: make UVD handle checking more strict") Reported-by: Yuhao Jiang Assisted-by: Claude:claude-opus-5 Cc: stable@vger.kernel.org Signed-off-by: Junrui Luo --- drivers/gpu/drm/amd/amdgpu/amdgpu_uvd.c | 14 ++++++++++++-- 1 file changed, 12 insertions(+), 2 deletions(-) diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_uvd.c b/drivers/gpu/drm/amd/= amdgpu/amdgpu_uvd.c index e8b0c62f72be..8d3e5435cf52 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_uvd.c +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_uvd.c @@ -918,9 +918,19 @@ static int amdgpu_uvd_cs_msg(struct amdgpu_uvd_cs_ctx = *ctx, =20 case 2: /* it's a destroy msg, free the handle */ - for (i =3D 0; i < adev->uvd.max_handles; ++i) - atomic_cmpxchg(&adev->uvd.handles[i], handle, 0); amdgpu_bo_kunmap(bo); + + for (i =3D 0; i < adev->uvd.max_handles; ++i) { + if (atomic_read(&adev->uvd.handles[i]) !=3D handle) + continue; + + if (adev->uvd.filp[i] !=3D ctx->parser->filp) { + DRM_ERROR("UVD handle collision detected!\n"); + return -EINVAL; + } + + atomic_cmpxchg(&adev->uvd.handles[i], handle, 0); + } return 0; =20 default: --=20 2.51.2 From nobody Tue Sep 29 08:22:34 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AC4CD4052B6; Mon, 10 Aug 2026 16:17:48 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786378668; cv=none; b=DPleSHhn3HMeT8z31JXuZfsVjHdYUzhV/eLDh8TPfCbPxO7yKJQm0iumGceHMLHbFjZLAq77yGiedFJfKU/eJ6HjP3wJh6DhVTNjGlbJy5ezSlTaDIJ9JZRmgVy/7FFpOobX4UyEhyGveUKyHX3REdcippB35fcepZtZN6VDdKk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786378668; c=relaxed/simple; bh=e7qU4DwQRhkwQZ7bRxaFZbYgsVTpDVtdRWvEtA7dFMA=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=FSgPN5o2edp+P2MtkrG28kHJ6/WrHRy6Y1DkNcqYh2gGvZO/c6JmgTRsE4tw7QhHOG8v+vRdAAHBbUGqfbFJu5BkkEn4jCEoZ30ho7ze7vV2js34nI5eBE8OnEL/6rn9w9nzAGSd5vjJuKfgWZH9RtdTGhnP95RcWF8MCKSiHl4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=N2yDUZHU; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="N2yDUZHU" Received: by smtp.kernel.org (Postfix) with ESMTPS id 0D024C2BCFB; Mon, 10 Aug 2026 16:17:48 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1786378668; bh=e7qU4DwQRhkwQZ7bRxaFZbYgsVTpDVtdRWvEtA7dFMA=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=N2yDUZHUC+IiL1ijHjfMJnnBZdHY8+FrZXfujPqROiUzMIRwWzXXQaV0pg+4b71R+ iLgvFXIoU9QJrpVoLy39ERqqLbjpwDV7w9vRyR7wmaCmjOrVoTVv27X0IgqzqfGGuQ tVztO48l8TE0ACnyKtMlo2J8zOolDMwREwqXaQSxmZ50QOF7DeB1yy8bE19qRyy0xN 9Z1wKjyXwfKZn3yXy5imLzmeRwaKSr5OysXS3NV7zG5u8OjY6gfbp6qp3TvNuV4k6E qmzIFa4OWQj3tUwc6QcvIXj+oYjrCaXeFigq31AsfE/DHtTiceqomoqw0V4pn8EAWJ kWCkQOvlzruJA== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id EDE7DC5B56A; Mon, 10 Aug 2026 16:17:47 +0000 (UTC) From: Junrui Luo via B4 Relay Date: Tue, 11 Aug 2026 00:13:14 +0800 Subject: [PATCH 5/5] drm/amdgpu: free userptr HMM ranges on the CS error path Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260811-amdgpu-fixes-v1-5-4954a417b8ff@outlook.com> References: <20260811-amdgpu-fixes-v1-0-4954a417b8ff@outlook.com> In-Reply-To: <20260811-amdgpu-fixes-v1-0-4954a417b8ff@outlook.com> To: Alex Deucher , =?utf-8?q?Christian_K=C3=B6nig?= , David Airlie , Simona Vetter , Sumit Semwal , Junwei Zhang , =?utf-8?q?Nicolai_H=C3=A4hnle?= , Prike Liang , Arvind Yadav , Shashank Sharma , Leo Liu , Felix Kuehling Cc: amd-gfx@lists.freedesktop.org, dri-devel@lists.freedesktop.org, linux-kernel@vger.kernel.org, linux-media@vger.kernel.org, linaro-mm-sig@lists.linaro.org, Junrui Luo , Yuhao Jiang , stable@vger.kernel.org X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=openpgp-sha256; l=2404; i=moonafterrain@outlook.com; h=from:subject:message-id; bh=L6QMfmjG08GEYgq6Lvgjn+rT+rJCq+Q7xLdhBvY9+GU=; b=owJ4nJvAy8zAJVb4wiKgu++DA+NptSSGrMqfy6RfzJuy89+X84usVeOvsxZ2dRd+fvO2Y3GL7 G0m597euRkdpSwMYlwMsmKKLMcLLn2z8N2iu8VnSzLMHFYmkCEMXJwCMJHHnxkZvjsd3j4/vThq p2rz6jjX/GVveLbFf9zL6NGSKFRhWnulg5Fh2dX9fdNlpNndDOxbPjf+kJn0MMVvVnP40Zi8st8 xCrz8AHwXTaI= X-Developer-Key: i=moonafterrain@outlook.com; a=openpgp; fpr=C770D2F6384DB42DB44CB46371E838508B8EF040 X-Endpoint-Received: by B4 Relay for moonafterrain@outlook.com/default with auth_id=909 X-Original-From: Junrui Luo Reply-To: moonafterrain@outlook.com From: Junrui Luo amdgpu_cs_parser_bos() allocates a struct amdgpu_hmm_range for every userptr entry of the BO list and returns with them live. They are only released in two places: the out_free_user_pages label in amdgpu_cs_parser_bos() itself, and the invalidation check loop in amdgpu_cs_submit(). Every error edge between those two points leaks. A failure in amdgpu_cs_patch_jobs(), amdgpu_cs_vm_handling() or amdgpu_cs_sync_rings(), or an early return from amdgpu_cs_submit() before its release loop, jumps to error_fini and falls into amdgpu_cs_parser_fini(), which never walks the BO list for userptr ranges. An IB address with no VM mapping is enough to get there: amdgpu_cs_patch_ibs() returns the -EINVAL that amdgpu_cs_find_mapping() hands back, so the leak is repeatable at will from an unprivileged render node fd. Each leaked entry costs a struct amdgpu_hmm_range plus its hmm_pfns array, a kvmalloc_array() of one entry per page of the userptr mapping, allocated with plain GFP_KERNEL and so not charged to the caller's memory cgroup. Release the ranges in amdgpu_cs_parser_fini(), which every path out of amdgpu_cs_ioctl() passes through. amdgpu_hmm_range_free() ignores a NULL range, so the success path, where amdgpu_cs_submit() has already freed and cleared them, is unaffected. Fixes: fec8fdb54e8f ("drm/amdgpu: fix userptr HMM range handling v2") Reported-by: Yuhao Jiang Assisted-by: Claude:claude-opus-5 Cc: stable@vger.kernel.org Signed-off-by: Junrui Luo --- drivers/gpu/drm/amd/amdgpu/amdgpu_cs.c | 10 ++++++++++ 1 file changed, 10 insertions(+) diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_cs.c b/drivers/gpu/drm/amd/a= mdgpu/amdgpu_cs.c index 617f53f135f3..17c4fec21402 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_cs.c +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_cs.c @@ -1416,6 +1416,16 @@ static void amdgpu_cs_parser_fini(struct amdgpu_cs_p= arser *parser) amdgpu_vm_bo_invalidate(bo, false); } } + + /* + * Release the ranges still live on the error paths; + * amdgpu_cs_submit() already freed and cleared them when it + * got far enough to check them for invalidation. + */ + amdgpu_bo_list_for_each_userptr_entry(e, parser->bo_list) { + amdgpu_hmm_range_free(e->range); + e->range =3D NULL; + } amdgpu_bo_list_put(parser->bo_list); } =20 --=20 2.51.2