From nobody Mon Sep 28 12:33:06 2026 Received: from mail-ej1-f50.google.com (mail-ej1-f50.google.com [209.85.218.50]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7532D4C9557 for ; Fri, 21 Aug 2026 15:23:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.218.50 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787325791; cv=none; b=HHws9ASJN6YUR1qy5anhUy6PUhHTLmYC+QxbY1fVV9KQy1uDx9+Z+RDf0SDbipeCCYnPi6/6Mrgo0AmmbjJwHuUXbrxVoq7/+7QXdDIzYXjW5Di32wJSMiihvnfsSxf0eHA5hFsII4GtXdS13NTQjKQwH4WxBbOBsDwUzSeQyj8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787325791; c=relaxed/simple; bh=ZErlHBib0LaSTq6e7b9h8OwF5+h7jSxTg3+J3DsZHqs=; h=From:Date:Subject:To:Cc:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=gthOhQJJ033sFxmmFvH0D5vnJoX/zdK1tWJpFzcU3lXmK4AX5nRyGau9+0FHlABJpO9BnHD5Fv4aojr+S5emZm6DkvftbgJ69Ixuo4emKRur4Etzcenk/0DXt6lhiDl4pV7bPdUCuJyhSeqTzhJ9dPFVXnCQkNs2nQDS6jWi0hg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=XarpU2sY; arc=none smtp.client-ip=209.85.218.50 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="XarpU2sY" Received: by mail-ej1-f50.google.com with SMTP id a640c23a62f3a-c15fc4707f5so10877466b.3 for ; Fri, 21 Aug 2026 08:23:09 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787325788; x=1787930588; darn=vger.kernel.org; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:cc:to:subject:date:from:from:to:cc:subject :date:message-id:reply-to:content-type; bh=YGnVjNNQSxXEhjGwV8/5Y2JvucgSZxjVxoUzS+hSzcQ=; b=XarpU2sYyAgRFSCsGdeL8vvJxaWiV08PDSvFpT52Cflf7dibEr+hbUw4/s0eCZliXv rKjDJHbGWB4sWy2xoOw8rkG282n/sQQicS+Yh7uMHa24KY0Jdn+NsYR/w+L0h8Gc7haQ iyJ8LUvYYXrdaeDvj4XqxOMd7sUyrOt2HQSpqOhGlBliFdhDfzZx+EcPtOKGdO0SPEDc YSfsPmZltJY+OzJ5BDFxy+kjSUEtXs4TApaAuxXd0wZHMHn+e6BGOFug5yz7uyRi4xWY 4juoqJv0VZ/fX1rqQAz7GlgvOEjJuE/e7QW5PC99cF2fZ3I/bY1QTX5VospCBW4V7aAh PEnQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787325788; x=1787930588; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:cc:to:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=YGnVjNNQSxXEhjGwV8/5Y2JvucgSZxjVxoUzS+hSzcQ=; b=YV5bt7p+aAd+VCRtsU20RTvSBN/V5XBOXjABf/3w6Dv2sfV0oDI7F9rhmGF9rn98MM Ol9l+/A9QGgItsr+4g02kJ5XcHuYmxWW1JtGKa7LEn/HxLcBkIbcZJoes1vQo+FvRBVr gWILu6Yy/zH3yFCN5+ZZCSUpds86oFAHY0hVYSy2ONx4AOquYHz4Wg0GgQc+avwlxPkq tNJc254pIVuaVA7JOjhL3QMxlGZU28y3YlaKryQs5zLYGdqFVm+c62Q7DRHmE+voXTtK buzV6fmPwuqwX1Xi3/HMbSuKsrwunMhsrfV5bnkwC+LDzX2e1up6pBNtA8K62VMjdOQ2 tLzw== X-Gm-Message-State: AFuF++kOpSv3dH+u/ZmpPFJXhOXf0OjG/8BvQZznCqbSgC/JueDe2MT0 tPiDcTKi01BYukHiyAgqieQ5oSvXXEMuLtIWASiHmHDhILV00bARadGI X-Gm-Gg: AR+sD13wsLnZ+k7OaiP3XTIE+JarkznqY5VHlnTBZptLtdWsgFFuDZ15OqSeoppiVY0 rEHBgYs/RjzlU4NfofW4dEv3OVL/K7WDUhMFHtLL9SdTCIpjMYGlahkM+6zbY2LuW+JtLucCoZe E9iYg4qJdVV8gpyGyOAEDh2oE/vCf7g1KVrm3tkkVrUoqyFXP7kLrkcjGC4zlGayyqd5i0OTXJ/ F2puiiTeANgCHhwDJvM0jxee4xv8iHYMXK4hvv+zYIH7DYoFd4/FLxfqjZv6ARXKRcp0J68lAaV Y4/MLKzQ3Bjwpq4bsUH92+M6nJwliQh60RYWg+daHbNVeEbBKRbcsOM0M5F8ZDJXmOrovvOv9E5 aTrGLDqawZDR56ipHHEYaH3j53ZP4bl2LVAcU3O+F5PnlIUxdJ/4NN3N1JS0nJX/vNFT1PQjPBU 1Xt/W/vcjBbgPKvoTjsiuJ0avzINOIVQ0+oNwTTMW5OLfCG9EjVTEojFT7tvKGXxaCqz2rSh2CQ fnDIzlvSJx3j5G3WpRJ5DGdomJX1q4qhT0n9Yd/9Q== X-Received: by 2002:a17:907:6d20:b0:c15:9f0c:8d31 with SMTP id a640c23a62f3a-c246a5f67f3mr391862066b.2.1787325787404; Fri, 21 Aug 2026 08:23:07 -0700 (PDT) Received: from [127.0.0.1] (ip-109-193-028-127.um39.pools.vodafone-ip.de. [109.193.28.127]) by smtp.gmail.com with ESMTPSA id a640c23a62f3a-c24591dc662sm506198566b.42.2026.08.21.08.23.06 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 21 Aug 2026 08:23:06 -0700 (PDT) From: Marek Czernohous X-Google-Original-From: Marek Czernohous Date: Fri, 21 Aug 2026 17:23:01 +0200 Subject: [PATCH v4 1/3] drm/nouveau: unsubscribe the channel-kill event before the fence context To: nouveau@lists.freedesktop.org, dri-devel@lists.freedesktop.org Cc: linux-kernel@vger.kernel.org, Lyude Paul , Danilo Krummrich , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , David Airlie , Simona Vetter , Ben Skeggs Message-ID: <178732578167.167481.14789825868710929167@gmail.com> In-Reply-To: <178732578167.167481.5619512544301226563@gmail.com> References: <178732578167.167481.5619512544301226563@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" nouveau_channel_del() tears the fence context down first and only drops the channel-kill subscription later, in the middle of the nvif object teardown: if (chan->fence) nouveau_fence(chan->cli->drm)->context_del(chan); ... nvif_object_dtor(&chan->vram); nvif_event_dtor(&chan->kill); The subscribed handler is nouveau_channel_killed(), which calls nouveau_channel_kill() and from there nouveau_fence_context_kill() on chan->fence. A kill event delivered in that window takes fctx->lock and walks fctx->pending on a fence context that context_del() has already freed. Nothing reaches this below Fermi today, because the subscription is gated on FERMI_CHANNEL_GPFIFO and nothing kills a channel there. On Fermi and newer the window is real but narrow, since a kill has to land exactly while the channel is being destroyed. That is reason enough on its own, which is why this carries a Fixes: tag. The last patch in this series subscribes Tesla channels as well; nothing kills those today, so it does not widen the exposure now, but it is the groundwork for a recovery path that would, and the ordering is better fixed before that lands than alongside it. Drop the subscription before anything it depends on is torn down. Fixes: ea13e5abf807 ("drm/nouveau: signal pending fences when channel has b= een killed") Cc: stable@vger.kernel.org Assisted-by: Claude:claude-opus-5 Signed-off-by: Marek Czernohous Reviewed-by: Lyude Paul --- drivers/gpu/drm/nouveau/nouveau_chan.c | 9 ++++++++- 1 file changed, 8 insertions(+), 1 deletion(-) diff --git a/drivers/gpu/drm/nouveau/nouveau_chan.c b/drivers/gpu/drm/nouve= au/nouveau_chan.c index 598513f60449..f142f6310596 100644 --- a/drivers/gpu/drm/nouveau/nouveau_chan.c +++ b/drivers/gpu/drm/nouveau/nouveau_chan.c @@ -90,6 +90,14 @@ nouveau_channel_del(struct nouveau_channel **pchan) { struct nouveau_channel *chan =3D *pchan; if (chan) { + /* + * Drop the kill-event subscription first. Its handler + * dereferences chan->fence, which the fence context teardown + * below frees, so leaving it armed across the teardown leaves + * a window for a use-after-free. + */ + nvif_event_dtor(&chan->kill); + if (chan->fence) nouveau_fence(chan->cli->drm)->context_del(chan); =20 @@ -100,7 +108,6 @@ nouveau_channel_del(struct nouveau_channel **pchan) nvif_object_dtor(&chan->nvsw); nvif_object_dtor(&chan->gart); nvif_object_dtor(&chan->vram); - nvif_event_dtor(&chan->kill); nvif_object_dtor(&chan->user); nvif_mem_dtor(&chan->mem_userd); nouveau_vma_del(&chan->sema.vma); --=20 2.54.0 From nobody Mon Sep 28 12:33:06 2026 Received: from mail-ej1-f44.google.com (mail-ej1-f44.google.com [209.85.218.44]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EAE644C9564 for ; Fri, 21 Aug 2026 15:23:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.218.44 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787325794; cv=none; b=N1ZuIyFBSTF+/vw9Lyva4Grs8D9iezYFQmZvnP4YX4maLuecRoCuqA3DdjFBpb/X9ce2caaR6nsuH9vddbvbi5Uf7KVOo5B7zVG+nu38tNHlYatXEPXhLZhgh/kVrkAAVJRkoUU+CboboBgv6mQtISM65M9nbvwxhVudBzh1xcI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787325794; c=relaxed/simple; bh=mDUYiCIFTwYuylrBDcfPAZgmy/Xj0GyFwnBLs10N0Wg=; h=From:Date:Subject:To:Cc:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=JV2mH4ObTmbO7CkKl8pxm/cIl2QZacvBluvwZgc5Ixd6+tgNRmIoHT5T3aYRYd0DegqptbhKe4+BjX5BcD7wYsFdz7M64dz/r5bzX75E+x28IyTJFem95txuvfigwkMCaR71CdtuWJhjLkFrN1bT4IkbdBgejlC3Cb2v2gbU22Q= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=a+R9ixyI; arc=none smtp.client-ip=209.85.218.44 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="a+R9ixyI" Received: by mail-ej1-f44.google.com with SMTP id a640c23a62f3a-c1676497000so12243366b.3 for ; Fri, 21 Aug 2026 08:23:12 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787325791; x=1787930591; darn=vger.kernel.org; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:cc:to:subject:date:from:from:to:cc:subject :date:message-id:reply-to:content-type; bh=vYeEFSz/we+8LtDsD3eJWBaG2LH+3C5IKZLmMvYVchk=; b=a+R9ixyIjLFQA7TMBYom+YBcNCouijIWyGVROeCBpBkJF59Qhz5A3b+LxV+G0+G0lK TMIUsF0dt4ui4NTvuScXanFqwMu625t6iR7I0gvTKwTWZ8NwdMhhV1w+QH2L67Af2W5h C02F8Zs6fDHemQb59nCbPewjNqjWP8Yu32TCTICTQ982RCaIAJeo7FjUp7YfmIVYXmN/ tz2Jg+8kh/XWg2kGoskhWEsNVRR8bwIjzNR8o+rAFlWJg4vAaX268Q+ZwjGJ7imfVZzC cYjYguwsH4w2kj32bcRDIBKwaB5ETN3gHkkvlwAxdOE+6LXFOEdZLyylfPE/PlgSl2nN D5xg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787325791; x=1787930591; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:cc:to:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=vYeEFSz/we+8LtDsD3eJWBaG2LH+3C5IKZLmMvYVchk=; b=Ol1MaP7y6taefXXrM5ctBPuJYhdF7ILtwdv/qUnW3K+T4qfewN8EPIT8t1/9fxcxIf uSglLsnuDJaWNAE6roEGl5P07cL0otykENR95LA5ZMmhn49EjiGOeHHUIdQC6XOk/svR x6q35UiyHr1PQ5cukcex7bOoBe/mPtYKpcoFfidHebreP5uS9v2ICeieK1QHYx2/yHe5 qhrUW687f9/rxGi2SBY9pTZ03+OVFMYAZIpysTT9av6C41mc8dtfsqvpz+LzdFdDsjai 5CuyDTiONa6MZBk/e8xLSUHsvnjJTYxausSA4WPZ3EGrZnFP4HQ+pxF0AZZauBhTxgrR 6O4Q== X-Gm-Message-State: AFuF++mEOp9NxYeqWq+7wkIeaE5etRESCWZoaABk24y64lz4xc1RrxW0 +6OxHXs82GUii/UICb27FufhBr/CZ48MYWBtSTg24QDtB0gx3oX17riE X-Gm-Gg: AR+sD11xhkN076aK57R7bkU0JidWRmSr95jfFOwc4IIZdZUmSwVKF98JCVznPMaKg1I Sir3AQTxAXkOCreMj8bwj1hf5qKET6u0tgFDyU8vEIDWEnwdnQ79QywYT64yi5db3atbV0OP0JF fojKPCvgOgAhijGVX3oVSxAeAXo+8AlwxlHA7v2Qc+vFmw622QTOdYFGBkMa1N76HIf/tjc2Zs3 /dkUyIbTeE6e0pX8kPzfuULkYZD1E9V5sJcMp6+eg0SZIGrEXcEDFsq3pfFzNKaGUocPrRfmj/2 Yl/y0IoTm+po1CYNK0fGvO6UO3biqrMAMZX3/JDXgVZc4DmXn6UeHjmnk/FO1q6x4WMEsNqkJ0e 8zBIasMZUdfyOb51hVw05saCS7b1G2l3Qe94mLcwtcqmu8xm7Qkxt7cIysVAcJGWAgPC9/MOM5H TTX5LcKFx8n2Q/VN5kQ3/8acplsqWgtY8d4j7CLue0P8P9KJUSPK4qoQs0wpmqS6fPlbXM5ZndD FVVniHkLMiaDDt25DXXzUX9Mgr//zUynczg0ve8fw== X-Received: by 2002:a17:907:84e:b0:c20:7897:bc67 with SMTP id a640c23a62f3a-c246a6a8e0bmr384314166b.3.1787325791134; Fri, 21 Aug 2026 08:23:11 -0700 (PDT) Received: from [127.0.0.1] (ip-109-193-028-127.um39.pools.vodafone-ip.de. [109.193.28.127]) by smtp.gmail.com with ESMTPSA id a640c23a62f3a-c24591dc662sm506198566b.42.2026.08.21.08.23.09 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 21 Aug 2026 08:23:10 -0700 (PDT) From: Marek Czernohous X-Google-Original-From: Marek Czernohous Date: Fri, 21 Aug 2026 17:23:01 +0200 Subject: [PATCH v4 2/3] drm/nouveau: don't kill a fence context that is not ready yet To: nouveau@lists.freedesktop.org, dri-devel@lists.freedesktop.org Cc: linux-kernel@vger.kernel.org, Lyude Paul , Danilo Krummrich , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , David Airlie , Simona Vetter , Ben Skeggs Message-ID: <178732578167.167481.4590474568281802870@gmail.com> In-Reply-To: <178732578167.167481.5619512544301226563@gmail.com> References: <178732578167.167481.5619512544301226563@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" nouveau_channel_init() arms the channel-kill subscription early, right after mapping userd, and only creates the fence context at the very end of the same function. The handler it installs, nouveau_channel_killed(), reaches nouveau_fence_context_kill(chan->fence). The NULL check in nouveau_channel_kill() does not cover the window in between. Every backend that can reach it publishes the pointer before the context is usable: fctx =3D chan->fence =3D kzalloc_obj(*fctx); if (!fctx) return -ENOMEM; nouveau_fence_context_new(chan, &fctx->base); and nouveau_fence_context_new() is what runs spin_lock_init(&fctx->lock) and INIT_LIST_HEAD(&fctx->pending). An event arriving after the assignment but before that call finds chan->fence non-NULL and unusable: nouveau_fence_context_kill() takes a lock that was never initialised and walks a list head whose next pointer is still the NULL left by kzalloc(). Give the fence context a ->ready flag and hand the kill over through it. nouveau_fence_context_arm() sets the flag once nouveau_channel_init() has finished building the context, and nouveau_channel_kill() leaves the context alone until it is set. A kill arriving while the context is still being built is no longer lost either: it is recorded in chan->killed, and nouveau_fence_context_arm() acts on it as soon as there is a context to kill. The two sides hand over rather than exclude each other, because the kill side must not touch fctx->lock at all before the context is built, which is the very bug being fixed. Each stores its own flag before it loads the other's, so at least one of them observes the other. Both observing it is harmless: nouveau_fence_context_kill() then walks a list the first caller has already emptied. This does not close the other window. A kill delivered before nouveau_channel_init() subscribes is still not observed at all, and nvkm_uchan_init() makes the channel schedulable before that point. Closing that one means subscribing before the channel becomes schedulable, which is a larger change than this fix. The approach is Lyude Paul's suggestion. It is implemented with two differences from the sketch, both following from the same detail. The sketch checks chan->killed before setting ->ready. Both sides have to store their own flag before loading the other's, or the interleaving loses the kill: arm() reads killed =3D=3D 0, kill() sets killed and reads ready =3D=3D false, arm() then sets ready, and neither calls nouveau_fence_context_kill(). That outcome is reachable under sequential consistency, so no barrier can forbid it and the two accesses have to be the other way round in program order. Swapped, and with the smp_mb() on each side, this is the store-buffering pattern of tools/memory-model/litmus-tests/SB+fencembonceonces.litmus. The sketch also holds fctx->lock across the handover. The kill side cannot join it, because reaching fctx->lock is exactly what has to be avoided until the context is built: on those backends chan->fence is published by the allocation, before nouveau_fence_context_new() calls spin_lock_init(). So ->ready is read outside the lock. That answers the open question in the sketch as well: it does not have to be atomic_t, but it does have to be published with release and read with acquire, so that a caller that sees it set also sees the initialised lock and list. Fixes: ea13e5abf807 ("drm/nouveau: signal pending fences when channel has b= een killed") Cc: stable@vger.kernel.org Suggested-by: Lyude Paul Assisted-by: Claude:claude-opus-5 Signed-off-by: Marek Czernohous --- drivers/gpu/drm/nouveau/nouveau_chan.c | 19 ++++++++++++++++--- drivers/gpu/drm/nouveau/nouveau_fence.c | 19 +++++++++++++++++++ drivers/gpu/drm/nouveau/nouveau_fence.h | 8 ++++++++ 3 files changed, 43 insertions(+), 3 deletions(-) diff --git a/drivers/gpu/drm/nouveau/nouveau_chan.c b/drivers/gpu/drm/nouve= au/nouveau_chan.c index f142f6310596..605ce74c0d15 100644 --- a/drivers/gpu/drm/nouveau/nouveau_chan.c +++ b/drivers/gpu/drm/nouveau/nouveau_chan.c @@ -43,9 +43,17 @@ module_param_named(vram_pushbuf, nouveau_vram_pushbuf, i= nt, 0400); void nouveau_channel_kill(struct nouveau_channel *chan) { + struct nouveau_fence_chan *fctx; + atomic_set(&chan->killed, 1); - if (chan->fence) - nouveau_fence_context_kill(chan->fence, -ENODEV); + + /* Pairs with the smp_mb() in nouveau_fence_context_arm(). */ + smp_mb(); + + fctx =3D READ_ONCE(chan->fence); + /* Pairs with the smp_store_release() there. */ + if (fctx && smp_load_acquire(&fctx->ready)) + nouveau_fence_context_kill(fctx, -ENODEV); } =20 static int @@ -494,7 +502,12 @@ nouveau_channel_init(struct nouveau_channel *chan, u32= vram, u32 gart) } =20 /* initialise synchronisation */ - return nouveau_fence(drm)->context_new(chan); + ret =3D nouveau_fence(drm)->context_new(chan); + if (ret) + return ret; + + nouveau_fence_context_arm(chan); + return 0; } =20 int diff --git a/drivers/gpu/drm/nouveau/nouveau_fence.c b/drivers/gpu/drm/nouv= eau/nouveau_fence.c index edbe9e08ba0f..2fed631d44ba 100644 --- a/drivers/gpu/drm/nouveau/nouveau_fence.c +++ b/drivers/gpu/drm/nouveau/nouveau_fence.c @@ -93,6 +93,25 @@ nouveau_fence_context_kill(struct nouveau_fence_chan *fc= tx, int error) spin_unlock_irqrestore(&fctx->lock, flags); } =20 +/* + * Declare a finished fence context killable. A kill can arrive while the + * caller is still building the context, so this and nouveau_channel_kill() + * hand over through fctx->ready and chan->killed. + */ +void +nouveau_fence_context_arm(struct nouveau_channel *chan) +{ + struct nouveau_fence_chan *fctx =3D chan->fence; + + /* Pairs with the smp_load_acquire() in nouveau_channel_kill(). */ + smp_store_release(&fctx->ready, true); + /* Pairs with the smp_mb() there: store-buffering, one side always sees t= he other. */ + smp_mb(); + + if (atomic_read(&chan->killed)) + nouveau_fence_context_kill(fctx, -ENODEV); +} + void nouveau_fence_context_del(struct nouveau_fence_chan *fctx) { diff --git a/drivers/gpu/drm/nouveau/nouveau_fence.h b/drivers/gpu/drm/nouv= eau/nouveau_fence.h index 183dd43ecfff..d9fede5dcba6 100644 --- a/drivers/gpu/drm/nouveau/nouveau_fence.h +++ b/drivers/gpu/drm/nouveau/nouveau_fence.h @@ -53,6 +53,13 @@ struct nouveau_fence_chan { struct work_struct uevent_work; struct nvif_event event; int notify_ref, dead, killed; + + /* + * Set by nouveau_fence_context_arm() once the context is complete. + * Read without fctx->lock, which nouveau_channel_kill() may not + * touch until it is set. + */ + bool ready; }; =20 struct nouveau_fence_priv { @@ -71,6 +78,7 @@ void nouveau_fence_context_new(struct nouveau_channel *, = struct nouveau_fence_ch void nouveau_fence_context_del(struct nouveau_fence_chan *); void nouveau_fence_context_free(struct nouveau_fence_chan *); void nouveau_fence_context_kill(struct nouveau_fence_chan *, int error); +void nouveau_fence_context_arm(struct nouveau_channel *chan); =20 int nv04_fence_create(struct nouveau_drm *); int nv04_fence_mthd(struct nouveau_channel *, u32, u32, u32); --=20 2.54.0 From nobody Mon Sep 28 12:33:06 2026 Received: from mail-ej1-f54.google.com (mail-ej1-f54.google.com [209.85.218.54]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 59FA24D8D80 for ; Fri, 21 Aug 2026 15:23:16 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.218.54 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787325798; cv=none; b=bpVgmH+WOPyxl6KJMgZlhfUGpG1deszK+68FPx0wIzTcbJD+3+UoSTXNtEBeMEVf2dhNx+Hx0d/uyS+5JykJKGfFXzVe1ekXjIo6dpH0qjjHeTLCmp7kFmVMZxot81+3xKccc4gyL62agIG5PKAp5fcPU3HnuexNV5TwKJayHhA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787325798; c=relaxed/simple; bh=lfEv8+TXGNaNlN4Myl3yrEVhPwe7zSuQawloB7wkKZo=; h=From:Date:Subject:To:Cc:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=CnpRAtOrkmzRS7OhhPFAYEuYGQWjKBrLGum1XMOeQeTqGaoE8enGhX3IgMJqkUf2eoEbeH9m90M8TT65GSp+G1DYd58Dh4iVzkG4RN7xajc9xek8i5M/UWOF5JmxKyFObNk3CNK7LpKGluMrXcYFemuwYR2+UIjrd6IoKktVGIc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=Dd8Q+ncQ; arc=none smtp.client-ip=209.85.218.54 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="Dd8Q+ncQ" Received: by mail-ej1-f54.google.com with SMTP id a640c23a62f3a-c160875e029so13301466b.3 for ; Fri, 21 Aug 2026 08:23:16 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787325794; x=1787930594; darn=vger.kernel.org; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:cc:to:subject:date:from:from:to:cc:subject :date:message-id:reply-to:content-type; bh=8sdbNeUe6ctrJEy+VKy4nn9z2qBfoh++mokomXoK9cc=; b=Dd8Q+ncQgppH4+oOpA/dFNPTl3OOaY1lzMj/mMwtCC+4vas1exkJpBiOmbc+Wz2NTu IvaoeZorpma3Oq7ZXesnMpq3cd5gjan/t6/Fi6JJ82z7qcQ3T1mpEvBZRyi1be9ku4Dn Y1AJSHKFmB0qzmZ0cPcjXNYj3ZmTpKri2eZlzrPcpcStXumEDCpbegU3PJ/uMhpa1g4G QA0y9fnNi2UxbgDx4FwTO++7X/lLCJO7R8tt4v1Z+UokEykN+GQtZNY38zW/mq9+xK9m 8e/22oewZDSZ2XVlwaktkQXpGfTM78dYEAO+5/7vDzAt96ZarKJjC1uaxqdt0pECNLB8 e0lg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787325794; x=1787930594; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:cc:to:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=8sdbNeUe6ctrJEy+VKy4nn9z2qBfoh++mokomXoK9cc=; b=OziBroei63lNd8XHfuEptl/1j+79xu47xA/4MUvT5/jJ3Gd2U2r8Qa5VzkV/Z31Y9x LSuKFYbTQZJJxm3mtISl62HbKdUZAcooZM1e8n6X1Y3Mg4yIRfKBQ09QKK1SqT1LTfG1 B1LW1mbL0geO87+TOtdx84zMFAcN3nkLBab4+6gxrQWdydS1CdZ8x5D6BMhvqKIxBBib LI94003tpY2UUsyvg74Tet6WC2QKCiISGWVCJ/tqLYh4BrgvIX8Bo2kdhdeL6jxCMb+z sOfwITz/mPHImZUhIZ0KXeuBWb1h05d0d9joGpAMf2StTogNOieDGSpti5XR2ECI7vKi 7upw== X-Gm-Message-State: AFuF++mtNWNk/h1U4YyF5aZNeW68YTxJfar4Wy/KcTOvNDAtDUnp+H1Z QMYnkP51nTGwCv/127HwB5XwUUoRppPi/TlauUkLbXy8Djl3cFTJg7CE X-Gm-Gg: AR+sD10AexhrsLWxx45sq4OAEWZrepHG14II+zUvTB/vEJHh/R5Ymu74iEy4txCL6YN EvceQuMTEnTubwJ4q4LLxzjdpCtHkrEWJW6jtjXEe+RVBbPuEtNyaDZn1Dgtm98LBIybkeE5mS2 kZGtHY09TmsLgjpEoFqN7lh04er1m4h9Tn/DuY5PrTN23qcUVt6EWznflcXfgeiD/Tk8yB2LEyb vNZSVevzvdDnnmF5GjxFOa7xH+uiKYJQDml3pY5AM+cROxpq4N5vR5e4EqPbPhWxRsNbROJ0YOe oHLw0SuOVOGW3zJ36/GJptO92MsXZ9df3R/2t8UCa4goZg+OqawI7RFwIerR3q0/3sLXglrFArR mMcoVsldAPRAYaMs8PYPm9W6LZ+H9KgQQofuYblDcUF+2VCU6gDD/A8vxsfxNNH0sRp3oMFlMtI RqVsREkpXqXJiYGwCCQl9GMonUeas79ESGJGokBQ3+4e61/AO+GcTtOx22GnyqEe0sHqArNN51t hHCVLPEQWpVcvJdDOXg2OClmc3DMtEbhC2wCHz1OA== X-Received: by 2002:a17:907:cd0e:b0:c15:ccda:26d5 with SMTP id a640c23a62f3a-c246a63d6b3mr388498966b.4.1787325794415; Fri, 21 Aug 2026 08:23:14 -0700 (PDT) Received: from [127.0.0.1] (ip-109-193-028-127.um39.pools.vodafone-ip.de. [109.193.28.127]) by smtp.gmail.com with ESMTPSA id a640c23a62f3a-c24591dc662sm506198566b.42.2026.08.21.08.23.13 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 21 Aug 2026 08:23:13 -0700 (PDT) From: Marek Czernohous X-Google-Original-From: Marek Czernohous Date: Fri, 21 Aug 2026 17:23:01 +0200 Subject: [PATCH v4 3/3] drm/nouveau: subscribe to channel-kill events on NV50 and newer To: nouveau@lists.freedesktop.org, dri-devel@lists.freedesktop.org Cc: linux-kernel@vger.kernel.org, Lyude Paul , Danilo Krummrich , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , David Airlie , Simona Vetter , Ben Skeggs Message-ID: <178732578167.167481.2178147998794444181@gmail.com> In-Reply-To: <178732578167.167481.5619512544301226563@gmail.com> References: <178732578167.167481.5619512544301226563@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" nouveau_channel_init() only subscribes to the channel-killed event for FERMI_CHANNEL_GPFIFO and newer. On NV50/Tesla the subscription therefore never happens, and nvkm_chan_error()'s NVKM_CHAN_EVENT_ERRORED is delivered into an empty notifier list. Today that is harmless, because nothing kills a channel on Tesla: the only nvkm_chan_error() callers are the Fermi and newer recovery paths. So this patch changes no observable behaviour on its own, and that is deliberate: it removes a latent trap before anything can fall into it. I am carrying a Tesla recovery path that does add such a caller and will send it separately once it is ready. Without a subscriber in place the consequences there are severe: nouveau_channel_killed() never runs, so nouveau_fence_context_kill() never runs either, and the pending fences of the killed channel are never signalled. Everything waiting on them waits forever: drm_atomic_helper_wait_for_fences() in the display commit tail waits uninterruptibly and without a timeout, and the TTM delayed delete workers wait in TASK_UNINTERRUPTIBLE. The user sees a frozen desktop on a machine that is otherwise alive, and nothing in the kernel ends that state: both waits pass MAX_SCHEDULE_TIMEOUT, so the fences cannot time out. They are signalled only when the fence context is torn down, that is when the DRM client owning the channel closes its fd and nouveau_fence_context_del() runs. Killing the client, or rebooting, clears it; waiting does not. That is also a dma-fence contract violation: a fence must always be signalled, with an error if necessary. Lower the class gate to NV50_CHANNEL_GPFIFO. The nvkm side is already class neutral: the KILLED case hangs the notifier on runl->chid->event, which every fifo owns since the runlist rework, and nvkm_uchan_uevent() does not discriminate by class. Pre-NV50 chips keep the old behaviour, so NV04 to NV40 are unaffected. Assisted-by: Claude:claude-opus-5 Signed-off-by: Marek Czernohous --- drivers/gpu/drm/nouveau/nouveau_chan.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/drivers/gpu/drm/nouveau/nouveau_chan.c b/drivers/gpu/drm/nouve= au/nouveau_chan.c index 605ce74c0d15..54e2202cb852 100644 --- a/drivers/gpu/drm/nouveau/nouveau_chan.c +++ b/drivers/gpu/drm/nouveau/nouveau_chan.c @@ -378,7 +378,7 @@ nouveau_channel_init(struct nouveau_channel *chan, u32 = vram, u32 gart) if (ret) return ret; =20 - if (chan->user.oclass >=3D FERMI_CHANNEL_GPFIFO) { + if (chan->user.oclass >=3D NV50_CHANNEL_GPFIFO) { DEFINE_RAW_FLEX(struct nvif_event_v0, args, data, sizeof(struct nvif_chan_event_v0)); struct nvif_chan_event_v0 *host =3D --=20 2.54.0