From nobody Mon Aug 24 04:17:54 2026 Received: from mail-wm1-f51.google.com (mail-wm1-f51.google.com [209.85.128.51]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7D7E627442 for ; Sun, 16 Aug 2026 12:58:56 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.51 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786885138; cv=none; b=RwavB66Wg/bpZNAueZc6zL5OyoThkY5Lf+i2K9dZ9vvNsJ6AMlk85WaayFTTT05za1GaFPCnZz6GysKaDL9AX4CYgs+Tot3D99EdXS+uTAzkgW2dg5+onxtDsVPzwI126ZEnxrvoO3L6jXH/VZPBhVgm/hltMFtVezRSAHULgA4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786885138; c=relaxed/simple; bh=kANdLgyx4ZUR2AY9MegqyOJ2Mm8STmGDf7bbKaSEPYo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: Content-Type:MIME-Version; b=OmBZyOFvbbjDYqf/1w13ropHqGus7WaCCVjRHBgUDKULKe6N2XCvO/Z7I7c3xoP7Xqe0HkJQtclgMozDYcPJwWUQNhH4Oh/SuXi+d4zj5dBosK0qn4Pc7bU3gWYoPqMDGEIcq0ZJnOPXyAKKNTSb0ER37/MXjdNYZAkbyQxTfrg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=kRrpkYw6; arc=none smtp.client-ip=209.85.128.51 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="kRrpkYw6" Received: by mail-wm1-f51.google.com with SMTP id 5b1f17b1804b1-4957952e0f8so2555345e9.2 for ; Sun, 16 Aug 2026 05:58:56 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786885135; x=1787489935; darn=vger.kernel.org; h=mime-version:content-transfer-encoding:content-type:references :in-reply-to:message-id:date:subject:cc:to:from:from:to:cc:subject :date:message-id:reply-to:content-type; bh=jTmygf1/JyUF8FN6We5SdcT/ctQ4+SGdYMt719yqvrw=; b=kRrpkYw6ZufDxZvoeZdBy/9Wr7plLwkTStwNeQMhzpT2cWFaToiA+fJbYcdqLK4G1W Wt1pvizzc0NNDUSKERd4H9CJsH1QU8Qfz7p7pYyu1jeqF48aLDnJjAznj3Y0dyCQ9ezS B/7T20znzR6Bt16TSz3b0CvETr7uwRn1cYoOSus/IfXmExDxJ3i/7kzw7P10YiSxd5zO x9mAo1WOmPp74BXqIVTiqSBKCTyqVOHz5P+L8H0TH2Fk4x1m7V90WJo7VpflUbDoUixX oG0nv2xmMt3I9XYHLYH9fur0Jz+Mofudu8yH6Q76L4HT9sOLRFv8BmgfqHEnA4wGKGGC ePBw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786885135; x=1787489935; h=mime-version:content-transfer-encoding:content-type:references :in-reply-to:message-id:date:subject:cc:to:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=jTmygf1/JyUF8FN6We5SdcT/ctQ4+SGdYMt719yqvrw=; b=Q+qmhNH2ZZkMmhmMpIQWOJDRJYOg0Mk+w19JW2yl4XY/66saKNcXKktrY8YfZf2aUq +fB/P6bkvLVN8LaS+FN449syBMJ+VwLHj0V/v1mtIee781eLybnNDYGtwPD4woPD6BXo sSS1y2O9mLCkt7FXZFoq/dsIEC2fVZT/LyUJtbqPGhbdVYw4wO3oHstr7XdjEGmYkY5A Bn6K3QA9iVY1P13xlBJEyhT34YUjTS0tjCciEgxd+qRx0anMvM2IYxaE9/YW6/UHar3/ 981+F0INphUoZHFi18VfT79T80o30Rw5OB5VW02P9cN4GzrCoPCJDLrplOepTwEoO35m x43w== X-Gm-Message-State: AOJu0YxFY1XsJygrEbABIJnp+0Bbz4YahSmwSzw/o19d6KcKtXkoAh8i v12WZt6PsTXd1jyR9a3pCESyvz9mU3zKp1Ieu94yCF9qb43uuyTEors+ X-Gm-Gg: AR+sD10jdU9PGSTZaIRkKaTwSEwrNZkFSdn0+auhch9F5hlK1q5af8jVe8ki+FPxZkl AkEmihO7UlTWnCBAARF776mS6xe6xWnq6C2+QFpH7nAmVqpcGyYWyRSFU0DVOhnq139GTRP4wBl hZuL1pZ1QlcbaVt7+jbSwHSv3Xg1KXvdIufhu1uthG1jrovhB2p5K0y3N2jClEDqwHPgTWRZ7EF 91K/4G7md4OE7SrN9nd2g5Xg0yZSnib+tZpIg8lJZ2weNG3U4XAxOsFYIsKLnmfZk8iMEEKbwHH O/yvCvdK/Llm3m74iPDKGs5mM4Rxjx98tPTKhT9U9Qj4VYysf1U7fk5zkZLk4x8PmY66bl/3p1n bufgrixlMqOnZl6S05wjGZsT1iII5J+RVUYsahA/FZlK8Mp0sigtOPJClDqZL5VZ9WqfDrlGjx5 XBO0pstROHJH5SjGe479VIqYM+ILDhAn+FJWjXvfgUEhPnWQnjKJDE25HnmryeHVoYRrgDYs5yN 3xQ2oYps4Cg68cNdPuBZ8TFg5o1BqJrV+4cHODO8+4= X-Received: by 2002:a05:600c:358d:b0:493:bea7:6b67 with SMTP id 5b1f17b1804b1-49987960f4bmr160582535e9.3.1786885134413; Sun, 16 Aug 2026 05:58:54 -0700 (PDT) Received: from [127.0.0.1] (ip-109-193-028-127.um39.pools.vodafone-ip.de. [109.193.28.127]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-49996188217sm108278675e9.13.2026.08.16.05.58.53 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sun, 16 Aug 2026 05:58:53 -0700 (PDT) From: Marek Czernohous To: nouveau@lists.freedesktop.org, dri-devel@lists.freedesktop.org Cc: linux-kernel@vger.kernel.org, Danilo Krummrich , Lyude Paul , David Airlie , Simona Vetter Subject: [PATCH v2] drm/nouveau: disable the fence uevent work instead of just cancelling it Date: Sun, 16 Aug 2026 14:58:52 +0200 Message-ID: <178688513268.513871.3468844663561639695@gmail.com> In-Reply-To: <178682366002.3748010.12779628082366287968@gmail.com> References: <178682366001.3748010.7798811159846779765@gmail.com> <178682366002.3748010.12779628082366287968@gmail.com> X-Mailer: python-smtplib Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 From: Marek Czernohous nouveau_fence_context_del() drains the uevent work while the event that feeds it is still armed: cancel_work_sync(&fctx->uevent_work); nouveau_fence_context_kill(fctx, 0); nvif_event_dtor(&fctx->event); nouveau_fence_wait_uevent_handler() queues the work unconditionally: schedule_work(&fctx->uevent_work); return NVIF_EVENT_KEEP; so a non-stall interrupt arriving after cancel_work_sync() has returned re-arms the work that was just drained. The window closes in two steps, neither of which is the drain. The kill blocks the event when the last fence holding a notify_ref is signalled, which reaches atomic_xchg(&ntfy->allowed, 0) (nvkm/core/event.c:104), and nvkm_event_ntfy() skips a ntfy that is not allowed (:183). That stops further handlers from starting, but not one that is already inside nvkm_event_ntfy(): the event is created with wait =3D false (nouveau_fence.c:201), so nvkm_event_ntfy_block_() leaves it on the list and never takes event->list_lock. Only nvif_event_dtor() waits that one out: nvkm_event_ntfy_del() (:141) goes through nvkm_event_ntfy_remove(), which takes write_lock_irq() on that same list_lock (:84). Either way the re-arm happens after the drain, and the caller drops its reference immediately afterwards, for example nv84_fence_context_del(): nouveau_fence_context_del(&fctx->base); chan->fence =3D NULL; nouveau_fence_context_free(&fctx->base); That is a kref_put() on fctx->fence_ref, so the context outlives the teardown only while emitted fences still hold a reference of their own. That is no safety net: whenever none do, the count reaches zero right there and nouveau_fence_context_put() kfree()s fctx while the work is still queued. &fctx->uevent_work is embedded in that allocation, so the workqueue already dereferences freed memory when it picks the item up, and nouveau_fence_uevent_work() can then take fctx->lock on it. With CONFIG_DEBUG_OBJECTS_WORK and CONFIG_DEBUG_OBJECTS_FREE, kfree() of a still-queued work item is reported as the free of an active object. On live memory the re-armed work has nothing left to do: nouveau_fence_context_kill() empties fctx->pending and sets fctx->killed under fctx->lock, nouveau_fence_emit() then returns -ENODEV rather than queueing anything new, and nouveau_fence_update() only reaches nvif_event_block() if it signalled something off that list. The defect is the access to freed memory, not what the work would have found. Only chips from G84 on can reach this at all: nouveau_fence_context_new() returns before nvif_event_ctor() when priv->uevent is clear, and nv84_fence_create() is the only place that sets it. nv84_fence_context_del() is the context_del for all of those, because nvc0_fence_create() and gv100_fence_create() build on nv84_fence_create() and override only context_new. Use disable_work_sync() instead. It drains the work exactly like cancel_work_sync() does, and additionally increments the work item's disable count, after which "any attempt to queue @work will fail and return %false" (kernel/workqueue.c, disable_work()). The handler's schedule_work() then has nothing to re-arm, and the teardown order stays as it is. Draining a second time after nvif_event_dtor() would close the window as well, and without the newer API: once nvkm_event_ntfy_remove() has returned, no handler can start or still be running, so nothing re-arms the work past that point. disable_work_sync() is preferred here because it needs one synchronisation point instead of two, it keeps the work from being queued at all rather than cleaning up after it, and it is what drm has settled on for this (drm/xe, drm/panthor, drm_pagemap). Blocking the event rather than the work is not an option: fctx->event is created with wait =3D false, so a handler already inside nvkm_event_ntfy() can still queue the work. Note for backports: disable_work_sync() arrived in v6.10 with commit 86898fa6b8cd ("workqueue: Implement disable/enable for (delayed) work items"), while the fix being corrected here reached 6.6.18 and 6.7.6. linux-6.6.y therefore carries this bug without the API, and this patch would apply there and then fail to build. A 6.6.y backport wants the second drain described above instead, as its own patch. Reported-by: sashiko-bot Closes: https://lore.kernel.org/nouveau/20260812231330.705425-1-mczernohous= @gmail.com/ Fixes: 39126abc5e20 ("nouveau: offload fence uevents work to workqueue") Cc: # 6.10.x Assisted-by: Claude:claude-opus-5 Signed-off-by: Marek Czernohous Reviewed-by: Lyude Paul --- Changes in v2: - Do not reorder the teardown. v1 moved nvif_event_dtor() ahead of nouveau_fence_context_kill(); with the fences still unsignalled that leaves nouveau_fence_enable_signaling() reachable, and nvif_event_constructed() is a plain unlocked read of object->client, so the dtor could race an nvif_event_allow() already past that check. Reported as [Critical] by the bot, and withdrawn: https://lore.kernel.org/all/20260815200914.8A1131F000E9@smtp.kernel.org/ - Change cancel_work_sync() to disable_work_sync() instead, which leaves every ordering alone. - Pin the damage down. Both versions call it a use-after-free; this one adds that &fctx->uevent_work is embedded in the freed allocation, and that the re-armed work has nothing left to do on a live context, so the access to freed memory is the whole of it. - Correct the backport note. v1 claimed no longterm tree sat in the gap between the bug and disable_work_sync(); 6.6.y does. The stable tag is annotated accordingly. This replaces 1/3 of https://lore.kernel.org/all/178682366002.3748010.12779628082366287968@gmail= .com/ 2/3 and 3/3 of that series are unaffected and still stand. drivers/gpu/drm/nouveau/nouveau_fence.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/drivers/gpu/drm/nouveau/nouveau_fence.c b/drivers/gpu/drm/nouv= eau/nouveau_fence.c index edbe9e08ba0f..11e95c37ce50 100644 --- a/drivers/gpu/drm/nouveau/nouveau_fence.c +++ b/drivers/gpu/drm/nouveau/nouveau_fence.c @@ -96,7 +96,7 @@ nouveau_fence_context_kill(struct nouveau_fence_chan *fct= x, int error) void nouveau_fence_context_del(struct nouveau_fence_chan *fctx) { - cancel_work_sync(&fctx->uevent_work); + disable_work_sync(&fctx->uevent_work); nouveau_fence_context_kill(fctx, 0); nvif_event_dtor(&fctx->event); fctx->dead =3D 1; base-commit: c21bb4193868a8de71fc4693fa741e195fdf5d86 --=20 2.54.0