From nobody Sat Jul 25 00:04:17 2026 Received: from mail-pf1-f181.google.com (mail-pf1-f181.google.com [209.85.210.181]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 904F7330D29 for ; Wed, 22 Jul 2026 02:31:28 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.181 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784687491; cv=none; b=uZb/DUWnvtcPHD9iJAH66Ds8daX4SLhVXM2X8FBe2QvwliondD1QQpBJL72Q8+TO3eo7LuXOIp64Ff4HU7CWFL7XenO25i8KGwfkpHFxEeyZpydlb2eZ8g9jgrfpZGplqhuIjI5Qgip+5xatOK32BpoB/pOU0xws/VuQ76Nue88= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784687491; c=relaxed/simple; bh=1PCskL1UZ2A7ti82OI7P8a9Ptxkfv5kMbkZgIW2SCEQ=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=RfVzic/TQfUqfuOOgX1SiCjkrg97c8oAQIUf+QATWhGZQR3II8a6/XXzqkIfU0VE7HsXPXyROZGXSfGx8h0ek+oNwOSrk+SrsycQOWrpU9aoAOjV00dbFKGWWO9bM4L+U24ZqkCdYZukQuE2W6m6oDojgV02JTZuLleUeBJcV3Y= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=f/Qsp4wk; arc=none smtp.client-ip=209.85.210.181 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="f/Qsp4wk" Received: by mail-pf1-f181.google.com with SMTP id d2e1a72fcca58-84830c774a0so6635804b3a.1 for ; Tue, 21 Jul 2026 19:31:28 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784687488; x=1785292288; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to:content-type; bh=J8D1Gp8pfnnkUQW4W1Yf773GP/LJEUPaf09dViFW6dk=; b=f/Qsp4wk7AFUn/ExCOfOt9tOwJsqpHu8FW7mXyONaWZdq8h4ADO0c6xrgj3OsOO3cJ HwwMW/i6ajZiAPB9DtZ4wnoDEfhGEr7M1LNG2HyUP5iIDWYoIPO2zEumZZ+5DVC7vcjG DpzWf1K5gWGrbP0SC/l2GQqvSgclcmIw3MArGw0nkhM0eoVc69nUoynviYjelD/SxQvY gKqI3AFTpLtaLpthws+4a+CWRt8Ij/5hXwlqKSafVcRdkFTuIOsDQ+OhzKMcndyWIYoE l/qeLfccsjHmW2JP1yyLudRQxbtetWRl8r4T1qd1yp8Bj7RaA960adZFNcS8A7w1S0PG kpRA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784687488; x=1785292288; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=J8D1Gp8pfnnkUQW4W1Yf773GP/LJEUPaf09dViFW6dk=; b=IIzUKC1z9wl56TQDOZC9db3Q/2WeRfd4Pt4d8SVQoeawpFBDw8BQUPZOeHZpgEidY5 RDiHz8G5JKJZR8k8dKcrUgnVuyrH3+TB6V4OFLlzc6DBnnX8ozcJR/hScHxbBNmDU4zv k8O2BSFUpVe51P8gSJzbDund1eIC7fQXxDdZYBlqdYPMvFJi6RBTU94xuAciEtAv9NNt jH1UMLmVoUtKKPYd3AhEkZUM8/cV/W+UHNsv+Jp7KZxnvjqmYrOP0kD2TMRuSEEYI/zU rceT7sk79hg4diPEMpwtTWvapvkI9IVfOprQVtKoiFlJIwJNKjNPjHI+ZFC5w9hEr72T PnaQ== X-Forwarded-Encrypted: i=1; AHgh+RqtIttL/9Aj4TA/Q2VySdvOMjZoiiqo0eqkJiPi2czIU6WPUaWTaYf2wON3NQ5PnS6eaIhcm7V/6xhBQNY=@vger.kernel.org X-Gm-Message-State: AOJu0Yzeg38mftf2RgVk/50HReOzVuYBSRy9KbtRJkMj0XWynB6MwgWl H/ZaDmu1KQy/KCpwT/Ut10POoiak7G09N4M1eoSCC7lmnQsfew3Ah9Zi X-Gm-Gg: AR+sD10o4adLwrK/KHC6ykKSCCKs5lY34OJ3J5jDipDUDG34ZSNdO/QAnh2qyLk3gro AOlSo3kifq8fpxA0I5KBzF6f0GlMQm0nOml6CU2W8gh+0VtWhxckE5ymqY2VoXwNWAbOI3WyXhH Yy7oCqXplsX9pw96wtLqC6R8Itxe3jFJX580e06ZWhL3jJu79D9OaS/HgcM+fVvMPYDUt0fLRfl oU1U042pltiA1fZnW4BdqzCyRqmhhkFZqecHrWmNeWgBAWMyiyyajmQEa/ho65WKTlH0vmlbmDE BR6jZoiWOkwXnsAW0X1G57UTBxjOAHNY2i+NG0pr3LIl8zcnmxSY/yGZRoZ+3/a4lR9HMCy3WmM /9q6AjrSLeclhiAg3FMcP0eUCJUOlXhJ5EU6wN14i9XTcA3ubJt0kMNk05ju+D9hz2SiROYc3YB /M7kv9PU9zwoA9GX4Stn/2dMktN9omH0ZjW1c/BNKmetj6kR5HjURCZ1OFWN4JJP4l/g== X-Received: by 2002:a05:6a00:1a8b:b0:847:99bb:b6d0 with SMTP id d2e1a72fcca58-84c29229f58mr20610651b3a.15.1784687487815; Tue, 21 Jul 2026 19:31:27 -0700 (PDT) Received: from secrnd-cstp.tailb7f510.ts.net ([125.131.91.97]) by smtp.gmail.com with ESMTPSA id 41be03b00d2f7-cbb8f0ee1e6sm356253a12.12.2026.07.21.19.31.24 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 21 Jul 2026 19:31:27 -0700 (PDT) From: Sanghyun Park To: Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , Eduard Zingerman Cc: Sanghyun Park , Sun Jian , Kumar Kartikeya Dwivedi , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Ihor Solodrai , Puranjay Mohan , bpf@vger.kernel.org, linux-kernel@vger.kernel.org, syzbot+cdd6c0925e12b0af60cc@syzkaller.appspotmail.com, sashiko-bot@kernel.org Subject: [PATCH bpf-next] bpf: Fix mmap_lock leak in irq_work path Date: Wed, 22 Jul 2026 11:30:03 +0900 Message-ID: <20260722023004.1497923-2-sanghyun.park.cnu@gmail.com> X-Mailer: git-send-email 2.48.1 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" stack_map_get_build_id_offset() introduced a per-CPU irq_work to defer mmap_read_unlock() from NMI context, and bpf_find_vma() later reused the same mmap_unlock_work. Both callers only check whether the work is busy before taking mmap_lock, so a nested caller can reuse the slot before the first caller queues it. Two read locks may then be acquired while only one deferred unlock runs, leaking a read lock and blocking exit_mmap(). Reserve the per-CPU slot before mmap_read_trylock(). Use the same wrapper in stackmap and bpf_find_vma() so both callers release the reservation on trylock failure. Release it after the irq_work callback unlocks the mm. Fixes: bae77c5eb5b2 ("bpf: enable stackmap with build_id in nmi context") Reported-by: syzbot+cdd6c0925e12b0af60cc@syzkaller.appspotmail.com Closes: https://syzkaller.appspot.com/bug?extid=3Dcdd6c0925e12b0af60cc Reported-by: sashiko-bot@kernel.org Closes: https://lore.kernel.org/r/20260630033745.B80201F000E9@smtp.kernel.o= rg Signed-off-by: Sanghyun Park Tested-by: Sun Jian --- kernel/bpf/mmap_unlock_work.h | 37 +++++++++++++++++++++++++++++++---- kernel/bpf/stackmap.c | 3 +-- kernel/bpf/task_iter.c | 7 +++---- 3 files changed, 37 insertions(+), 10 deletions(-) diff --git a/kernel/bpf/mmap_unlock_work.h b/kernel/bpf/mmap_unlock_work.h index 5d18d7d85bef9..d416e4337635f 100644 --- a/kernel/bpf/mmap_unlock_work.h +++ b/kernel/bpf/mmap_unlock_work.h @@ -4,12 +4,14 @@ =20 #ifndef __MMAP_UNLOCK_WORK_H__ #define __MMAP_UNLOCK_WORK_H__ +#include #include =20 /* irq_work to run mmap_read_unlock() in irq_work */ struct mmap_unlock_irq_work { struct irq_work irq_work; struct mm_struct *mm; + atomic_t active; }; =20 DECLARE_PER_CPU(struct mmap_unlock_irq_work, mmap_unlock_work); @@ -18,8 +20,8 @@ DECLARE_PER_CPU(struct mmap_unlock_irq_work, mmap_unlock_= work); * We cannot do mmap_read_unlock() when the irq is disabled, because of * risk to deadlock with rq_lock. To look up vma when the irqs are * disabled, we need to run mmap_read_unlock() in irq_work. We use a - * percpu variable to do the irq_work. If the irq_work is already used - * by another lookup, we fall over. + * percpu variable to do the irq_work. The active flag reserves the slot + * before mmap_read_trylock() and until the irq_work callback consumes mm. */ static inline bool bpf_mmap_unlock_get_irq_work(struct mmap_unlock_irq_wor= k **work_ptr) { @@ -29,9 +31,10 @@ static inline bool bpf_mmap_unlock_get_irq_work(struct m= map_unlock_irq_work **wo if (irqs_disabled()) { if (!IS_ENABLED(CONFIG_PREEMPT_RT)) { work =3D this_cpu_ptr(&mmap_unlock_work); - if (irq_work_is_busy(&work->irq_work)) { - /* cannot queue more up_read, fallback */ + if (irq_work_is_busy(&work->irq_work) || + atomic_cmpxchg_acquire(&work->active, 0, 1)) { irq_work_busy =3D true; + work =3D NULL; } } else { /* @@ -46,6 +49,32 @@ static inline bool bpf_mmap_unlock_get_irq_work(struct m= map_unlock_irq_work **wo return irq_work_busy; } =20 +static inline void bpf_mmap_unlock_put_irq_work(struct mmap_unlock_irq_wor= k *work) +{ + if (work) + atomic_set_release(&work->active, 0); +} + +/* + * Try to take mm->mmap_lock for reading on behalf of a BPF helper that may + * run with IRQs disabled. On success, *work is the slot to hand to + * bpf_mmap_unlock_mm() (NULL when the unlock can be done inline); on fail= ure + * no slot stays reserved and the caller must fall back. + */ +static inline bool bpf_mmap_read_trylock(struct mm_struct *mm, + struct mmap_unlock_irq_work **work) +{ + if (bpf_mmap_unlock_get_irq_work(work)) + return false; + + if (!mmap_read_trylock(mm)) { + bpf_mmap_unlock_put_irq_work(*work); + return false; + } + + return true; +} + static inline void bpf_mmap_unlock_mm(struct mmap_unlock_irq_work *work, s= truct mm_struct *mm) { if (!work) { diff --git a/kernel/bpf/stackmap.c b/kernel/bpf/stackmap.c index 41fe87d7302f2..166fc8efab8b2 100644 --- a/kernel/bpf/stackmap.c +++ b/kernel/bpf/stackmap.c @@ -415,7 +415,6 @@ static void stack_map_get_build_id_offset(struct bpf_st= ack_build_id *id_offs, u32 trace_nr, bool user, bool may_fault) { struct mmap_unlock_irq_work *work =3D NULL; - bool irq_work_busy =3D bpf_mmap_unlock_get_irq_work(&work); bool has_user_ctx =3D user && current && current->mm; struct stack_map_build_id_cache cache =3D {}; struct vm_area_struct *vma; @@ -430,7 +429,7 @@ static void stack_map_get_build_id_offset(struct bpf_st= ack_build_id *id_offs, * fallback is used for kernel stack (!user) on a stackmap with * build_id. */ - if (!has_user_ctx || irq_work_busy || !mmap_read_trylock(current->mm)) { + if (!has_user_ctx || !bpf_mmap_read_trylock(current->mm, &work)) { /* cannot access current->mm, fall back to ips */ for (i =3D 0; i < trace_nr; i++) stack_map_build_id_set_ip(&id_offs[i]); diff --git a/kernel/bpf/task_iter.c b/kernel/bpf/task_iter.c index b256fb9c1214e..e66e3484f923d 100644 --- a/kernel/bpf/task_iter.c +++ b/kernel/bpf/task_iter.c @@ -755,7 +755,6 @@ BPF_CALL_5(bpf_find_vma, struct task_struct *, task, u6= 4, start, { struct mmap_unlock_irq_work *work =3D NULL; struct vm_area_struct *vma; - bool irq_work_busy =3D false; bool __maybe_unused mmput_needed =3D false; struct mm_struct *mm; int ret =3D -ENOENT; @@ -792,9 +791,7 @@ BPF_CALL_5(bpf_find_vma, struct task_struct *, task, u6= 4, start, if (!mm) return -ENOENT; =20 - irq_work_busy =3D bpf_mmap_unlock_get_irq_work(&work); - - if (irq_work_busy || !mmap_read_trylock(mm)) { + if (!bpf_mmap_read_trylock(mm, &work)) { ret =3D -EBUSY; goto out; } @@ -1191,6 +1188,8 @@ static void do_mmap_read_unlock(struct irq_work *entr= y) =20 work =3D container_of(entry, struct mmap_unlock_irq_work, irq_work); mmap_read_unlock_non_owner(work->mm); + work->mm =3D NULL; + bpf_mmap_unlock_put_irq_work(work); } =20 static int __init task_iter_init(void) --=20 2.48.1