From nobody Tue Sep 29 13:20:32 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AA0AE30D3E9; Fri, 7 Aug 2026 13:50:47 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786110657; cv=none; b=KpgUhWX8KofHrXDwsH6qBT3MzUZbI16YYz6icK1Rf/bCRAOBzQUgDxEvXqH5eqnbJxMeyRqKReCEFrWfY6/H9kBZHDU8G7TQaoaIZNL++BsW/DtGYj/9v0bpBFFc/r4cLuJVtvjlFNt+8CAW5xMpmUgRgWYO186rGs4fT0Ashss= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786110657; c=relaxed/simple; bh=ky9oI+Rh3DA1pLSbyL4Qrvo5EzgRjfvlbv7xCHQoY/g=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=OUfQ3I2xYvhns6MO/xZuX9+go34F7+AFoo2YZs+kKQBGhj7gBGBatNxREwpqa3M/AR+ac7BRSrSMl6a2qtsgPI37SG1qHUi3alt88oRgNhnxzTRfLY9v2wTMdlUzoKvZaET7jaUDEoFh1o+uBB4dgzY0YMhrThGRw7FYol4lpbs= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=MRqvRaU0; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="MRqvRaU0" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 1F9781F00A3E; Fri, 7 Aug 2026 13:50:36 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786110643; bh=EwPefllU6d1ZOhpuz2HmPQx5BKCFxUJHa6Nznn6LyRE=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=MRqvRaU0RAWLPqYnu2hVF4J5A/YLaMcCc9zGyL2HSxIqa/Mj0SL0oKlKbwt3k85om /K+mY2Imhr15J0MQpohqmZx3hyI5PEo905yB/Iz/8GW23+RGf0MgL09nPRi/ZX4aCQ Trxl3o2SIdim/srZsS61ZsLQcc67m5Oe8hHOAwiSkt02Ou6uM8FsSASn+yez70ZATR N6FQ3VtTRJea/r6sSGDDFMCUJz76Mt6lR4H0cJeQq3OMQkbhrzJLLKULVBoysbLf+c 9kP9g90IN29MxfINOrWOUfpOVoqKYxH2oqKURDtpC+GxAE3HGVlr6SDAen7o2JOfdv 8v2yKob6okQxw== From: "Vlastimil Babka (SUSE)" Date: Fri, 07 Aug 2026 15:50:27 +0200 Subject: [PATCH RFC 1/5] mm/slab: cleanup deferred free handling Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260807-kfree_nolock_kmalloc-v1-1-ba993cbf7a60@kernel.org> References: <20260807-kfree_nolock_kmalloc-v1-0-ba993cbf7a60@kernel.org> In-Reply-To: <20260807-kfree_nolock_kmalloc-v1-0-ba993cbf7a60@kernel.org> To: Harry Yoo , Alexei Starovoitov , Alexander Potapenko , Marco Elver , Sumit Semwal , =?utf-8?q?Christian_K=C3=B6nig?= , Catalin Marinas , Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Sebastian Andrzej Siewior , Clark Williams , Steven Rostedt Cc: Andrew Morton , Hao Li , Christoph Lameter , David Rientjes , Roman Gushchin , bpf@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, Dmitry Vyukov , linux-media@vger.kernel.org, dri-devel@lists.freedesktop.org, linaro-mm-sig@lists.linaro.org, kasan-dev@googlegroups.com, Dietmar Eggemann , Ben Segall , Mel Gorman , Valentin Schneider , K Prateek Nayak , linux-rt-devel@lists.linux.dev, "Vlastimil Babka (SUSE)" X-Mailer: b4 0.15.2 In deferred_percpu_work_fn() we have a bunch of single-use local variables for the various llists. Remove them and access the lists directly. In defer_free() make it more obvious and documented what we are doing. Also restrict guard(preempt) to only the necessary part. Signed-off-by: Vlastimil Babka (SUSE) Reviewed-by: Hao Li --- mm/slub.c | 24 +++++++++++++----------- 1 file changed, 13 insertions(+), 11 deletions(-) diff --git a/mm/slub.c b/mm/slub.c index b9aeb02a880f..044db93d64a0 100644 --- a/mm/slub.c +++ b/mm/slub.c @@ -6371,16 +6371,12 @@ static void free_to_pcs_bulk(struct kmem_cache *s, = size_t size, void **p) static void deferred_percpu_work_fn(struct irq_work *work) { struct deferred_percpu_work *dpw; - struct llist_head *objs, *objs_by_rcu, *rcu_sheaves; struct llist_node *llnode, *pos, *t; struct slab_sheaf *sheaf, *next; =20 dpw =3D container_of(work, struct deferred_percpu_work, work); - rcu_sheaves =3D &dpw->rcu_sheaves; - objs =3D &dpw->objects; - objs_by_rcu =3D &dpw->objects_by_rcu; =20 - llnode =3D llist_del_all(objs); + llnode =3D llist_del_all(&dpw->objects); llist_for_each_safe(pos, t, llnode) { struct kmem_cache *s; struct slab *slab; @@ -6403,7 +6399,7 @@ static void deferred_percpu_work_fn(struct irq_work *= work) stat(s, FREE_SLOWPATH); } =20 - llnode =3D llist_del_all(objs_by_rcu); + llnode =3D llist_del_all(&dpw->objects_by_rcu); llist_for_each_safe(pos, t, llnode) { void *head =3D pos; void *objp =3D kvmalloc_obj_start_addr(head); @@ -6411,21 +6407,27 @@ static void deferred_percpu_work_fn(struct irq_work= *work) kvfree_call_rcu(head, objp); } =20 - llnode =3D llist_del_all(rcu_sheaves); + llnode =3D llist_del_all(&dpw->rcu_sheaves); llist_for_each_entry_safe(sheaf, next, llnode, llnode) call_rcu(&sheaf->rcu_head, rcu_free_sheaf); } =20 -static void defer_free(struct kmem_cache *s, void *head) +static void defer_free(struct kmem_cache *s, void *obj) { struct deferred_percpu_work *dpw; + struct llist_node *llnode; =20 - guard(preempt)(); + /* + * Place the llist node where the freepointer would be if we freed the + * object immediately. That means we can write there safely, only need + * to remove kasan tag first. + */ + llnode =3D kasan_reset_tag(obj) + s->offset; =20 - head =3D kasan_reset_tag(head); + guard(preempt)(); =20 dpw =3D this_cpu_ptr(&deferred_percpu_work); - if (llist_add(head + s->offset, &dpw->objects)) + if (llist_add(llnode, &dpw->objects)) irq_work_queue(&dpw->work); } =20 --=20 2.55.0 From nobody Tue Sep 29 13:20:32 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 450962BE03C; Fri, 7 Aug 2026 13:50:51 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786110660; cv=none; b=glt1XvvzAoPcKtwqmvZmGQQiQZKlVN8g7MjuXo1fsxVYYC1HMtQ6suRCstQtfcbqCnzH20BhdtPdchvegG0WZLNMLMJVj5L8pFSL48D+1d0OBnTLl7c0DrV5XoHPrw+4K3y/fND7bBsVSiWqQJ1X7vpfwoAYfdiR/GHZ6ZUYT94= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786110660; c=relaxed/simple; bh=JkbYZ6MqoMwwegvFna5plnwe/Q32yMBbRptc6AqLJ1o=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=P/ewVrd2tztw+FVzz4xmTQ83pYA8kUC/cFDhB81h1OTFHIMDbWBZctFcfWjjdcQGiFvqjQIBvx+VtevwJ3nWoNsTXAI+ye0RvYulEMOkySQJzb8toXo8J3Qt0XF1pQ1rbQZCBtVNzqKG88NO+9dOEpRpAvSQ8Ue7xEtxyWAZiUs= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=oroht8SH; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="oroht8SH" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 29AEC1F01558; Fri, 7 Aug 2026 13:50:43 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786110650; bh=N6heLdTrKkhpyc80UCW2aSDkkttYCK0SQ8+JM1HNgJk=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=oroht8SHBVezTmTAIeRMFFVdtkT5L7PiN82AxE59EDAH+uMGTOVeqysT+Us04uBTO 8uYl+S9Bn9qEz7CfN6vbRvdXj4UIOdwo+2EPIacVW6JpTKVpLN1c+44qGIqSS2iv5H sHF+5Wipg4mVOrsmWLSiXnyqhhBYYOIFmyciRVEuA/itJ/q/qY+l2O1uSeql9PgrdQ UIzGsrnmyS8y/g0Epf8oJI1eA8jfaTqY6+a1qRaHkwvwkZw7XcarCvrA1IJloLxqHv o+eVhWqd/GufTn8hwSVSjYG6AbyRY7qQBiZRi/NfxlK6xlq7OsZFWUiFV8p9tTMpVf u4FXzDxZqFE3A== From: "Vlastimil Babka (SUSE)" Date: Fri, 07 Aug 2026 15:50:28 +0200 Subject: [PATCH RFC 2/5] mm/slab, kfence: support kfence objects in kfree_nolock() Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260807-kfree_nolock_kmalloc-v1-2-ba993cbf7a60@kernel.org> References: <20260807-kfree_nolock_kmalloc-v1-0-ba993cbf7a60@kernel.org> In-Reply-To: <20260807-kfree_nolock_kmalloc-v1-0-ba993cbf7a60@kernel.org> To: Harry Yoo , Alexei Starovoitov , Alexander Potapenko , Marco Elver , Sumit Semwal , =?utf-8?q?Christian_K=C3=B6nig?= , Catalin Marinas , Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Sebastian Andrzej Siewior , Clark Williams , Steven Rostedt Cc: Andrew Morton , Hao Li , Christoph Lameter , David Rientjes , Roman Gushchin , bpf@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, Dmitry Vyukov , linux-media@vger.kernel.org, dri-devel@lists.freedesktop.org, linaro-mm-sig@lists.linaro.org, kasan-dev@googlegroups.com, Dietmar Eggemann , Ben Segall , Mel Gorman , Valentin Schneider , K Prateek Nayak , linux-rt-devel@lists.linux.dev, "Vlastimil Babka (SUSE)" X-Mailer: b4 0.15.2 KFENCE objects are one of the reasons why kfree_nolock() cannot currently handle kmalloc() objects. They are however rare so we can simply defer their freeing to irq_work. The only complication is where to put the llist node. We cannot use the freepointer location like in defer_free() because for some caches it may be outside the object area and KFENCE would detect writes there. Since KFENCE already solves a similar situation when freeing objects from SLAB_TYPESAFE_BY_RCU caches with an rcu_head in its internal metadata, reuse that rcu_head also for the llist node. Introduce kfence_obj_to_llnode() and kfence_llnode_to_obj() so SLAB can work with this llist node without being exposed to KFENCE internals. Signed-off-by: Vlastimil Babka (SUSE) --- include/linux/kfence.h | 5 +++++ mm/kfence/core.c | 14 ++++++++++++++ mm/kfence/kfence.h | 5 ++++- mm/slub.c | 39 ++++++++++++++++++++++++++++++++++++--- 4 files changed, 59 insertions(+), 4 deletions(-) diff --git a/include/linux/kfence.h b/include/linux/kfence.h index e5822f6e7f27..00721c85258d 100644 --- a/include/linux/kfence.h +++ b/include/linux/kfence.h @@ -188,6 +188,9 @@ static __always_inline __must_check bool kfence_free(vo= id *addr) return true; } =20 +struct llist_node *kfence_obj_to_llnode(void *addr); +void *kfence_llnode_to_obj(struct llist_node *llnode); + /** * kfence_handle_page_fault() - perform page fault handling for KFENCE pag= es * @addr: faulting address @@ -235,6 +238,8 @@ static inline size_t kfence_ksize(const void *addr) { r= eturn 0; } static inline void *kfence_object_start(const void *addr) { return NULL; } static inline void __kfence_free(void *addr) { } static inline bool __must_check kfence_free(void *addr) { return false; } +static inline struct llist_node *kfence_obj_to_llnode(void *addr) { return= NULL; } +static inline void *kfence_llnode_to_obj(struct llist_node *llnode) { retu= rn NULL; } static inline bool __must_check kfence_handle_page_fault(unsigned long add= r, bool is_write, struct pt_regs *regs) { diff --git a/mm/kfence/core.c b/mm/kfence/core.c index 6577bd76954e..42519d24687f 100644 --- a/mm/kfence/core.c +++ b/mm/kfence/core.c @@ -1271,6 +1271,20 @@ void __kfence_free(void *addr) } } =20 +struct llist_node *kfence_obj_to_llnode(void *addr) +{ + struct kfence_metadata *meta =3D addr_to_metadata((unsigned long)addr); + + return &meta->llnode; +} + +void *kfence_llnode_to_obj(struct llist_node *llnode) +{ + struct kfence_metadata *meta =3D container_of(llnode, struct kfence_metad= ata, llnode); + + return (void *)meta->addr; +} + bool kfence_handle_page_fault(unsigned long addr, bool is_write, struct pt= _regs *regs) { const int page_index =3D (addr - (unsigned long)__kfence_pool) / PAGE_SIZ= E; diff --git a/mm/kfence/kfence.h b/mm/kfence/kfence.h index 1f618f9b0d12..0fca1dc2c794 100644 --- a/mm/kfence/kfence.h +++ b/mm/kfence/kfence.h @@ -58,7 +58,10 @@ struct kfence_track { /* KFENCE metadata per guarded allocation. */ struct kfence_metadata { struct list_head list __guarded_by(&kfence_freelist_lock); /* Freelist no= de. */ - struct rcu_head rcu_head; /* For delayed freeing. */ + union { + struct rcu_head rcu_head; /* For delayed freeing. */ + struct llist_node llnode; /* For kfree_nolock(). */ + }; =20 /* * Lock protecting below data; to ensure consistency of the below data, diff --git a/mm/slub.c b/mm/slub.c index 044db93d64a0..2d7648b96bfa 100644 --- a/mm/slub.c +++ b/mm/slub.c @@ -4050,6 +4050,7 @@ static void flush_all(struct kmem_cache *s) =20 struct deferred_percpu_work { struct llist_head objects; + struct llist_head objects_kfence; struct llist_head objects_by_rcu; struct llist_head rcu_sheaves; struct irq_work work; @@ -4059,6 +4060,7 @@ static void deferred_percpu_work_fn(struct irq_work *= work); =20 static DEFINE_PER_CPU(struct deferred_percpu_work, deferred_percpu_work) = =3D { .objects =3D LLIST_HEAD_INIT(objects), + .objects_kfence =3D LLIST_HEAD_INIT(objects_kfence), .objects_by_rcu =3D LLIST_HEAD_INIT(objects_by_rcu), .rcu_sheaves =3D LLIST_HEAD_INIT(rcu_sheaves), .work =3D IRQ_WORK_INIT(deferred_percpu_work_fn), @@ -6399,6 +6401,13 @@ static void deferred_percpu_work_fn(struct irq_work = *work) stat(s, FREE_SLOWPATH); } =20 + llnode =3D llist_del_all(&dpw->objects_kfence); + llist_for_each_safe(pos, t, llnode) { + void *obj =3D kfence_llnode_to_obj(pos); + + __kfence_free(obj); + } + llnode =3D llist_del_all(&dpw->objects_by_rcu); llist_for_each_safe(pos, t, llnode) { void *head =3D pos; @@ -6431,6 +6440,21 @@ static void defer_free(struct kmem_cache *s, void *o= bj) irq_work_queue(&dpw->work); } =20 +static void defer_free_kfence(void *obj) +{ + struct deferred_percpu_work *dpw; + struct llist_node *llnode; + + /* kasan_reset_tag() is not necessary, kfence objects are not tagged */ + llnode =3D kfence_obj_to_llnode(obj); + + guard(preempt)(); + + dpw =3D this_cpu_ptr(&deferred_percpu_work); + if (llist_add(llnode, &dpw->objects_kfence)) + irq_work_queue(&dpw->work); +} + void defer_kfree_rcu(struct kvfree_rcu_head *head) { struct deferred_percpu_work *dpw; @@ -6758,10 +6782,13 @@ EXPORT_SYMBOL(kfree); /* * Can be called while holding raw_spinlock_t or from IRQ and NMI, * but ONLY for objects allocated by kmalloc_nolock(). - * Debug checks (like kmemleak and kfence) were skipped on allocation, - * hence + * + * In case kmemleak is enabled, + * * obj =3D kmalloc(); kfree_nolock(obj); - * will miss kmemleak/kfence book keeping and will cause false positives. + * + * will miss kmemleak book keeping and will cause false positives. + * * large_kmalloc is not supported either. */ void kfree_nolock(const void *object) @@ -6793,6 +6820,12 @@ void kfree_nolock(const void *object) * since they take spinlocks or not safe from any context. */ kmsan_slab_free(s, x); + + if (is_kfence_address(x)) { + defer_free_kfence(x); + return; + } + /* * If KASAN finds a kernel bug it will do kasan_report_invalid_free() * which will call raw_spin_lock_irqsave() which is technically --=20 2.55.0 From nobody Tue Sep 29 13:20:32 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8C97E3264C4; Fri, 7 Aug 2026 13:50:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786110666; cv=none; b=mXgdkyuOZWCi+8iynum0ikgjL2XgiCmObCSD8GedZkiDapDKYrnKGp8I/yreln9QW48geGuCBXSZczUQA1+QqY9Dx/80Uo1SsB/vl6Bqa2BfFD5FyzVI9hIoALrGMXc7ZCukFnlWukP0eCiHsfdAWFXaPbSPwwvp6DCtJuVvR1w= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786110666; c=relaxed/simple; bh=c2/aImCi2qV2RtDsuU+SgXIPunHKLmZ4n0PCybH7n/s=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=bmraSllt3AQW3VkyuSpB2KyGqmhcCE9RuGUm2EuByPeS5s7letSVfxMGmvSyT+r2Wkb6ihB9CJctPUJ1f33zX80O0/78hk/cEHCJCD6VUWxRewPya4N4LiKQY95Se26CbPpe2CzlOPybgWWMhHJzwJUseGr0J2Lb5SO4ypunPIE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=bftfFwcV; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="bftfFwcV" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 343681F00A3A; Fri, 7 Aug 2026 13:50:51 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786110657; bh=Qw5EXLRgG7MVbyyzo5Hfu/12GgcU/fQ9naW0jlMMCd0=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=bftfFwcV3aDHjtMAwILI5O79RrtunZvph2GfZ9MMZn6kCchr0mEVG/ybaswejvRLv zlukEzJFtGZZeEBuRq2bouOGb9Md9UY4OzBhzrlTHCeJ8hnJZUHx31OveMSFd0WEo8 rytPjymjr0QPgGh7kEBi40uKhi3kytwMtwq4QfMFB1uMw+YtlNdyy3fb5bKX0u+u+A SfIie0u/oyt6/TO3TDYcftfy7mMve8wXr9GFyWhWakLhp9E4HgylEkZK+glgFSYYJq gytQRdLKrenFLW6NRiZI69kZ6tKkt0gh9M16PQvzC3cd40dRpfFij+/C6fgbpCIkSK LPNLY9CjGjJqQ== From: "Vlastimil Babka (SUSE)" Date: Fri, 07 Aug 2026 15:50:29 +0200 Subject: [PATCH RFC 3/5] mm/slab, kmemleak: handle kmemleak freeing in kfree_nolock() Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260807-kfree_nolock_kmalloc-v1-3-ba993cbf7a60@kernel.org> References: <20260807-kfree_nolock_kmalloc-v1-0-ba993cbf7a60@kernel.org> In-Reply-To: <20260807-kfree_nolock_kmalloc-v1-0-ba993cbf7a60@kernel.org> To: Harry Yoo , Alexei Starovoitov , Alexander Potapenko , Marco Elver , Sumit Semwal , =?utf-8?q?Christian_K=C3=B6nig?= , Catalin Marinas , Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Sebastian Andrzej Siewior , Clark Williams , Steven Rostedt Cc: Andrew Morton , Hao Li , Christoph Lameter , David Rientjes , Roman Gushchin , bpf@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, Dmitry Vyukov , linux-media@vger.kernel.org, dri-devel@lists.freedesktop.org, linaro-mm-sig@lists.linaro.org, kasan-dev@googlegroups.com, Dietmar Eggemann , Ben Segall , Mel Gorman , Valentin Schneider , K Prateek Nayak , linux-rt-devel@lists.linux.dev, "Vlastimil Babka (SUSE)" X-Mailer: b4 0.15.2 Kmemleak handling is one of the reasons why kfree_nolock() cannot currently handle kmalloc() objects, because calling kmemleak_free() would involve spinning on its internal raw spinlocks. Kmemleak is a debugging mechanism so we could simply defer all kfree_nolock() to irq_work if it's enabled, and eat the extra cost. But that would be unnecessary pessimistic. We expect kfree_nolock() will be still mostly called on objects from kmalloc_nolock() that are not registered in kmemleak so they still don't need any deferred freeing. Thus introduce kmemleak_may_need_free() that can check if the object is registered. This is done using __lookup_object() performed under a raw_spin_trylock_irqsave(), which is safe to attempt from kfree_nolock() (except from a NMI on a !CONFIG_SMP system). When that trylock fails or can't be attempted, we however must assume the object might be registered, and defer the freeing. The ordering of kmsan/kasan handling and kmemleak is also different from what kfree() is doing, but as explained in the comment, it should be OK. Signed-off-by: Vlastimil Babka (SUSE) Reviewed-by: Catalin Marinas --- include/linux/kmemleak.h | 17 +++++++++++++++++ mm/kmemleak.c | 42 ++++++++++++++++++++++++++++++++++++++++++ mm/slub.c | 34 ++++++++++++++++++++++++++-------- 3 files changed, 85 insertions(+), 8 deletions(-) diff --git a/include/linux/kmemleak.h b/include/linux/kmemleak.h index fbd424b2abb1..52f75f10a9ce 100644 --- a/include/linux/kmemleak.h +++ b/include/linux/kmemleak.h @@ -22,6 +22,7 @@ extern void kmemleak_alloc_percpu(const void __percpu *pt= r, size_t size, extern void kmemleak_vmalloc(const struct vm_struct *area, size_t size, gfp_t gfp) __ref; extern void kmemleak_free(const void *ptr) __ref; +bool kmemleak_may_need_free(const void *ptr) __ref; extern void kmemleak_free_part(const void *ptr, size_t size) __ref; extern void kmemleak_free_percpu(const void __percpu *ptr) __ref; extern void kmemleak_update_trace(const void *ptr) __ref; @@ -50,6 +51,14 @@ static inline void kmemleak_free_recursive(const void *p= tr, slab_flags_t flags) kmemleak_free(ptr); } =20 +static inline bool kmemleak_may_need_free_recursive(const void *ptr, slab_= flags_t flags) +{ + if (!(flags & SLAB_NOLEAKTRACE)) + return kmemleak_may_need_free(ptr); + + return false; +} + static inline void kmemleak_erase(void **ptr) { *ptr =3D NULL; @@ -86,6 +95,14 @@ static inline void kmemleak_free_part(const void *ptr, s= ize_t size) static inline void kmemleak_free_recursive(const void *ptr, slab_flags_t f= lags) { } +static inline bool kmemleak_may_need_free(const void *ptr) +{ + return false; +} +static inline bool kmemleak_may_need_free_recursive(const void *ptr, slab_= flags_t flags) +{ + return false; +} static inline void kmemleak_free_percpu(const void __percpu *ptr) { } diff --git a/mm/kmemleak.c b/mm/kmemleak.c index 7c7ba17ce7af..e3560ce82632 100644 --- a/mm/kmemleak.c +++ b/mm/kmemleak.c @@ -1168,6 +1168,48 @@ void __ref kmemleak_free(const void *ptr) } EXPORT_SYMBOL_GPL(kmemleak_free); =20 +/** + * kmemleak_may_need_free - check if object is registered + * @ptr: pointer to beginning of the object + * + * This function is called from the kernel allocator when an object should= be + * freed but the caller context might be unsafe to spin on the internal lo= cks. + * + * It will therefore only use trylock and thus might return a false positi= ve + * if the trylock fails and the status cannot be determined. + * + * For objects that (might) need free, the allocator has to defer the actu= al + * freeing to a safe context. + * + * The assumption is that most objects freed from the unsafe context are a= lso + * allocated in such context and thus are not registered in kmemleak, so i= t's + * unlikely the defered freeing will be necessary just because kmemleak is + * enabled. + */ +bool __ref kmemleak_may_need_free(const void *ptr) +{ + unsigned long flags; + struct kmemleak_object *object; + + pr_debug("%s(0x%px)\n", __func__, ptr); + + if (!kmemleak_free_enabled || !ptr || IS_ERR(ptr)) + return false; + + /* On UP, raw_spin_trylock() always succeeds even when it is locked */ + if (!IS_ENABLED(CONFIG_SMP) && in_nmi()) + return true; + + if (!raw_spin_trylock_irqsave(&kmemleak_lock, flags)) + return true; + + object =3D __lookup_object((unsigned long)ptr, 0, 0); + + raw_spin_unlock_irqrestore(&kmemleak_lock, flags); + + return !!object; +} + /** * kmemleak_free_part - partially unregister a previously registered object * @ptr: pointer to the beginning or inside the object. This also diff --git a/mm/slub.c b/mm/slub.c index 2d7648b96bfa..423b5bdb910b 100644 --- a/mm/slub.c +++ b/mm/slub.c @@ -6390,6 +6390,8 @@ static void deferred_percpu_work_fn(struct irq_work *= work) /* Point 'x' back to the beginning of allocated object */ x -=3D s->offset; =20 + kmemleak_free_recursive(x, s->flags); + /* * We used freepointer in 'x' to link 'x' into df->objects. * Clear it to NULL to avoid false positive detection @@ -6403,8 +6405,15 @@ static void deferred_percpu_work_fn(struct irq_work = *work) =20 llnode =3D llist_del_all(&dpw->objects_kfence); llist_for_each_safe(pos, t, llnode) { + struct kmem_cache *s; + struct slab *slab; void *obj =3D kfence_llnode_to_obj(pos); =20 + slab =3D virt_to_slab(obj); + s =3D slab->slab_cache; + + kmemleak_free_recursive(obj, s->flags); + __kfence_free(obj); } =20 @@ -6781,15 +6790,10 @@ EXPORT_SYMBOL(kfree); =20 /* * Can be called while holding raw_spinlock_t or from IRQ and NMI, - * but ONLY for objects allocated by kmalloc_nolock(). - * - * In case kmemleak is enabled, + * but may defer freeing to irq_work() in some cases. * - * obj =3D kmalloc(); kfree_nolock(obj); - * - * will miss kmemleak book keeping and will cause false positives. - * - * large_kmalloc is not supported either. + * Intended mainly for objects allocated from kmalloc_nolock(), but can ha= ndle + * also kmem_cache_alloc() and kmalloc() objects, except large_kmalloc. */ void kfree_nolock(const void *object) { @@ -6844,9 +6848,23 @@ void kfree_nolock(const void *object) */ kasan_slab_free(s, x, false, false, /* skip quarantine */true); =20 + /* + * with kfree() the kmemleak handling happens much sooner, but for + * defering we need to write llnode to the object's freepointer so + * we should have it in the state when it's no longer treated as + * allocated by kasan etc. + * + * defer_free will also reset the pointer tag, but it's ok to do a + * deferred kmemleak_free() using the untagged pointer, because + * __lookup_object() resets the tag anyway + */ + if (unlikely(kmemleak_may_need_free_recursive(x, s->flags))) + goto defer; + if (likely(can_free_to_pcs(slab)) && likely(free_to_pcs(s, x, false))) return; =20 +defer: /* * __slab_free() can locklessly cmpxchg16 into a slab, but then it might * need to take spin_lock for further processing. --=20 2.55.0 From nobody Tue Sep 29 13:20:32 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0CC7B36920C; Fri, 7 Aug 2026 13:51:08 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786110672; cv=none; b=FsryyupgjXKoGZlBPysHMHrideeM7tcWt0HYUoY0UE+tK+EB0aUE2EoTm5CTICU2dZiP/akuwN86Q+zQLzshja4WfbVvPTgOP7g3GYVxQUn9a4dQ7t4SB4XeoPAxqJtXUo7C6vVQcJ0z+SWvqEHVLTyI+RCzjz12biW3BjOPEXI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786110672; c=relaxed/simple; bh=pOlHJ19D5l1vM3gKIor39Psb3JZKgPFdPbJedFSXZ3A=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=XIqAc1nyjtm1mlB/ncvKZ6cgGK1EqwJaGVrLybMsuf2R8RYhXWQQFEIRTGz1Ji5YL7u/qK/5M4tEcM9N2abjZvPvl4NMH04rr8G2ShSSWuHTfjz82B+D+lBf0uRKNvbngabfdQrkGpk78q2tcvlrJp7GFCm/K2/QMDtw7M/kqHQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=e4VSmJSt; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="e4VSmJSt" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 3E7291F00ACF; Fri, 7 Aug 2026 13:50:58 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786110664; bh=UGBELfa0PNJRa3PvbOBLlp8SGyFGd7t+kM65vGIAZwc=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=e4VSmJSt73tA02FcE2cwM/eyfhDNj1IybKNAoQq4cuYxrXygc2XXmaK6ICGsSOOuU XmsxUDzNigL3kFl+StJRA1b/0LqYThJmhrHZJYv0q+HgjLVDqkaqaOnqt2eEV3kHAH eS53QQHpPBkK2d5djTzlbNJOR6puKjMDn3qplP1y5aX0E7LA5p/yt3AsewlXAoE+gi 0NAmhX9xVb6MpvGI8dK5dxg4I/Vnyr9RRkEgxqW8IOHCUgQL72FE88sJU+4wEbiZDp VMBnUWwtd+bbroav/X3QGzfJXLxMuZKj05+itt98UINUISZpJ6fHQZO4VXRWTIX8zv N0Ng6BuX6uueQ== From: "Vlastimil Babka (SUSE)" Date: Fri, 07 Aug 2026 15:50:30 +0200 Subject: [PATCH RFC 4/5] mm/slab: handle large_kmalloc objects in kfree_nolock() Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260807-kfree_nolock_kmalloc-v1-4-ba993cbf7a60@kernel.org> References: <20260807-kfree_nolock_kmalloc-v1-0-ba993cbf7a60@kernel.org> In-Reply-To: <20260807-kfree_nolock_kmalloc-v1-0-ba993cbf7a60@kernel.org> To: Harry Yoo , Alexei Starovoitov , Alexander Potapenko , Marco Elver , Sumit Semwal , =?utf-8?q?Christian_K=C3=B6nig?= , Catalin Marinas , Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Sebastian Andrzej Siewior , Clark Williams , Steven Rostedt Cc: Andrew Morton , Hao Li , Christoph Lameter , David Rientjes , Roman Gushchin , bpf@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, Dmitry Vyukov , linux-media@vger.kernel.org, dri-devel@lists.freedesktop.org, linaro-mm-sig@lists.linaro.org, kasan-dev@googlegroups.com, Dietmar Eggemann , Ben Segall , Mel Gorman , Valentin Schneider , K Prateek Nayak , linux-rt-devel@lists.linux.dev, "Vlastimil Babka (SUSE)" X-Mailer: b4 0.15.2 Large kmalloc objects is the only remaining case that kfree_nolock() cannot handle from kmalloc() allocations. Note kmalloc_nolock() does not return large kmalloc objects. Supporting them is however mostly straigtforward. free_large_kmalloc() calls kmsan and kasan hooks that should be safe and similar to those called in kfree_nolock(). Freeing the pages can be handled by free_frozen_pages_nolock(). The only obstacle is kmemleak_free(), which we can solve by deferring when necessary, the same way as done for small kmalloc objects. Signed-off-by: Vlastimil Babka (SUSE) --- mm/slub.c | 74 ++++++++++++++++++++++++++++++++++++++++++++++++++++-------= ---- 1 file changed, 62 insertions(+), 12 deletions(-) diff --git a/mm/slub.c b/mm/slub.c index 423b5bdb910b..3be98faa9f0d 100644 --- a/mm/slub.c +++ b/mm/slub.c @@ -4051,6 +4051,7 @@ static void flush_all(struct kmem_cache *s) struct deferred_percpu_work { struct llist_head objects; struct llist_head objects_kfence; + struct llist_head objects_large_kmalloc; struct llist_head objects_by_rcu; struct llist_head rcu_sheaves; struct irq_work work; @@ -4061,6 +4062,7 @@ static void deferred_percpu_work_fn(struct irq_work *= work); static DEFINE_PER_CPU(struct deferred_percpu_work, deferred_percpu_work) = =3D { .objects =3D LLIST_HEAD_INIT(objects), .objects_kfence =3D LLIST_HEAD_INIT(objects_kfence), + .objects_large_kmalloc =3D LLIST_HEAD_INIT(objects_large_kmalloc), .objects_by_rcu =3D LLIST_HEAD_INIT(objects_by_rcu), .rcu_sheaves =3D LLIST_HEAD_INIT(rcu_sheaves), .work =3D IRQ_WORK_INIT(deferred_percpu_work_fn), @@ -6365,6 +6367,21 @@ static void free_to_pcs_bulk(struct kmem_cache *s, s= ize_t size, void **p) } } =20 +static inline void +__free_large_kmalloc_page(struct page *page, unsigned int free_flags) +{ + unsigned int order =3D compound_order(page); + + mod_lruvec_page_state(page, NR_SLAB_UNRECLAIMABLE_B, + -(PAGE_SIZE << order)); + __ClearPageLargeKmalloc(page); + + if (free_flags & SLAB_FREE_NOLOCK) + free_frozen_pages_nolock(page, order); + else + free_frozen_pages(page, order); +} + /* * In PREEMPT_RT irq_work runs in per-cpu kthread, so it's safe * to take sleeping spin_locks from __slab_free(). @@ -6417,6 +6434,15 @@ static void deferred_percpu_work_fn(struct irq_work = *work) __kfence_free(obj); } =20 + llnode =3D llist_del_all(&dpw->objects_large_kmalloc); + llist_for_each_safe(pos, t, llnode) { + struct page *page =3D virt_to_page(pos); + + kmemleak_free(pos); + + __free_large_kmalloc_page(page, SLAB_FREE_DEFAULT); + } + llnode =3D llist_del_all(&dpw->objects_by_rcu); llist_for_each_safe(pos, t, llnode) { void *head =3D pos; @@ -6464,6 +6490,24 @@ static void defer_free_kfence(void *obj) irq_work_queue(&dpw->work); } =20 +static void defer_free_large_kmalloc(void *obj) +{ + struct deferred_percpu_work *dpw; + struct llist_node *llnode; + + /* + * we can simply use the first word of the large kmalloc object + * for the llnode, as there's no ctor or TYPESAFE_BY_RCU + */ + llnode =3D kasan_reset_tag(obj); + + guard(preempt)(); + + dpw =3D this_cpu_ptr(&deferred_percpu_work); + if (llist_add(llnode, &dpw->objects_large_kmalloc)) + irq_work_queue(&dpw->work); +} + void defer_kfree_rcu(struct kvfree_rcu_head *head) { struct deferred_percpu_work *dpw; @@ -6712,9 +6756,11 @@ size_t ksize(const void *objp) } EXPORT_SYMBOL(ksize); =20 -static void free_large_kmalloc(struct page *page, void *object) +static void free_large_kmalloc(struct page *page, void *object, + unsigned int free_flags) { unsigned int order =3D compound_order(page); + bool nolock =3D free_flags & SLAB_FREE_NOLOCK; =20 if (WARN_ON_ONCE(!PageLargeKmalloc(page))) { dump_page(page, "Not a kmalloc allocation"); @@ -6724,14 +6770,16 @@ static void free_large_kmalloc(struct page *page, v= oid *object) if (WARN_ON_ONCE(order =3D=3D 0)) pr_warn_once("object pointer: 0x%p\n", object); =20 - kmemleak_free(object); + if (!nolock) + kmemleak_free(object); + kasan_kfree_large(object); kmsan_kfree_large(object); =20 - mod_lruvec_page_state(page, NR_SLAB_UNRECLAIMABLE_B, - -(PAGE_SIZE << order)); - __ClearPageLargeKmalloc(page); - free_frozen_pages(page, order); + if (unlikely(nolock && kmemleak_may_need_free(object))) + defer_free_large_kmalloc(object); + else + __free_large_kmalloc_page(page, free_flags); } =20 /* @@ -6753,7 +6801,7 @@ void kvfree_rcu_cb(struct rcu_head *head) if (slab) slab_free(slab->slab_cache, slab, obj, _RET_IP_); else - free_large_kmalloc(page, obj); + free_large_kmalloc(page, obj, SLAB_FREE_DEFAULT); } } =20 @@ -6779,7 +6827,7 @@ void kfree(const void *object) slab =3D page_slab(page); if (!slab) { /* kmalloc_nolock() doesn't support large kmalloc */ - free_large_kmalloc(page, (void *)object); + free_large_kmalloc(page, (void *)object, SLAB_FREE_DEFAULT); return; } =20 @@ -6793,10 +6841,11 @@ EXPORT_SYMBOL(kfree); * but may defer freeing to irq_work() in some cases. * * Intended mainly for objects allocated from kmalloc_nolock(), but can ha= ndle - * also kmem_cache_alloc() and kmalloc() objects, except large_kmalloc. + * also kmem_cache_alloc() and kmalloc() objects, including large_kmalloc. */ void kfree_nolock(const void *object) { + struct page *page; struct slab *slab; struct kmem_cache *s; void *x =3D (void *)object; @@ -6804,9 +6853,10 @@ void kfree_nolock(const void *object) if (unlikely(ZERO_OR_NULL_PTR(object))) return; =20 - slab =3D virt_to_slab(object); + page =3D virt_to_page(object); + slab =3D page_slab(page); if (unlikely(!slab)) { - WARN_ONCE(1, "large_kmalloc is not supported by kfree_nolock()"); + free_large_kmalloc(page, (void *)object, SLAB_FREE_NOLOCK); return; } =20 @@ -7167,7 +7217,7 @@ int build_detached_freelist(struct kmem_cache *s, siz= e_t size, if (!s) { /* Handle kalloc'ed objects */ if (!slab) { - free_large_kmalloc(page, object); + free_large_kmalloc(page, object, SLAB_FREE_DEFAULT); df->slab =3D NULL; return size; } --=20 2.55.0 From nobody Tue Sep 29 13:20:32 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 36E4E312807; Fri, 7 Aug 2026 13:51:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786110679; cv=none; b=uCHXvSXeF3FJa3SSYIeQKReY64SfMFYI/23csSa29Lp1SN51KTFdl1mhw4rF/pRJUDzGf5azj2x2ZdZSsI5cQWijCKDUBhpFMWmSAtGm3zZNjGikQdZvvMvyoqRewpce81z74hVXqgh+2Fk9JV+Wi0Oy6+wzzcYS/5VY/C9cME8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786110679; c=relaxed/simple; bh=ULv5EHNv0omDKMBobosi5APF+equtZsKapOv2iw4VQs=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=jbrx+yOS6XC6ospOIVQV6vXASUfDO8Chw/xQ52TNUSX9cjPeaIPON66mJhFZnY8H15XJrDWXYy1XpHOCMXLxc4LS5LM6lnsKD8ln2T0x9hN8lKQ7rBLzCuL5ikMwukC884iyWTnWgOrfT9UkWiAxcr2tRJfntFXCs1hQsdO6QRo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=PXYh6M2K; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="PXYh6M2K" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 4D01C1F00A3D; Fri, 7 Aug 2026 13:51:05 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786110671; bh=JsBlNoLXapzvhaOZfFVgJmB3wffwQ8ZUd0cI3plKWLk=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=PXYh6M2KkgQHGOcPs4VHfkvwxJyfSKuED6zAeHwMz8mmUUgtOHTbSiQM5BJieAoLO H/WV7XvvSbkldcw2ezKoX6DKUZtAcjDvgjCWVLh1TbnJnQlGY+u4pRC+AIJ1cr/4s+ JByHYpe+ebHxFnreqVNPxaZ5n/PFOCKp9FCVEt1gVdNWAjj4teAKaCf/xFOIYta2uQ 234zl/9rw3uirz1ArxlCvIIddHjOso0Sf2s3/7xSRv4/CUeZyw0WEXcN1ojHc4WL4H w3jDfaKIq5AcnXo1qggs2ROLMdMYAzKmTRqpoIncaxXUWYTKO89frDEQNSedcwHZs4 lxnBzFqghs4Ig== From: "Vlastimil Babka (SUSE)" Date: Fri, 07 Aug 2026 15:50:31 +0200 Subject: [PATCH RFC 5/5] sched: use kfree_nolock() instead of kfree_rcu() Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260807-kfree_nolock_kmalloc-v1-5-ba993cbf7a60@kernel.org> References: <20260807-kfree_nolock_kmalloc-v1-0-ba993cbf7a60@kernel.org> In-Reply-To: <20260807-kfree_nolock_kmalloc-v1-0-ba993cbf7a60@kernel.org> To: Harry Yoo , Alexei Starovoitov , Alexander Potapenko , Marco Elver , Sumit Semwal , =?utf-8?q?Christian_K=C3=B6nig?= , Catalin Marinas , Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Sebastian Andrzej Siewior , Clark Williams , Steven Rostedt Cc: Andrew Morton , Hao Li , Christoph Lameter , David Rientjes , Roman Gushchin , bpf@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, Dmitry Vyukov , linux-media@vger.kernel.org, dri-devel@lists.freedesktop.org, linaro-mm-sig@lists.linaro.org, kasan-dev@googlegroups.com, Dietmar Eggemann , Ben Segall , Mel Gorman , Valentin Schneider , K Prateek Nayak , linux-rt-devel@lists.linux.dev, "Vlastimil Babka (SUSE)" X-Mailer: b4 0.15.2 In set_cpus_allowed_force() we use kfree_rcu() because kfree() is unsafe under p->pi_lock. With kfree_nolock() now being able to free arbitrary kmalloc() objects, we can switch to kfree_nolock() and avoid the unnecessary rcu grace period delay. Only in some cases the freeing might be deferred to irq_work(). Signed-off-by: Vlastimil Babka (SUSE) --- kernel/sched/core.c | 9 ++------- kernel/sched/sched.h | 7 +------ 2 files changed, 3 insertions(+), 13 deletions(-) diff --git a/kernel/sched/core.c b/kernel/sched/core.c index 96226707c2f6..d2929e4e23f1 100644 --- a/kernel/sched/core.c +++ b/kernel/sched/core.c @@ -2807,20 +2807,15 @@ void set_cpus_allowed_force(struct task_struct *p, = const struct cpumask *new_mas .user_mask =3D NULL, .flags =3D SCA_USER, /* clear the user requested mask */ }; - union cpumask_rcuhead { - cpumask_t cpumask; - struct rcu_head rcu; - }; =20 scoped_guard (__task_rq_lock, p) do_set_cpus_allowed(p, &ac); =20 /* * Because this is called with p->pi_lock held, it is not possible - * to use kfree() here (when PREEMPT_RT=3Dy), therefore punt to using - * kfree_rcu(). + * to use kfree() here (when PREEMPT_RT=3Dy), thus use kfree_nolock() */ - kfree_rcu((union cpumask_rcuhead *)ac.user_mask, rcu); + kfree_nolock(ac.user_mask); } =20 int dup_user_cpus_ptr(struct task_struct *dst, struct task_struct *src, diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h index 56acf502ba26..6a8d0578e963 100644 --- a/kernel/sched/sched.h +++ b/kernel/sched/sched.h @@ -2886,12 +2886,7 @@ static inline bool task_allowed_on_cpu(struct task_s= truct *p, int cpu) =20 static inline cpumask_t *alloc_user_cpus_ptr(int node) { - /* - * See set_cpus_allowed_force() above for the rcu_head usage. - */ - int size =3D max_t(int, cpumask_size(), sizeof(struct rcu_head)); - - return kmalloc_node(size, GFP_KERNEL, node); + return kmalloc_node(cpumask_size(), GFP_KERNEL, node); } =20 static inline struct task_struct *get_push_task(struct rq *rq) --=20 2.55.0