From nobody Fri Oct 2 11:40:43 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5FC9138D; Sat, 1 Aug 2026 12:25:42 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785587143; cv=none; b=Ck4Ceom9z8DO7gk8BdO0KJr47kLa9m5XXWNG3lEEjaWEEe57INdUI9iWkwqbmOgnpezWT9e+uYpH4z5JmdUNad2HMZDzKQv7qBh6jOs1wiIOUsKfxEJNr2aQxWakPH9hieB1ILr+QeYEyqCTG1My4Klv+NYQb+deqOMcM7LSxNo= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785587143; c=relaxed/simple; bh=4gG5xB2QZ5YE8hLuOHmyuKaQd3LO5mKZnrgRUZNxWJ4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=IGJabZwZxQEPTjUaLbYEybMknWpCgKC/kbyfzr87wIUs4znHZwyEk1jbrhyIGojzFjLEdkH7h2XI6as4zaQotGW3DpR8uDZOc10rFbGvWakpPCHa5W7NMloYdN/p0J9cPWy4qHy9JGPxKZlrtkkCfo++GdyM7YHJZcv5dsNKeUQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=XgBVE3Gb; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="XgBVE3Gb" Received: by smtp.kernel.org (Postfix) with ESMTPSA id C63901F00ACF; Sat, 1 Aug 2026 12:25:38 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785587142; bh=fR5GQklAEmMIRAxF5ZINhcf10poKg1/4k56TIs+euS8=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=XgBVE3GbYz9QXwCf2ZRYlPhwPyUGUCO6C6tk9Pc2sQnSY35UCRgBrOJw4nNPSBjJe 7mKzbtf9C50CiWxQ+T6Ch+6bbkQUy64hdggkkV1/GPNmT9VegCkXhWu1RnPPl9AOGh JwTWoJKtKGgQsh7iwJKb/A3T4FRUr1gvYiQOPe9eRlvISnD38j5w7BZiUyc/asvyxk jwMlqIeV556EQIBILYc4nBH6N7PoML3cOtJeTYYA7nkEEsj9pMA66ZY92+kc9ncQl0 h8DQx1B9KOEObFLTFXfe/gYKYXAoodT2LqnpKvKNsVcWzClr7BOtVsn0dTipceSeuQ rGy2YU4jTfIDQ== From: Harry Yoo To: Danielle Costantino , Shakeel Butt , Suren Baghdasaryan , Hao Ge , Vlastimil Babka , Harry Yoo , Andrew Morton , Hao Li , Christoph Lameter , David Rientjes , Roman Gushchin , Kees Cook , Pedro Falcato , "Liam R . Howlett" Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org, stable@vger.kernel.org Subject: [7.1.y 1/3] mm/slab: decouple SLAB_NO_SHEAVES from SLAB_NO_OBJ_EXT Date: Sat, 1 Aug 2026 12:25:31 +0000 Message-ID: <20260801122533.319272-2-harry@kernel.org> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260801122533.319272-1-harry@kernel.org> References: <20260801122533.319272-1-harry@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Harry Yoo (Oracle)" commit 982e31382d9a1a3c8c4e6a13702a53711f4efe9f upstream. Bootstrap caches are created with SLAB_NO_OBJ_EXT to disallow sheaves and obj_exts. To allow disabling obj_exts while allowing sheaves, decouple SLAB_NO_SHEAVES from SLAB_NO_OBJ_EXT. Bootstrap caches now have both SLAB_NO_SHEAVES and SLAB_NO_OBJ_EXT. No functional change intended. Reviewed-by: Vlastimil Babka (SUSE) Signed-off-by: Harry Yoo (Oracle) Reviewed-by: Suren Baghdasaryan Link: https://patch.msgid.link/20260713-kmalloc-no-objext-v3-2-47c7bd138de7= @kernel.org Signed-off-by: Vlastimil Babka (SUSE) Signed-off-by: Harry Yoo --- include/linux/slab.h | 13 +++++++++++-- mm/slub.c | 10 ++++++---- 2 files changed, 17 insertions(+), 6 deletions(-) diff --git a/include/linux/slab.h b/include/linux/slab.h index 1a69255bb87f..fda7e24a762b 100644 --- a/include/linux/slab.h +++ b/include/linux/slab.h @@ -58,10 +58,13 @@ enum _slab_flag_bits { #endif _SLAB_OBJECT_POISON, _SLAB_CMPXCHG_DOUBLE, +#ifdef CONFIG_SLAB_OBJ_EXT _SLAB_NO_OBJ_EXT, -#if defined(CONFIG_SLAB_OBJ_EXT) && defined(CONFIG_64BIT) +#ifdef CONFIG_64BIT _SLAB_OBJ_EXT_IN_OBJ, #endif +#endif + _SLAB_NO_SHEAVES, _SLAB_FLAGS_LAST_BIT }; =20 @@ -239,8 +242,14 @@ enum _slab_flag_bits { #endif #define SLAB_TEMPORARY SLAB_RECLAIM_ACCOUNT /* Objects are short-lived */ =20 -/* Slab created using create_boot_cache */ +/* Slab caches without obj_exts array */ +#ifdef CONFIG_SLAB_OBJ_EXT #define SLAB_NO_OBJ_EXT __SLAB_FLAG_BIT(_SLAB_NO_OBJ_EXT) +#else +#define SLAB_NO_OBJ_EXT __SLAB_FLAG_UNUSED +#endif + +#define SLAB_NO_SHEAVES __SLAB_FLAG_BIT(_SLAB_NO_SHEAVES) =20 #if defined(CONFIG_SLAB_OBJ_EXT) && defined(CONFIG_64BIT) #define SLAB_OBJ_EXT_IN_OBJ __SLAB_FLAG_BIT(_SLAB_OBJ_EXT_IN_OBJ) diff --git a/mm/slub.c b/mm/slub.c index fa24b985a783..33e64f21ecc1 100644 --- a/mm/slub.c +++ b/mm/slub.c @@ -7726,12 +7726,12 @@ static unsigned int calculate_sheaf_capacity(struct= kmem_cache *s, return 0; =20 /* - * Bootstrap caches can't have sheaves for now (SLAB_NO_OBJ_EXT). + * Bootstrap caches can't have sheaves for now (SLAB_NO_SHEAVES). * SLAB_NOLEAKTRACE caches (e.g., kmemleak's object_cache) must not * have sheaves to avoid recursion when sheaf allocation triggers * kmemleak tracking. */ - if (s->flags & (SLAB_NO_OBJ_EXT | SLAB_NOLEAKTRACE)) + if (s->flags & (SLAB_NO_SHEAVES | SLAB_NOLEAKTRACE)) return 0; =20 /* @@ -8513,7 +8513,8 @@ void __init kmem_cache_init(void) =20 create_boot_cache(kmem_cache_node, "kmem_cache_node", sizeof(struct kmem_cache_node), - SLAB_HWCACHE_ALIGN | SLAB_NO_OBJ_EXT, 0, 0); + SLAB_HWCACHE_ALIGN | SLAB_NO_SHEAVES | SLAB_NO_OBJ_EXT, + 0, 0); =20 hotplug_node_notifier(slab_memory_callback, SLAB_CALLBACK_PRI); =20 @@ -8523,7 +8524,8 @@ void __init kmem_cache_init(void) create_boot_cache(kmem_cache, "kmem_cache", offsetof(struct kmem_cache, per_node) + nr_node_ids * sizeof(struct kmem_cache_per_node_ptrs), - SLAB_HWCACHE_ALIGN | SLAB_NO_OBJ_EXT, 0, 0); + SLAB_HWCACHE_ALIGN | SLAB_NO_SHEAVES | SLAB_NO_OBJ_EXT, + 0, 0); =20 kmem_cache =3D bootstrap(&boot_kmem_cache); kmem_cache_node =3D bootstrap(&boot_kmem_cache_node); --=20 2.53.0 From nobody Fri Oct 2 11:40:43 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DEE3F280035; Sat, 1 Aug 2026 12:25:45 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785587147; cv=none; b=rF9APQ3chrUgvlFYiPgdKfRLo0INPMJxJqu1TMJZtUgR7KryHDrd6gkze8LMHmUa3uBgr5NlbV//+i4qI2zQtTok48eM4PPRx8l+u0DJG/WqPVq+wvjmutj3vV2upAkyPa7x7NEiTvyyuuq7L6xyS36cRDndaQQ3wzUpGQJDuUM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785587147; c=relaxed/simple; bh=PR3me3Pq9mra7Mow3JXcUqquBDuQsrVeTSKm2iP3Ils=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=LAL+M5DE1BtQHTikeujk2FOtXtrzpNYVqbYLnDC0Cp3a+kniH4X8yH51O4RTfYzeawpFNqL0yBY0lnxuAfmQGP/cOfZNYh5LUKaPIQy7YvL7D7/iNUP0rSXYj125MUpCDIFXSNcqbaVtUOTNLSy1K9NXl1/rXKe2TFtFFb8n1Zo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=RzOwbCuD; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="RzOwbCuD" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 6AD8B1F00AC4; Sat, 1 Aug 2026 12:25:42 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785587145; bh=oCy5bIVtYQxaWyXI8Dl7cmjAZ6Z+ut7Uf+oQ6OT3GMU=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=RzOwbCuDSnCP913cQqRNIX5YpQ1FBuMuir92WwGtgL2MMHMBN03zq19hltowL6rTo 2coSPOwO4F21Ey6kPAMSl6fmWglrpAk+lgK0JbL59Um9Y4lVwHAV1Lc0AF1gNrAWyX J9EWAoZ/I3YVkXendqI1gGoXLn5b2u3dziNG6zH/b+FgVJkrzT8RITz9xRHDJp3wWH 0Oas6siKzXp7RMUbZiwu1AoAk9VMC6cnIaYi8PB86UtY30gb3Ketia+sWgcx2GriiU v5m86ov0+kx8KxnICb8rXPxAjsq9ehQ0MkzTm1dlcU8++KvszyjJSlDkGsF1kpl4W4 bLHkmhAwfOnLw== From: Harry Yoo To: Danielle Costantino , Shakeel Butt , Suren Baghdasaryan , Hao Ge , Vlastimil Babka , Harry Yoo , Andrew Morton , Hao Li , Christoph Lameter , David Rientjes , Roman Gushchin , Kees Cook , Pedro Falcato , "Liam R . Howlett" Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org, stable@vger.kernel.org Subject: [7.1.y 2/3] lib/alloc_tag: introduce mem_alloc_profiling_permanently_disabled() Date: Sat, 1 Aug 2026 12:25:32 +0000 Message-ID: <20260801122533.319272-3-harry@kernel.org> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260801122533.319272-1-harry@kernel.org> References: <20260801122533.319272-1-harry@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Harry Yoo (Oracle)" commit a37b0066a10aabf3c968b4566706fb866eaf9a85 upstream. mem_alloc_profiling_enabled() tells whether memalloc profiling is currently enabled. However, even when this function returns false, it can be enabled later. However, this is not enough. Some optimizations can be applied only when memalloc profiling is permanently disabled. For example, to skip the creation of KMALLOC_NO_OBJ_EXT caches at boot time, mem_profiling must be set to "never", "0" w/ debugging on, or have been shutdown so that it can no longer be enabled. Introduce mem_alloc_profiling_permanently_disabled() for this purpose. Signed-off-by: Harry Yoo (Oracle) Acked-by: Suren Baghdasaryan Link: https://patch.msgid.link/20260713-kmalloc-no-objext-v3-3-47c7bd138de7= @kernel.org Signed-off-by: Vlastimil Babka (SUSE) Signed-off-by: Harry Yoo --- include/linux/alloc_tag.h | 3 +++ lib/alloc_tag.c | 9 +++++++++ 2 files changed, 12 insertions(+) diff --git a/include/linux/alloc_tag.h b/include/linux/alloc_tag.h index 02de2ede560f..7e7cdc7612be 100644 --- a/include/linux/alloc_tag.h +++ b/include/linux/alloc_tag.h @@ -134,6 +134,8 @@ static inline bool mem_alloc_profiling_enabled(void) &mem_alloc_profiling_key); } =20 +bool mem_alloc_profiling_permanently_disabled(void); + static inline struct alloc_tag_counters alloc_tag_read(struct alloc_tag *t= ag) { struct alloc_tag_counters v =3D { 0, 0 }; @@ -239,6 +241,7 @@ static inline bool alloc_tag_is_inaccurate(struct alloc= _tag *tag) =20 #define DEFINE_ALLOC_TAG(_alloc_tag) static inline bool mem_alloc_profiling_enabled(void) { return false; } +static inline bool mem_alloc_profiling_permanently_disabled(void) { return= true; } static inline void alloc_tag_add(union codetag_ref *ref, struct alloc_tag = *tag, size_t bytes) {} static inline void alloc_tag_sub(union codetag_ref *ref, size_t bytes) {} diff --git a/lib/alloc_tag.c b/lib/alloc_tag.c index a9ab88f416b9..e3a923594604 100644 --- a/lib/alloc_tag.c +++ b/lib/alloc_tag.c @@ -26,6 +26,15 @@ static bool mem_profiling_support =3D true; static bool mem_profiling_support; #endif =20 +/* + * Memory allocation profiling is permanently disabled and cannot be enabl= ed. + * Must be called after setup_early_mem_profiling(). + */ +bool mem_alloc_profiling_permanently_disabled(void) +{ + return !mem_profiling_support; +} + static struct codetag_type *alloc_tag_cttype; =20 #ifdef CONFIG_ARCH_MODULE_NEEDS_WEAK_PER_CPU --=20 2.53.0 From nobody Fri Oct 2 11:40:43 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8D4F9266B46; Sat, 1 Aug 2026 12:25:49 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785587151; cv=none; b=VjzFMePWRtiluslURE6fcDhdiabe7YLMeKnrgwSYr45DWV4FWvoUNdIknSHRe9lSIBpEXPx2l9XkPK1zfxly6eQrI5Mg+aj8E8sXHhW7DRddB/NQPMriJwdDN45JHBLVQsVPIWpaTuOX9YBSFaRohl3CaeHLYQ/KrxUJe3VwgR0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785587151; c=relaxed/simple; bh=Vcqtx6Te1lxc0luE7KQcQUSSsH/Peo1vvpjhqX/CFlk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=LahYBiU8xdguaGp5SOFUpQ2cDhYb1qEzQylozkJ+c9HmG0LBMBGIe4ZUtIEDbKwiulF1z9RHlG3MKx6fYE/Dhaf2LiPvRV1UhWuBvxH5xEotUI03OeZavfgaCfDaI51qx1/xOUs0KsXHY66l+uk4KNTeADSm9Pv0e1FiStYR5BI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=STKiaHXo; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="STKiaHXo" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 0C4581F00ACF; Sat, 1 Aug 2026 12:25:45 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785587149; bh=XK4NNGYugEZGIiWxHOIeQgY29Y5r4aVTlZbmh3ehY5s=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=STKiaHXorkmOTLe5F4LPzOpkhJL5RD8LndokVEJAWZf32DwZEgmyIFj/SY/DyeynX Q8jxaW+JqAnEsPrShw6rhTfaelziRUvMYPPpgUHRDUYiOM9S83AIcMxpWzlxDMN27w /51FxR2XoqVXNst5pOZrwRjoATMCLrwCexi1fCytp4LgTQQlOfUmDLTQAn9cjpQboQ jP8zKlOYcpJk5gLxtnA72ADcOwlrH7GAIyFTs9eLEcBCPehuRxWZYHrcsjrdJ3aJF2 RjO5thrcEiW72ddjav3B/nJ/dz1jM8VfujHYnpbz4l9oByiByxijCNyu4qAbRFD/gy iiubV/2mCYNPA== From: Harry Yoo To: Danielle Costantino , Shakeel Butt , Suren Baghdasaryan , Hao Ge , Vlastimil Babka , Harry Yoo , Andrew Morton , Hao Li , Christoph Lameter , David Rientjes , Roman Gushchin , Kees Cook , Pedro Falcato , "Liam R . Howlett" Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org, stable@vger.kernel.org Subject: [7.1.y 3/3] mm/slab: prevent unbounded recursion in free path with new kmalloc type Date: Sat, 1 Aug 2026 12:25:33 +0000 Message-ID: <20260801122533.319272-4-harry@kernel.org> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260801122533.319272-1-harry@kernel.org> References: <20260801122533.319272-1-harry@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Harry Yoo (Oracle)" commit d9e6a7623938968e3752b67e37eaff097e559a54 upstream. Commit 280ea9c3154b ("mm/slab: avoid allocating slabobj_ext array from its own slab") avoided recursive allocation of obj_exts from kmalloc caches of the same size, by bumping the obj_exts array's allocation size whenever the array size equals the size of the object being allocated. However, as reported by Danielle Costantino and Shakeel Butt, even slabs from kmalloc caches of different sizes can form a cycle by allocating obj_exts arrays from each other [1]: What happened: a KMALLOC_NORMAL slab's obj_exts array (used by allocation profiling / memcg accounting) is itself kmalloc()'d from a KMALLOC_NORMAL cache, so the "slab holds another slab's obj_exts array" relation can form cycles. With sizeof(struct slabobj_ext) =3D=3D 16 and the host's geometry: - kmalloc-512 has 64 objects/slab -> array is 64*16 =3D=3D 1024 bytes, served from kmalloc-1k; - kmalloc-1k has 32 objects/slab -> array is 32*16 =3D=3D 512 bytes, served from kmalloc-512. A kmalloc-512 slab and a kmalloc-1k slab therefore hold each other's obj_exts array. Discarding one frees the other's array, which empties and discards that slab, which frees the first's array, and so on: __free_slab() -> free_slab_obj_exts() -> kfree() -> discard_slab() -> __free_slab() recurses along the cycle until the stack is exhausted. With memory allocation profiling, this allows unbounded recursion in the free path and led to a stack overflow on a production host in the Meta fleet [1]: BUG: TASK stack guard page was hit Oops: stack guard page RIP: 0010:kfree+0x8/0x5d0 Call Trace: __free_slab+0x66/0xc0 kfree+0x3f0/0x5d0 ... ( ~125x __free_slab <-> kfree ) ... do_syscall_64 It is proposed [1] to resolve this issue by always serving the obj_exts array allocation from kmalloc caches (or large kmalloc) of sizes larger than the object size. However, as pointed out by Vlastimil Babka [2], this can waste an excessive amount of memory as slabs from large kmalloc sizes (e.g. kmalloc-8k) generally need obj_exts arrays much smaller than the object size. Therefore, rather than bumping the size, let us take a different approach; disallow formation of cycles between kmalloc types when allocating obj_exts arrays. Currently, all obj_exts arrays are served from normal kmalloc caches. Cycles cannot be created if obj_exts arrays of normal kmalloc caches are served from a special kmalloc type that can never have obj_exts arrays. To achieve this, create a new kmalloc type called KMALLOC_NO_OBJ_EXT. KMALLOC_NO_OBJ_EXT caches are created with SLAB_NO_OBJ_EXT flag when either 1) memory allocation profiling is not permanently disabled, or 2) kmalloc types with a priority higher than KMALLOC_CGROUP are aliased with KMALLOC_NORMAL. Sheaf bootstrapping for KMALLOC_NO_OBJ_EXT caches now must be deferred because allocation of a barn can trigger obj_exts array allocation of normal kmalloc caches when the KMALLOC_NO_OBJ_EXT cache for that size is not ready yet. For simplicity, perform bootstrapping of sheaves for all kmalloc caches later. Introduce a new slab alloc flag, SLAB_ALLOC_NO_OBJ_EXT, to prevent allocation of obj_exts arrays, and let kmalloc_slab() override the type to KMALLOC_NO_OBJ_EXT when specified. Note that kmalloc_type() remains unchanged because kmalloc_flags() bypasses the kmalloc fastpath. Do not pass SLAB_ALLOC_NO_RECURSE to kmalloc_flags() in alloc_slab_obj_exts() and instead use SLAB_ALLOC_NO_OBJ_EXT only when the objects are allocated from normal kmalloc caches. While this prevents unbounded recursive allocation of obj_exts, it allows KMALLOC_NO_OBJ_EXT caches to have sheaves. Since sheaf allocations specify SLAB_ALLOC_NO_RECURSE that prevents allocation of both sheaves and obj_exts arrays, the recursion depth is bounded. obj_exts arrays for non-kmalloc-normal caches can now have a valid tag. Do not call mark_obj_codetag_empty() when freeing an obj_exts array to avoid false warnings. KMALLOC_NO_OBJ_EXT don't need this as they never allocate those arrays. Reported-by: Danielle Costantino Reported-by: Shakeel Butt Closes: https://lore.kernel.org/linux-mm/20260625230029.703750-1-shakeel.bu= tt@linux.dev [1] Fixes: 4b8736964640 ("mm/slab: add allocation accounting into slab allocati= on and free paths") Cc: stable@vger.kernel.org Link: https://lore.kernel.org/linux-mm/c5c4208d-a6f0-413e-bad9-49be12f12d55= @kernel.org [2] Signed-off-by: Harry Yoo (Oracle) Reviewed-by: Suren Baghdasaryan Link: https://patch.msgid.link/20260713-kmalloc-no-objext-v3-4-47c7bd138de7= @kernel.org Signed-off-by: Vlastimil Babka (SUSE) [harry@kernel.org: Backport notes: - Fix a minor conflict due to missing partitioned kmalloc caches in 7.1. - Use __GFP_NO_OBJ_EXT instead of SLAB_ALLOC_NO_OBJ_EXT since slab's internal alloc_flags do not exist in 7.1. - Since there is no way to distinguish between no obj_exts vs. no recursion in 7.1 due to lack of SLAB_ALLOC_* flags, sheaves for normal kmalloc caches are allocated from kmalloc-no-objext-* unlike upstream. Also, allocations from kmalloc-no-objext cannot allocate sheaves and may observe slightly higher contention. Only 7.1 has this problem as kmalloc caches don't have sheaves in LTS kernels. - Apply the __GFP_NO_OBJ_EXT flag to the !allow_spin path in alloc_slab_obj_exts(). ] Signed-off-by: Harry Yoo --- include/linux/slab.h | 6 +++ mm/slab.h | 28 +++++++++++++- mm/slab_common.c | 13 +++++++ mm/slub.c | 87 +++++++++++++++----------------------------- 4 files changed, 75 insertions(+), 59 deletions(-) diff --git a/include/linux/slab.h b/include/linux/slab.h index fda7e24a762b..739f474c5d92 100644 --- a/include/linux/slab.h +++ b/include/linux/slab.h @@ -642,6 +642,9 @@ enum kmalloc_cache_type { #endif #ifndef CONFIG_MEMCG KMALLOC_CGROUP =3D KMALLOC_NORMAL, +#endif +#ifndef CONFIG_SLAB_OBJ_EXT + KMALLOC_NO_OBJ_EXT =3D KMALLOC_NORMAL, #endif KMALLOC_RANDOM_START =3D KMALLOC_NORMAL, KMALLOC_RANDOM_END =3D KMALLOC_RANDOM_START + RANDOM_KMALLOC_CACHES_NR, @@ -655,6 +658,9 @@ enum kmalloc_cache_type { #endif #ifdef CONFIG_MEMCG KMALLOC_CGROUP, +#endif +#ifdef CONFIG_SLAB_OBJ_EXT + KMALLOC_NO_OBJ_EXT, #endif NR_KMALLOC_TYPES }; diff --git a/mm/slab.h b/mm/slab.h index bf2f87acf5e3..50414bd2b919 100644 --- a/mm/slab.h +++ b/mm/slab.h @@ -365,9 +365,13 @@ static inline struct kmem_cache * kmalloc_slab(size_t size, kmem_buckets *b, gfp_t flags, unsigned long call= er) { unsigned int index; + enum kmalloc_cache_type type =3D kmalloc_type(flags, caller); + + if (flags & __GFP_NO_OBJ_EXT) + type =3D KMALLOC_NO_OBJ_EXT; =20 if (!b) - b =3D &kmalloc_caches[kmalloc_type(flags, caller)]; + b =3D &kmalloc_caches[type]; if (size <=3D 192) index =3D kmalloc_size_index[size_index_elem(size)]; else @@ -402,7 +406,8 @@ static inline bool is_kmalloc_normal(struct kmem_cache = *s) { if (!is_kmalloc_cache(s)) return false; - return !(s->flags & (SLAB_CACHE_DMA|SLAB_ACCOUNT|SLAB_RECLAIM_ACCOUNT)); + + return !(s->flags & (SLAB_CACHE_DMA|SLAB_ACCOUNT|SLAB_RECLAIM_ACCOUNT|SLA= B_NO_OBJ_EXT)); } =20 bool __kfree_rcu_sheaf(struct kmem_cache *s, void *obj); @@ -505,6 +510,25 @@ static inline void metadata_access_disable(void) kasan_enable_current(); } =20 +/* + * Return true if KMALLOC_NORMAL caches may need obj_exts arrays. + * + * Memory allocation profiling requires obj_exts for all caches. + * Memcg usually doesn't need them for normal kmalloc caches, but kmalloc = types + * with a priority higher than KMALLOC_CGROUP can be aliased with KMALLOC_= NORMAL. + */ +static inline bool need_kmalloc_no_objext(void) +{ + if (!mem_alloc_profiling_permanently_disabled()) + return true; + + if (!mem_cgroup_kmem_disabled() && + (KMALLOC_NORMAL =3D=3D KMALLOC_RECLAIM)) + return true; + + return false; +} + #ifdef CONFIG_SLAB_OBJ_EXT =20 /* diff --git a/mm/slab_common.c b/mm/slab_common.c index 8b661fff5eed..214374233fb8 100644 --- a/mm/slab_common.c +++ b/mm/slab_common.c @@ -843,6 +843,12 @@ EXPORT_SYMBOL(kmalloc_size_roundup); #define KMALLOC_RANDOM_NAME(N, sz) #endif =20 +#ifdef CONFIG_SLAB_OBJ_EXT +#define KMALLOC_NO_OBJ_EXT_NAME(sz) .name[KMALLOC_NO_OBJ_EXT] =3D "kmalloc= -no-objext-" #sz, +#else +#define KMALLOC_NO_OBJ_EXT_NAME(sz) +#endif + #define INIT_KMALLOC_INFO(__size, __short_size) \ { \ .name[KMALLOC_NORMAL] =3D "kmalloc-" #__short_size, \ @@ -850,6 +856,7 @@ EXPORT_SYMBOL(kmalloc_size_roundup); KMALLOC_CGROUP_NAME(__short_size) \ KMALLOC_DMA_NAME(__short_size) \ KMALLOC_RANDOM_NAME(RANDOM_KMALLOC_CACHES_NR, __short_size) \ + KMALLOC_NO_OBJ_EXT_NAME(__short_size) \ .size =3D __size, \ } =20 @@ -957,6 +964,12 @@ new_kmalloc_cache(int idx, enum kmalloc_cache_type typ= e) return; } flags |=3D SLAB_ACCOUNT; + } else if (IS_ENABLED(CONFIG_SLAB_OBJ_EXT) && type =3D=3D KMALLOC_NO_OBJ_= EXT) { + if (!need_kmalloc_no_objext()) { + kmalloc_caches[type][idx] =3D kmalloc_caches[KMALLOC_NORMAL][idx]; + return; + } + flags |=3D SLAB_NO_OBJ_EXT | SLAB_NO_MERGE; } else if (IS_ENABLED(CONFIG_ZONE_DMA) && (type =3D=3D KMALLOC_DMA)) { flags |=3D SLAB_CACHE_DMA; } diff --git a/mm/slub.c b/mm/slub.c index 33e64f21ecc1..d83c3622ccde 100644 --- a/mm/slub.c +++ b/mm/slub.c @@ -2106,42 +2106,6 @@ static inline void init_slab_obj_exts(struct slab *s= lab) slab->obj_exts =3D 0; } =20 -/* - * Calculate the allocation size for slabobj_ext array. - * - * When memory allocation profiling is enabled, the obj_exts array - * could be allocated from the same slab cache it's being allocated for. - * This would prevent the slab from ever being freed because it would - * always contain at least one allocated object (its own obj_exts array). - * - * To avoid this, increase the allocation size when we detect the array - * may come from the same cache, forcing it to use a different cache. - */ -static inline size_t obj_exts_alloc_size(struct kmem_cache *s, - struct slab *slab, gfp_t gfp) -{ - size_t sz =3D sizeof(struct slabobj_ext) * slab->objects; - struct kmem_cache *obj_exts_cache; - - if (sz > KMALLOC_MAX_CACHE_SIZE) - return sz; - - if (!is_kmalloc_normal(s)) - return sz; - - obj_exts_cache =3D kmalloc_slab(sz, NULL, gfp, 0); - /* - * We can't simply compare s with obj_exts_cache, because random kmalloc - * caches have multiple caches per size, selected by caller address. - * Since caller address may differ between kmalloc_slab() and actual - * allocation, bump size when sizes are equal. - */ - if (s->object_size =3D=3D obj_exts_cache->object_size) - return obj_exts_cache->object_size + 1; - - return sz; -} - int alloc_slab_obj_exts(struct slab *slab, struct kmem_cache *s, gfp_t gfp, bool new_slab) { @@ -2150,13 +2114,17 @@ int alloc_slab_obj_exts(struct slab *slab, struct k= mem_cache *s, unsigned long new_exts; unsigned long old_exts; struct slabobj_ext *vec; - size_t sz; + size_t sz =3D sizeof(struct slabobj_ext) * slab->objects; =20 gfp &=3D ~OBJCGS_CLEAR_MASK; - /* Prevent recursive extension vector allocation */ - gfp |=3D __GFP_NO_OBJ_EXT; =20 - sz =3D obj_exts_alloc_size(s, slab, gfp); + /* + * In most cases, obj_exts arrays are allocated from normal kmalloc. + * However, normal kmalloc caches must allocate them from + * KMALLOC_NO_OBJ_EXT caches to prevent recursion. + */ + if (is_kmalloc_normal(s)) + gfp |=3D __GFP_NO_OBJ_EXT; =20 /* * Note that allow_spin may be false during early boot and its @@ -2165,10 +2133,11 @@ int alloc_slab_obj_exts(struct slab *slab, struct k= mem_cache *s, * very early allocations on those. */ if (unlikely(!allow_spin)) - vec =3D kmalloc_nolock(sz, __GFP_ZERO | __GFP_NO_OBJ_EXT, + vec =3D kmalloc_nolock(sz, __GFP_ZERO | (gfp & __GFP_NO_OBJ_EXT), slab_nid(slab)); else - vec =3D kmalloc_node(sz, gfp | __GFP_ZERO, slab_nid(slab)); + vec =3D kmalloc_node(sz, gfp | __GFP_ZERO, + slab_nid(slab)); =20 if (!vec) { /* @@ -2183,8 +2152,21 @@ int alloc_slab_obj_exts(struct slab *slab, struct km= em_cache *s, return -ENOMEM; } =20 - VM_WARN_ON_ONCE(virt_to_slab(vec) !=3D NULL && - virt_to_slab(vec)->slab_cache =3D=3D s); + if (IS_ENABLED(CONFIG_DEBUG_VM)) { + struct kmem_cache *exts_cache; + struct slab *exts_slab; + + exts_slab =3D virt_to_slab(vec); + if (exts_slab) { + /* + * The vector must be allocated from either normal or + * KMALLOC_NO_OBJ_EXT kmalloc caches to avoid cycles. + */ + exts_cache =3D exts_slab->slab_cache; + WARN_ON_ONCE(!is_kmalloc_normal(exts_cache) && + !(exts_cache->flags & SLAB_NO_OBJ_EXT)); + } + } =20 new_exts =3D (unsigned long)vec; #ifdef CONFIG_MEMCG @@ -2207,7 +2189,6 @@ int alloc_slab_obj_exts(struct slab *slab, struct kme= m_cache *s, * assign slabobj_exts in parallel. In this case the existing * objcg vector should be reused. */ - mark_obj_codetag_empty(vec); if (unlikely(!allow_spin)) kfree_nolock(vec); else @@ -2243,14 +2224,6 @@ static inline void free_slab_obj_exts(struct slab *s= lab, bool allow_spin) return; } =20 - /* - * obj_exts was created with __GFP_NO_OBJ_EXT flag, therefore its - * corresponding extension will be NULL. alloc_tag_sub() will throw a - * warning if slab has extensions but the extension of an object is - * NULL, therefore replace NULL with CODETAG_EMPTY to indicate that - * the extension for obj_exts is expected to be NULL. - */ - mark_obj_codetag_empty(obj_exts); if (allow_spin) kfree(obj_exts); else @@ -7906,10 +7879,10 @@ static int calculate_sizes(struct kmem_cache_args *= args, struct kmem_cache *s) s->allocflags |=3D __GFP_RECLAIMABLE; =20 /* - * For KMALLOC_NORMAL caches we enable sheaves later by - * bootstrap_kmalloc_sheaves() to avoid recursion + * For kmalloc caches we enable sheaves later by + * bootstrap_kmalloc_sheaves() to avoid recursion. */ - if (!is_kmalloc_normal(s)) + if (!is_kmalloc_cache(s)) s->sheaf_capacity =3D calculate_sheaf_capacity(s, args); =20 /* @@ -8475,7 +8448,7 @@ static void __init bootstrap_kmalloc_sheaves(void) { enum kmalloc_cache_type type; =20 - for (type =3D KMALLOC_NORMAL; type <=3D KMALLOC_RANDOM_END; type++) { + for (type =3D KMALLOC_NORMAL; type < NR_KMALLOC_TYPES; type++) { for (int idx =3D 0; idx < KMALLOC_SHIFT_HIGH + 1; idx++) { struct kmem_cache *s =3D kmalloc_caches[type][idx]; =20 --=20 2.53.0