From nobody Tue Sep 29 11:19:36 2026 Received: from m16.mail.163.com (m16.mail.163.com [220.197.31.2]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 057883B1B3; Sat, 8 Aug 2026 03:16:38 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=220.197.31.2 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786159004; cv=none; b=TtFDcKPGkp+Sfs64/QtrhGD6EAi/94fjATYVdtFEeg9AW/+oReuUL5aaJkIHJRRY9dEEDoIHDwbp3VyPEtCNq7j5VEdk1oCfz0XjhwXR5AoFFoxUaJsyqRfM9QLrZeMLi7r4S9/zG7jLDqzV9UvLOxcJLuZkdePGvN7LTJdcePY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786159004; c=relaxed/simple; bh=Ola9iQHMJnsFlT9upmL02ADrEDC3KX5FrrmNfmmz75g=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=X+wzLvkBAHT5mZ5dHsX1whHYV4Cj7YTuxNahYRO+dNRvgQeNiBLssqYz/xsGbNLIX3AFCrM3d+O1rQDxbtADxYQZj2al5+s8Sd4SCgSqSnVrsUDSCRsn+Jc4cXEZH0thD6gOynPM02Y0SzBzge9j53dtkaR+I9ZeITgEBgL+Hhk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com; spf=pass smtp.mailfrom=163.com; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b=d1wj/x2O; arc=none smtp.client-ip=220.197.31.2 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=163.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b="d1wj/x2O" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=163.com; s=s110527; h=From:To:Subject:Date:Message-ID:MIME-Version; bh=yF xoYkm6JojXNUVuKu6qzrezBLUlcTSSEZm/pGV4e6I=; b=d1wj/x2OXuvoqNF1TM tpseKZpWuR5WJcjlhcyDfL0FTyNiVFd8JDkWVBXHlfhlbPE9slIEt6s7uZs0HMCX +LNzByKybG9+L1g21mSXU/U38dgMWMqoHxqRmVjegMkR/WeeAXpHyck3nA4/7Maf VkHP5YAVQppnyASLVMX+rtPS0= Received: from localhost (unknown []) by gzsmtp5 (Coremail) with SMTP id QCgvCgC3QCBin3ZqScbJLA--.19934S2; Sat, 08 Aug 2026 11:15:47 +0800 (CST) From: Hui Su To: Andrey Ryabinin , Andrew Morton Cc: Alexander Potapenko , Andrey Konovalov , Dmitry Vyukov , Vincenzo Frascino , Zqiang , kasan-dev@googlegroups.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, stable@vger.kernel.org, Hui Su Subject: [PATCH v2] kasan: fix cache shrink race with CPU hotplug Date: Sat, 8 Aug 2026 11:14:59 +0800 Message-ID: <20260808031459.3032812-1-sh_def@163.com> X-Mailer: git-send-email 2.43.0 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: QCgvCgC3QCBin3ZqScbJLA--.19934S2 X-Coremail-Antispam: 1Uf129KBjvJXoWxurW3XF45WF1rtw48AF1DWrg_yoWrtw47pF y3Ka47Jr18Wr18tanrGa1UWryrAFZ8ZFnxXwsagw1FyF45u34DW34UKrW0vFWUWryrJa15 t3Wvya45WFWqyF7anT9S1TB71UUUUU7qnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDUYxBIdaVFxhVjvjDU0xZFpf9x0piHa0PUUUUU= X-CM-SenderInfo: xvkbvvri6rljoofrz/xtbC6gU5mGp2n2XpBAAA3F Content-Type: text/plain; charset="utf-8" kasan_quarantine_remove_cache() first invokes per_cpu_remove_cache() on all online CPUs. Each callback moves objects belonging to the cache from cpu_quarantine to the CPU's shrink_qlist, where they can later be freed from task context. kmem_cache_destroy() invokes the quarantine removal path while holding cpus_read_lock(), but kmem_cache_shrink() does not. The latter can therefore race with CPU offlining as follows: kmem_cache_shrink() CPU hotplug ------------------- ----------- on_each_cpu() CPU1 moves objects to CPU1's shrink_qlist on_each_cpu() returns CPU1 goes offline kasan_cpu_offline() drains cpu_quarantine leaves shrink_qlist untouched for_each_online_cpu() skips CPU1 The objects left on CPU1's shrink_qlist are not returned to the slab allocator. This may prevent kmem_cache_shrink() from releasing slabs that would otherwise become empty. If CPU1 remains offline, a later kmem_cache_destroy() also skips the list and can report that the cache still contains objects. An intermittent occurrence was observed with a virtio-9p filesystem. The mount and umount commands both returned 0, but the kernel logged the following during the userspace-triggered teardown: [ 2994.380134][ T111] BUG 9p-fcall-cache-1 (Tainted: G B = ): Objects remaining on __kmem_cache_shutdown() [ 2994.381140][ T111] Object 0xff11000004361118 @offset=3D4376 [ 2994.381607][ T111] Allocated in p9_fcall_init+0x201/0x400 age=3D1956= 4 cpu=3D1 pid=3D104 [ 2994.382591][ T111] p9_fcall_init+0x201/0x400 [ 2994.382810][ T111] p9_tag_alloc+0x12f/0x700 [ 2994.382982][ T111] p9_client_prepare_req+0x102/0x3e0 [ 2994.383165][ T111] p9_client_rpc+0x1ab/0xa50 [ 2994.383334][ T111] p9_client_getattr_dotl+0xb0/0x1a0 [ 2994.383515][ T111] v9fs_vfs_getattr_dotl+0x115/0x360 [ 2994.383719][ T111] vfs_getattr_nosec+0x22c/0x3a0 [ 2994.383910][ T111] vfs_statx+0xd7/0x170 [ 2994.384062][ T111] vfs_fstatat+0x45/0x80 [ 2994.384215][ T111] __do_sys_newfstatat+0x84/0xe0 [ 2994.384386][ T111] do_syscall_64+0x115/0x6a0 [ 2994.384566][ T111] entry_SYSCALL_64_after_hwframe+0x77/0x7f [ 2994.399720][ T111] WARNING: mm/slub.c:1244 at __kmem_cache_shutdown+= 0x363/0x500, CPU#0: busybox/111 [ 2994.405655][ T111] Call Trace: [ 2994.406325][ T111] kmem_cache_destroy+0x73/0x1b0 [ 2994.406630][ T111] p9_client_destroy+0x271/0x3c0 [ 2994.407210][ T111] v9fs_session_close+0x3c/0x260 [ 2994.407409][ T111] v9fs_kill_super+0x48/0x90 [ 2994.407584][ T111] deactivate_locked_super+0xa3/0x160 [ 2994.407778][ T111] cleanup_mnt+0x1dd/0x3e0 Thus, a successful umount left objects in the 9p fcall cache and prevented the cache from being destroyed cleanly. Per-CPU shrink_qlist storage exists for every possible CPU, and each list is protected by its own raw spinlock. Iterate over possible CPUs so that a list populated before its CPU went offline is drained as well. for_each_possible_cpu() can do more work than for_each_online_cpu(), but this change only affects CONFIG_KASAN_GENERIC kernels. The extra work is limited to cache shrink and cache destruction paths and does not affect the normal allocation/free fast path. It adds one raw-spinlock-protected scan of each possible CPU's shrink list. These lists are normally empty; a non-empty list is traversed to remove objects belonging to the cache being shrunk or destroyed. Fixes: 07d067e4f2ce ("kasan: fix sleeping function called from invalid cont= ext on RT kernel") Cc: stable@vger.kernel.org Signed-off-by: Hui Su Reviewed-by: Andrey Ryabinin --- Changes in v2: - Document the intermittent virtio-9p slab shutdown log and user-visible ef= fect. - Explain the runtime cost of scanning all possible CPUs. - No code changes. Link to v1: https://lore.kernel.org/r/20260806121006.1642946-1-sh_def@163.c= om mm/kasan/quarantine.c | 7 ++++++- 1 file changed, 6 insertions(+), 1 deletion(-) diff --git a/mm/kasan/quarantine.c b/mm/kasan/quarantine.c index 6958aa713c67..16f4e67beee8 100644 --- a/mm/kasan/quarantine.c +++ b/mm/kasan/quarantine.c @@ -355,7 +355,12 @@ void kasan_quarantine_remove_cache(struct kmem_cache *= cache) */ on_each_cpu(per_cpu_remove_cache, cache, 1); =20 - for_each_online_cpu(cpu) { + /* + * A CPU can go offline after on_each_cpu() returns, leaving cache + * objects on that CPU's shrink list. Scan all possible CPUs to + * drain those lists. + */ + for_each_possible_cpu(cpu) { sq =3D per_cpu_ptr(&shrink_qlist, cpu); raw_spin_lock_irqsave(&sq->lock, flags); qlist_move_cache(&sq->qlist, &to_free, cache); --=20 2.43.0