From nobody Sat Oct 3 03:45:27 2026 Received: from mail-pf1-f179.google.com (mail-pf1-f179.google.com [209.85.210.179]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D2DE3353A90 for ; Wed, 5 Aug 2026 14:11:56 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.179 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785939118; cv=none; b=F08e7NEkpsRBNWYl1Ba1n03nYqd4Iuk++umEcu7yS1kGeNrXUeBPHCDPt0xhu2LXZlNFw8f+sBKXK0wUESouEKotGeJHkvMAX4uRnHEutMNrhctlMqS3YPrASeBWAlCdiY8VFufRlHondTSZ2MKLa/1AM9KMNiNbNAAPtWKG2gg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785939118; c=relaxed/simple; bh=+jzbVjfcBjRIhuu80Gk/L3OMqK79L8nBqHPUawRt7Sg=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=RH4Geq9BsOYFfCY9g8DNDpCyX7X6E0jOa6AEoKbn6h//hPoDIlyGO0USBG/zdk/QyvKkNV7OofHWPdU+EC2TjZZOSaTaEZSCqI+IYdFKq02nsGby8ipzuNOKpkGmxz8g/3IXvxoNAE/9Ms5noePTe2Sb2y0k+WIZims8KUIHhh4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=jjzi80X7; arc=none smtp.client-ip=209.85.210.179 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="jjzi80X7" Received: by mail-pf1-f179.google.com with SMTP id d2e1a72fcca58-8485ef63b68so1549468b3a.1 for ; Wed, 05 Aug 2026 07:11:56 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1785939116; x=1786543916; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=BYzPnD7BiNKqZ5IgmbP05MNUjhzCdP6nAUYQEGSSwl4=; b=jjzi80X7/v7/TGGmebGUTS02x6/h8kC1fUelndfBx+dxK+bG09NH6K7Cxpp+XFGzrb 9VTDQfmTvcan3XSFeigGCCxWmBmPzIsuspbGclvel787vfxGHukBzYMsrbzvI9+UkISA KeWGso86gCgp1yHf+e11kHBWu7NLK+sAIOv6BcKCOFmiSkWZSBy1zl855t86KcmECKse cMxaOgvk+YzhwFW7qX/4QLLs2Ib/4RQutvKHa85xKG3yzKlCel7MAWgCXklLIr7zGpwA rFLKJFjCb1JgT8m6XlUzq9+95kVs+fs+njNLk+BViADoCcJhGuUb/gPzBZtKAdav1+J2 Tapg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785939116; x=1786543916; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=BYzPnD7BiNKqZ5IgmbP05MNUjhzCdP6nAUYQEGSSwl4=; b=qv2tRnXkzTkzFyhP8r2ddazQlwEb40voXEXVVboNfoAThH+eHu+nxyPt+GLkFV8vEX 1WZArRwqLmmwEVqutR4eYre+RtM4WiDcJHBgAUpzUxg5six0gg6rpDzxQVJRF/WUWeiO 3JXD+T6hraYwB53fT1u7B3lIycFMsw9zdV8oeDiEZYLLXvMl7tFLBKxFAANwSCxMZmCk XuosZfzM8oa3IiO8Lj+Wg3xeKGIdzSgshb4tQRQh/CTcPo61SmRG9FOoIzTnP5ndtROQ Fn0Toq7vnYgg2nfwvETczZXGQxfRVZg8EDZyFNYmg7I03TWuBgg0xxLOxNorHeka94QH 0b9Q== X-Forwarded-Encrypted: i=1; AHgh+RoqIlCQcyznouunnb4xoWmige9HPiSdYHXUElJ6ty/TydEtz2cTpBIUevanKk+HLq0KWaAeVzJ4HyiCsRA=@vger.kernel.org X-Gm-Message-State: AOJu0Yw/QPar2mfLMOQuEu8s4TQa6cu2NJKxJP8BBJICZBkXKdDZRUop VWldG0qPDxn+oKiM4hEVtJWN7ug8pI+9LZqru1/yi0ghvJ1kDBn5+ywl X-Gm-Gg: AR+sD11p88HTvfiJwy8BrgPVkoXpMwNiwGbtroS2JVa62tsm+ZSGDoacwx7oXiAAxDI G9S2XiprCG02IJf9utAEQLgp5zll941yIoJaXvJDngas/HYQNw874wYj4217b0nOy67i5zm2Szy c9Gkzwn6QWwNsbUpOL0nuVbe0vtHz3ch0piiyx94LW/hj91m+J6xgRfxQYI/R99cJFqFINKnk4u 44TFveDUlY4mpCP4Qbj+BKzyp3BNPMvmgc/IrbxJflnNTtJT0tlRC3KasYycXn+DBRxWA1xMNFp ZHiOTMGm9DtPb6HI1bDf9HQf/YHWircBLj1I+hLiYtEdy6KoP9R9ELVaaketPtEfCOvDzJLuAkn tDQKDi2fHE7jSBsCXjUMGjbgIjTHJYSPvj4ORLL/Z4CgU79RVL/v/Xr6360lGtIBYvvWth2hcSA h5TGVzeWteRSZvEDylykweZvx0ntBilqZulSn1bxrGFluo8KSisQuxXzIgJBIpwtL1qEHn4dz9B rrfA8J8Y2s398u/f7afZeTnHry1aiaSnGt/uztqbHFSr8xaEQ== X-Received: by 2002:a05:6a00:c86:b0:845:e440:d0ca with SMTP id d2e1a72fcca58-84f2dfc8b86mr6689864b3a.8.1785939116040; Wed, 05 Aug 2026 07:11:56 -0700 (PDT) Received: from localhost.localdomain ([220.85.166.190]) by smtp.gmail.com with ESMTPSA id d2e1a72fcca58-84f2e50846esm913940b3a.46.2026.08.05.07.11.52 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 05 Aug 2026 07:11:55 -0700 (PDT) From: Youngjun Park X-Google-Original-From: Youngjun Park To: Andrew Morton Cc: Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , her0gyugyu@gmail.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org Subject: [PATCH v2 1/2] mm/swap: fix stale comment on swap_info_struct::cluster_info Date: Wed, 5 Aug 2026 23:11:45 +0900 Message-ID: <20260805141146.127776-2-youngjun.park@lge.com> X-Mailer: git-send-email 2.48.1 In-Reply-To: <20260805141146.127776-1-youngjun.park@lge.com> References: <20260805141146.127776-1-youngjun.park@lge.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" setup_swap_clusters_info() allocates cluster_info for every swap area, not only for SSDs. Signed-off-by: Youngjun Park Acked-by: Kairui Song Reviewed-by: Barry Song --- include/linux/swap.h | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/include/linux/swap.h b/include/linux/swap.h index 2cb1d29307c5..2b14e2e9673b 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -246,7 +246,7 @@ struct swap_info_struct { struct plist_node list; /* entry in swap_active_head */ signed char type; /* strange name for an index */ unsigned int max; /* size of this swap device */ - struct swap_cluster_info *cluster_info; /* cluster info. Only for SSD */ + struct swap_cluster_info *cluster_info; /* array, one entry per cluster */ struct list_head free_clusters; /* free clusters list */ struct list_head full_clusters; /* full clusters list */ struct list_head nonfull_clusters[SWAP_NR_ORDERS]; base-commit: 0b53bff4fa05ff0d3ffbd3d3bb10fae69dfab498 --=20 2.48.1 From nobody Sat Oct 3 03:45:27 2026 Received: from mail-pf1-f171.google.com (mail-pf1-f171.google.com [209.85.210.171]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6126D34D398 for ; Wed, 5 Aug 2026 14:12:00 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.171 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785939121; cv=none; b=UzzUjaZZsjzSCUX3pVBTVD+ViBC9ALypLMLX7T5yY8HCguGl9gj4XqvnkqlKainqAQhVgwKvJBKQl4qDA6OuH/wVjNt2O0M6mQ/Rxcpb6/7Ttyw8eY3jBDNea0BP3op3oTqBQa2xxW8PZhubS3KDCd5f4A+JVG6kpD7kNH2bM6Q= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785939121; c=relaxed/simple; bh=KDy18u4ZvBz6clEFGIIOXq3IPMoYRfbEIk9Rjzlq5B4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=KXQqU6nkUsYpmopbp9lKjybIhK6Lo6Lg+isRiduLcM0R6jZdA8yACAOrF//wFofh2bdhrXb+P59vYl1Amxod8u4j5IiWBgzVmoeJz492K1CSZtOEADUfsaJbA4MkXM/s+Yi1zJmSYTtN6+c2ntUB0Xv3viBUVRFgOmwG5qPoc4g= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=FewkM3wG; arc=none smtp.client-ip=209.85.210.171 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="FewkM3wG" Received: by mail-pf1-f171.google.com with SMTP id d2e1a72fcca58-84a2dcede83so1519466b3a.3 for ; Wed, 05 Aug 2026 07:12:00 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1785939120; x=1786543920; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=nSVkiHxOCVF8oGGFv8KCXv/ezUKTNu+0csWqVWaxMlc=; b=FewkM3wGs8D2VjnaEmeEvWjZnUlyROLQq5JfSYh+wL2ZINITTU62y8N97vSIOCbS3F LLgvLkJcB7juT3bKUmTprQnhzOAv7mpBMHFJNyShK1ujijqbiro/x2t/RnDCPlgVaYlT 3yDg2YiMy8oER7meIfhZBssIkg3OYC0x8vVYmae6zxANnIOKWEJaB09k8w5aNxWP02jE qRNbXt5hdc0Ce9dCRF/DZsj0GmgBEc9OkAWpFFUpcaKVSJMrfP6NfYmO7yhFzpVqCxFY VfNPct8NfPZkaUMkY4Z7PkxVUZQ92DqoDY1XREekFxdS2WDVNbxKl6RXenOYdAheoIBg 9slQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785939120; x=1786543920; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=nSVkiHxOCVF8oGGFv8KCXv/ezUKTNu+0csWqVWaxMlc=; b=KO/AwiGGwkV2naT56fPtrT+TdC26xXQvJuEXJxGvq9JvKQg8xu1z9+2p64PI9d/s+d dn/BC9W20MRubg/4MeMAXb+FdXjxb/qrC1rtZaSAAetYu+kB5Yw9qKaIDD5OwM5aIKN1 lS9vlg0SEQUhk7jXRnPP10eGuhqO6gBU7C11Bo5ah1qesDxf4zD2fl85F6fxBjLqlp65 48WJV3dvNAXVPkk8joorUceAktMWuqr4OQ4/qdceULtjE+MFAZ6rWFnUJ4sdEB/10FTo PYymThgQAjKlmbrJnMf7JLHya+uR4/2wRQNywLl4bwKzNQ5tdkeREAOHcb+if8cMhyi3 ylmQ== X-Forwarded-Encrypted: i=1; AHgh+Rrr1ikrH+gNzg40pVNO3LfYTqkkEQnQs/3asXjvDmuol1Jzbz+srFMURalA7USTHxU+IDfAXdI3ZNcC6Gw=@vger.kernel.org X-Gm-Message-State: AOJu0YxEJqaDE90Lw5ABht1C3ekYeUd+h6pG4ZufCGn2VjBL9mW8OqLg Mew/0jBeyyLo594cSPZYGL9kmW5mar32YqcSWbbkWujOza1h+ADoFNM2 X-Gm-Gg: AR+sD12Ir6QGZsMoFMmIZ5Kgzkcm05h2uvicYqSzc1WECg2ZXusILN11FmrtSmwrgJ7 hRDOrgGwlS7eOLACvJz/TLojId/BLMBEm6LJVs+/j1hqHwsDy7IlK0DAkzqB7nk1l2aecHzkMwz /vfm3E2JK9C3oYUbQ3/fPPacRM9VtzCpLRpRsfhGEu3VFcGLeUREf7YIewK5Wdnmz6d4Uxn5VTB +5J/Ag23tZ2s3xnh32MKf6WLmTeGCQ8/UICqvcn6lozHcQRdoEisdPN0ngbOh4C/quw1DIMFvoT FtVNVht4taoRHv3pymywskKouE3rkY7+XrCo2eJ8GcB5oD1KkOz3JMIOf1ccO614EnXRFNPvKzZ WyQrJlP++1AyE77tjHGJSnPg+4noqbOqYGvkQGHcIWHfbXBM8xn3aout1tH9IqcSayeJ3KOUXE2 /3xdyPdVOc/3uZGxj2mwCOSiu2tDpikH1oxquxrZsY9AZw0UEzDuyqCBSQrBrqsTZ0oTFfTwoWi VqsV3j1Gp0ZBJk/cGdyPRTWt3+5HUXVGR7IB01qFp3lSKVXeA== X-Received: by 2002:a05:6a00:3996:b0:84e:5e4:b9a2 with SMTP id d2e1a72fcca58-84f2e0edf09mr7646388b3a.36.1785939119798; Wed, 05 Aug 2026 07:11:59 -0700 (PDT) Received: from localhost.localdomain ([220.85.166.190]) by smtp.gmail.com with ESMTPSA id d2e1a72fcca58-84f2e50846esm913940b3a.46.2026.08.05.07.11.56 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 05 Aug 2026 07:11:59 -0700 (PDT) From: Youngjun Park X-Google-Original-From: Youngjun Park To: Andrew Morton Cc: Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , her0gyugyu@gmail.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org Subject: [PATCH v2 2/2] mm/swap: scan by cluster in find_next_to_unuse() Date: Wed, 5 Aug 2026 23:11:46 +0900 Message-ID: <20260805141146.127776-3-youngjun.park@lge.com> X-Mailer: git-send-email 2.48.1 In-Reply-To: <20260805141146.127776-1-youngjun.park@lge.com> References: <20260805141146.127776-1-youngjun.park@lge.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" find_next_to_unuse() walks every offset from 0 to si->max, and swapoff restarts that walk on each retry, so the cost scales with the size of the device rather than with the few slots the shmem and mmlist passes could not free. It has caused stalls before. The flat walk predates the swap table. Slot state now lives in a per cluster table, and wait_for_allocation() stops all allocation before try_to_unuse() runs, so a cluster that holds no slot in use stays that way. Skip such a cluster instead of reading all of its entries. Commit dc644a073769 ("mm: add three more cond_resched() in swapoff") answered those stalls with a cond_resched() every 256 offsets. A walk bounded by one cluster no longer needs that counter. The loop now runs at most SWAPFILE_CLUSTER times before it returns or reschedules, the same bound swap_reclaim_full_clusters() already scans between cond_resched() calls. The inner loop runs to the end of the cluster rather than to si->max. The swap table is always SWAPFILE_CLUSTER entries and swapon() masks [si->max, round_up(si->max, SWAPFILE_CLUSTER)) as bad, so the tail of a partial last cluster is rejected by swp_tb_is_bad() and never returned. ci->count is read without ci->lock, so READ_ONCE() marks the read for KCSAN. Allocation is already stopped, so the count can only drop, and a slot stops being counted only after its folio has left the swap cache. An empty cluster therefore holds nothing for try_to_unuse() to act on. Signed-off-by: Youngjun Park Reviewed-by: Barry Song --- mm/swapfile.c | 41 ++++++++++++++++++++++++++++------------- 1 file changed, 28 insertions(+), 13 deletions(-) diff --git a/mm/swapfile.c b/mm/swapfile.c index dea2d3b36e06..36e4e8884b76 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -370,8 +370,6 @@ static void discard_swap_cluster(struct swap_info_struc= t *si, } } =20 -#define LATENCY_LIMIT 256 - static inline bool cluster_is_empty(struct swap_cluster_info *info) { return info->count =3D=3D 0; @@ -2763,7 +2761,9 @@ static int unuse_mm(struct mm_struct *mm, unsigned in= t type) static unsigned int find_next_to_unuse(struct swap_info_struct *si, unsigned int prev) { - unsigned int i; + struct swap_cluster_info *ci; + unsigned long i, end; + unsigned int ci_off; unsigned long swp_tb; =20 /* @@ -2772,19 +2772,34 @@ static unsigned int find_next_to_unuse(struct swap_= info_struct *si, * hits are okay, and sys_swapoff() has already prevented new * allocations from this area (while holding swap_lock). */ - for (i =3D prev + 1; i < si->max; i++) { - swp_tb =3D swap_table_get(__swap_offset_to_cluster(si, i), - i % SWAPFILE_CLUSTER); - if (!swp_tb_is_null(swp_tb) && !swp_tb_is_bad(swp_tb)) - break; - if ((i % LATENCY_LIMIT) =3D=3D 0) + i =3D prev + 1; + while (i < si->max) { + ci =3D __swap_offset_to_cluster(si, i); + ci_off =3D i % SWAPFILE_CLUSTER; + end =3D i - ci_off + SWAPFILE_CLUSTER; + + /* + * An empty cluster has no slot in use, so skip it whole. + * A slot is uncounted only after its folio left the swap + * cache, so there is nothing here for try_to_unuse() to act on. + * Count only drops here, so a READ_ONCE() without ci->lock is + * enough, unlike in every other cluster_is_empty() caller. + */ + if (!READ_ONCE(ci->count)) { + i =3D end; cond_resched(); - } + continue; + } =20 - if (i =3D=3D si->max) - i =3D 0; + for (; i < end; ci_off++, i++) { + swp_tb =3D swap_table_get(ci, ci_off); + if (!swp_tb_is_null(swp_tb) && !swp_tb_is_bad(swp_tb)) + return i; + } + cond_resched(); + } =20 - return i; + return 0; } =20 static int try_to_unuse(unsigned int type) --=20 2.48.1