From nobody Sat Sep 26 22:02:19 2026 Received: from mail-pf1-f178.google.com (mail-pf1-f178.google.com [209.85.210.178]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CEA633264C1 for ; Sat, 29 Aug 2026 07:46:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.178 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787989620; cv=none; b=CcZAcu2xcGKE2PybUGCCOXHfaSAELHH5dYBcAJ/bkHAjQPyNPCvdA+zB8XINOtHPFODOklENlF2DUXs4mfnMOhK9VxeRKlFv5s2j6JcqBlDc/6tTnrUgPIOwVnmfcT1/3jfwk9FdaNEMlJR3XVm2JbNnBr4CAf77eQkkemsdIHY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787989620; c=relaxed/simple; bh=xHcpGAf/1y+7HzUaEjIKyAZ/SYP2aVkR+0WjUUOVHak=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=f7ALZbpTZnKGI5aOcPey+DipTIoUhOUT2JDMUyq66aD21yNFVQObBWJUmqPS7tVf3g+RmMe0C+tZYvfwdvl8lOEdk2mniVkdFpwL2NFHUaNQE5pBx6s9wWxS1N+GDtypTNE/6WL/vfxl0sYPGVrBiW6CUjLDkJf0vR9T1ODXlYQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=BlQ7cl3t; arc=none smtp.client-ip=209.85.210.178 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="BlQ7cl3t" Received: by mail-pf1-f178.google.com with SMTP id d2e1a72fcca58-8518b3ff3e9so1901019b3a.2 for ; Sat, 29 Aug 2026 00:46:58 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787989618; x=1788594418; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=UPBa7pHbP53WE2M6DUPslMFnjHFqwGfpa+1GCbbSTEE=; b=BlQ7cl3tZDEE/m/md//p+G1JPPyJ96+hCWyxHFEq34rkVPhUq9Vfd39+XPAAs/Lm8V iuHrOK9iMnUrNYYX+He/6KqZWY9DXhXwtRPKeWotx6/4B4iLImU5d18lVimWV3giU+dk AoncfZSg4B5p7zq1vguWqPjIk1wSUL4z+zapzLD7uIf/QmdTYYdcKW+HMgooIBMj4f2y K3HR0IAs9TgSRLoiVfYMfKxyaIFiJNPi70TGFMj8wLxXs+WLqngKP6oMr9CPW6GRfs3J 1ybuehJQpuuM0xVJ8bdp96YsnpycHZe7d9xwdhdeUNn5qBIXeOlaizoWbNgcvmDez6P/ pJNA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787989618; x=1788594418; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=UPBa7pHbP53WE2M6DUPslMFnjHFqwGfpa+1GCbbSTEE=; b=gdNf2LfuG7BTfo4yXZ/KJ9RSU6YQcgxxA+XNMJf61OA92xsGuvbsmChnAIBklQQltE njE4LKQzEPuwogQfMYkpk87eGtDcxOQyMbjssubiiK/eGbGBjOuqjMwTVHZ8Uwa5xmys Ps7zr+AlW1mgRZvnEpOR5odi//7y5RWcEUJ/xjy1fV/ADtPVYwjJgrrtkModaW8Z8cSh dVqznTr3bdQSg/IgvJ/Xb48lvjHGMErhc262HRB1Y/R1sa3MWCpj5EVnTitfkJV/Txnh 6GF7y4mLdnSrfKXO3OJXYn8Cu0Du9ZmLXygGYray+Ph0IxYrmQfEsixe6xGFUvoo8/L+ sUhA== X-Gm-Message-State: AFuF++khFrZPrwxeQV9kmGJzMMcAkxdF7aVAB6gYjD82fSzdLNkBIETw teDDUfTXc9W6tK/aHLbyw+rXFLKL171TIIh0Whtd6EnsIqF1xEVLLdVL X-Gm-Gg: AR+sD11hkjVrf+gsTawiGFASAiN3WZ5303BNBwDbKBKkdAo6B4NULGU+zB1Ua2v6llv e836OMWZ7/eyFy55U/dQV4Yvyuz0X3WpXfRoqAUDi/C9rms1UvWCbpengQQn9ebiYJ38zUj85vC RG8FSrHmzj1lo99qBiqt4Mnk4USphthX0+e5qHJ41x1j5+jWpDixOVaCUkt7mGMWLJ7Wa6UmrFz qJ0G8C33619s+IBnjVpNWIM1uFsVrlkiN1pFhxjqTJmLflclwX8cBfkrP5iGD1ZFN/bXXZ6rfXS kKhCb45gRgTSBHJiI/gan61jivzGoodKKoXL3Y09Sv9hwNspExN7HJUKC0VFpa0s+N+aawo5CKN A+SZowoZJqf5v7w/MIKpPhtAsT9GCJ0hhuwu4ajBBmwxounUJwlLFw8pr8Y8hAT3XxXLTyhHIAP eKlYCGLYzbBbNGKmmF1Zx4tiEQeSgSYnWO4+eythkLHI9w/PZlwi9KFuYGlHbbNPaSvJjGUDLME TXPOflqji/+I8IBo9j+zy4+l17CPNI0+LXwprk= X-Received: by 2002:a05:6a00:7607:b0:857:72ba:ff12 with SMTP id d2e1a72fcca58-85772baffdfmr7291395b3a.26.1787989618074; Sat, 29 Aug 2026 00:46:58 -0700 (PDT) Received: from localhost.localdomain (vmi2317720.contaboserver.net. [84.247.152.65]) by smtp.gmail.com with ESMTPSA id d2e1a72fcca58-8569fa79d0asm1310439b3a.14.2026.08.29.00.46.53 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Sat, 29 Aug 2026 00:46:57 -0700 (PDT) From: "Lian Wang (ProcessMission)" To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, Andrew Morton , Chris Li , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , Kemeng Shi , Kairui Song , Jihan LIN , Kunwu Chan Subject: [RESEND RFC PATCH v2 01/13] mm/swap: remove unused parameter for reading swap header Date: Sat, 29 Aug 2026 15:46:49 +0800 Message-ID: <20260829-swap-pcp-priq-v2-resend-1-68d3d925578c@gmail.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260829-swap-pcp-priq-v2-resend-0-68d3d925578c@gmail.com> References: <20260829-swap-pcp-priq-v2-resend-0-68d3d925578c@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Kairui Song No feature change, just a minor cleanup. Signed-off-by: Kairui Song Signed-off-by: Lian Wang (ProcessMission) Tested-by: Kunwu Chan Tested-by: Kunwu Chan --- mm/swapfile.c | 7 +++---- 1 file changed, 3 insertions(+), 4 deletions(-) diff --git a/mm/swapfile.c b/mm/swapfile.c index 53bf01d5f7f1..0f962cdfa5c0 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -3467,9 +3467,8 @@ __weak unsigned long arch_max_swapfile_size(void) return generic_max_swapfile_size(); } =20 -static unsigned long read_swap_header(struct swap_info_struct *si, - union swap_header *swap_header, - struct inode *inode) +static unsigned long read_swap_header(union swap_header *swap_header, + struct inode *inode) { int i; unsigned long maxpages; @@ -3700,7 +3699,7 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialf= ile, int, swap_flags) } swap_header =3D kmap_local_folio(folio, 0); =20 - maxpages =3D read_swap_header(si, swap_header, inode); + maxpages =3D read_swap_header(swap_header, inode); if (unlikely(!maxpages)) { error =3D -EINVAL; goto bad_swap_unlock_inode; base-commit: aeddb4d52acfcc5ce5e988acd48f2906fe966ca3 --=20 2.55.0 From nobody Sat Sep 26 22:02:19 2026 Received: from mail-pl1-f179.google.com (mail-pl1-f179.google.com [209.85.214.179]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2D77A3264C1 for ; Sat, 29 Aug 2026 07:47:07 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.179 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787989629; cv=none; b=afNagiAzL/FbN3Cnpf+jruqkbwV7fPif5RuqlO3+8qpyVXww+WgjmxlUr2y+c0knHp3fG3kH0RPxT6QqXea90zrHfLv7/ZvYh4Y/kMuhXwkvcTrJOV4y7FGhESqghs+x6iUK5flaftxuXf3UR2r5kZ1CgLHCQaMve///0abuKQk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787989629; c=relaxed/simple; bh=F+k64I/whcdTvBpi75E6T6u1RKpAOmS/TI4FGXuJfvI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=WZZqviNshe2VcsknMbmgNmdK4GY7jOXk8oYgyBgwSJh7NXOXACSU/pujEmdGFBz88WFeIGKV1yUBUBzvscJiM+wKSfWZO3+qSb24PKCIICiUrtmB47F6UvVvwWRe6XfGH3XDvp3XRpc02huN2+YjTpFiqE/tDgu2/WZXVa8Gf7M= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=IeGaky8L; arc=none smtp.client-ip=209.85.214.179 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="IeGaky8L" Received: by mail-pl1-f179.google.com with SMTP id d9443c01a7336-2d6efd73032so24756695ad.0 for ; Sat, 29 Aug 2026 00:47:07 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787989627; x=1788594427; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=ZlfOvgBEtySKIOUVBBOTKpFAbjQZaUCTZO6eqUdoNDE=; b=IeGaky8Lo/o8ALrS0LKBlKF5uUWMn3pZkRjKtfd2l00d0WR8yPcafWCLievg5+zj7Z mdm4VlgBpd3DciIUfgtcB1F8ImFLhnIcxxEsM68II4jrbmP02DXKyx2ybcOPdCEKN3PW KMpSmaEjdFx6caJYwfHjCeV0mJk9HrFqErP0WTDOaeQ3qPHttLEIvEhdGpd3ug77KJKx PuMJavbY9zK2oo81riXmF0E9Irx6h2lSZ9/V0U/ah8pNVXKgzW1OdIM1KpCkPcHmZkvc ubWkvdXSCCs6mqvAFD6Fhf2yTXO4TEm57dMhFQ1MP+UjpcDpa2FY5i0uUBHrH82dP112 c4Gg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787989627; x=1788594427; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=ZlfOvgBEtySKIOUVBBOTKpFAbjQZaUCTZO6eqUdoNDE=; b=NlZD3M09jNrlrIcUDt/VEOEBnWxKGkAEBWDc35siFp7tNdo0T3CJBTEuvJmT/FJHqV de55Cdc+32/Wgf43BJHqC/R00EeGawCcV8fiHzMu6XJgjzA60W8/kWEl2jWNVi5WL0HN /O1xf/rliKLG9bQgsPpfnT8qd07KbR1k9b17mAT6gcpWWFYZRfd6wXvwJdau5aVZQh/Y BRCerdfFfLUpecDg4dQ/13aDzaLAVp5o6tK/kWYjoQE55Uu0E3NidvB3o96nMZEEq4kN H52/A1ZXGopg1JGkbXra50I4D/KYtiOtQ02g04UCXhArk6nZlOjSyiA1e4Unre8gRVOr k9AA== X-Gm-Message-State: AFuF++l1o1jeol9Jy6R9giz7iH1daD+wApx6maLq8KzhnfL2VbT/nYwb AFBV9cbsKC1/A9qX8ZlZ5gEQcJAkgAngaoWUrp474vVKGUblooE93V1M X-Gm-Gg: AYBFou0C6QOzH8SfGSzIE2rTSvQIzNU2YbpKxv0wNkQhHX51BC+Hf+zaxeugzatkpKj 08eVTDiU6M9O8vVUZOmnuxjJEN3rCNHPsdcqsnAeSngORbMRo62nAY6AfO5d/oBWDBZrt/l9tIh TcR+6J7+H+g8Yltr4jrcPT7jhnFw9KSMyhnb8rl/wgi7xclCAsQamyc47SKNbiw1MZCoNyFvWk2 YcfN/ZboLWJMkjaFBP8DLbUsTY1xCzbC7PJAbP9bfETjxyUjobZse8llcNq1LlJSLMWk31uYr9x 2NpMxELJv4nttG+CuZdRDZFsTF8CTpbl9gZmvQX+owvupDJrd8ItAkq0X4iQY+EilzfBR0BdDQQ Jn4/WVl0QV+bJqYKFr4q/spAyoDhrHz4WBrlPERoVv8h9a2N6DPRT4WLa81kj37GqaLjeKdUmOm hGIWqV5vMTS4L31PQR6mf4k2V3k0elSurKdE49xDwZh+Hsj0sFOi3XnQUyfCw3f/0Fu71ztOc67 44ghM48dJKgAxGs5D8ZnxKmivol X-Received: by 2002:a17:90b:2243:b0:384:927f:3db9 with SMTP id 98e67ed59e1d1-3989b3097b0mr3060892a91.1.1787989627434; Sat, 29 Aug 2026 00:47:07 -0700 (PDT) Received: from localhost.localdomain (vmi2317720.contaboserver.net. [84.247.152.65]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-396b1992435sm11159127a91.13.2026.08.29.00.47.02 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Sat, 29 Aug 2026 00:47:07 -0700 (PDT) From: "Lian Wang (ProcessMission)" To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, Andrew Morton , Chris Li , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , Kemeng Shi , Kairui Song , Jihan LIN , Kunwu Chan Subject: [RESEND RFC PATCH v2 02/13] mm/swap: slightly cleanup the code for hibernation error handling Date: Sat, 29 Aug 2026 15:46:58 +0800 Message-ID: <20260829-swap-pcp-priq-v2-resend-2-68d3d925578c@gmail.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260829-swap-pcp-priq-v2-resend-0-68d3d925578c@gmail.com> References: <20260829-swap-pcp-priq-v2-resend-0-68d3d925578c@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Kairui Song Restructure the error paths in pin_hibernation_swap_type() using a goto out pattern to avoid repeated spin_unlock() calls and simplify the return flow. Also simplify unpin_hibernation_swap_type() by making the si NULL check inline, and clean up find_first_swap() to avoid an early return inside the loop. Signed-off-by: Kairui Song Signed-off-by: Lian Wang (ProcessMission) Tested-by: Kunwu Chan Tested-by: Kunwu Chan --- mm/swapfile.c | 39 +++++++++++++++++---------------------- 1 file changed, 17 insertions(+), 22 deletions(-) diff --git a/mm/swapfile.c b/mm/swapfile.c index 0f962cdfa5c0..46772d0e3e68 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -2260,21 +2260,18 @@ static int __find_hibernation_swap_type(dev_t devic= e, sector_t offset) */ int pin_hibernation_swap_type(dev_t device, sector_t offset) { - int type; + int ret; struct swap_info_struct *si; =20 spin_lock(&swap_lock); + ret =3D __find_hibernation_swap_type(device, offset); + if (ret < 0) + goto out; =20 - type =3D __find_hibernation_swap_type(device, offset); - if (type < 0) { - spin_unlock(&swap_lock); - return type; - } - - si =3D swap_type_to_info(type); + si =3D swap_type_to_info(ret); if (WARN_ON_ONCE(!si)) { - spin_unlock(&swap_lock); - return -ENODEV; + ret =3D -ENODEV; + goto out; } =20 /* @@ -2283,14 +2280,15 @@ int pin_hibernation_swap_type(dev_t device, sector_= t offset) * the same session. */ if (WARN_ON_ONCE(si->flags & SWP_HIBERNATION)) { - spin_unlock(&swap_lock); - return -EBUSY; + ret =3D -EBUSY; + goto out; } =20 si->flags |=3D SWP_HIBERNATION; =20 +out: spin_unlock(&swap_lock); - return type; + return ret; } =20 /** @@ -2309,11 +2307,8 @@ void unpin_hibernation_swap_type(int type) =20 spin_lock(&swap_lock); si =3D swap_type_to_info(type); - if (!si) { - spin_unlock(&swap_lock); - return; - } - si->flags &=3D ~SWP_HIBERNATION; + if (si) + si->flags &=3D ~SWP_HIBERNATION; spin_unlock(&swap_lock); } =20 @@ -2348,7 +2343,7 @@ int find_hibernation_swap_type(dev_t device, sector_t= offset) =20 int find_first_swap(dev_t *device) { - int type; + int type, ret =3D -ENODEV; =20 spin_lock(&swap_lock); for (type =3D 0; type < nr_swapfiles; type++) { @@ -2357,11 +2352,11 @@ int find_first_swap(dev_t *device) if (!(sis->flags & SWP_WRITEOK)) continue; *device =3D sis->bdev->bd_dev; - spin_unlock(&swap_lock); - return type; + ret =3D type; + break; } spin_unlock(&swap_lock); - return -ENODEV; + return ret; } =20 /* --=20 2.55.0 From nobody Sat Sep 26 22:02:19 2026 Received: from mail-pj1-f48.google.com (mail-pj1-f48.google.com [209.85.216.48]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 57FBD356755 for ; Sat, 29 Aug 2026 07:47:17 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.48 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787989639; cv=none; b=SRZoVCVCgCfNyk6WQFKIJyUvaFOrDcrqgWmXa4ni8Z/pXbpRKU2XYnZ/JSWdzDO6K1kKlAXM1OhcQFuP1yp/3nITX5mD4DAzJ/FjeXQYO/CzlCYxksfRkQ0xJrax1A92ttHNuuwbfFAV3appH+G45yG/QbTO/mjKFyYVkm4L95Q= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787989639; c=relaxed/simple; bh=r3AtZeFgwL645pkYtEc2ViBr/e/CD7Jo8bgjnq80Rd0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=k2RPScQQ7v4Gz0ddDS3k54l/PxDInYxl/eO8bddVptI5raJ/rjnGbj2McDYV0ot7/VLEPj2dRGZgzRfs4stUX45h7FkWxgQ24LLJW4kqVeE4gE67CYVI9R3ThIAdEKKrT+7sf2uLxwakpcil/3AN7m/rOZtGIwxcFwQT5Yf+YUg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=aDQvIw2l; arc=none smtp.client-ip=209.85.216.48 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="aDQvIw2l" Received: by mail-pj1-f48.google.com with SMTP id 98e67ed59e1d1-3965d3d9ab8so1340779a91.3 for ; Sat, 29 Aug 2026 00:47:17 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787989637; x=1788594437; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=wi+OGZ51DE6sVJ1HATvP7JiivAEZJcFvS/FmT0stnCw=; b=aDQvIw2lzF7/1Y7VeVXC+46Ksggw+KS6ZOtNJ0vNMDtyqqRDDInd9+MVA/DdzEhbqm VoE6u4LiTn+fCpyvowE4GScB1SfluFOGYJsRirdKjk9iIpdCmR7aQe0jv4TQPrUGGHRR YonnJyeZTuxsyrRc/9JjRx51yuFWddTmDCSrNipM0xabX39t079PODGMxmWbUccljCm0 TkrxYYmKyIjNT6eJPwCHzom0eaq0Mon8tvSkV3OHw/oZXcEVq+QxS4zTOzpXdB2MI8kR WfvieS+ZjhZQbkASgZJuffOq5s2g5wjNFkj55ctPI2QRI6YO3XoCa/MXLvdycsjzct5c 1U2Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787989637; x=1788594437; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=wi+OGZ51DE6sVJ1HATvP7JiivAEZJcFvS/FmT0stnCw=; b=DZLzIjRXlO1LOsPQdgTq6t64omYRBLIYnhdSRHhGTnb1kIMaVAEUyRFbQT7mKynjgt dPVpWNnuXFYK4F/oPe/uFqDQ0AdlpSKwPN0E4b+1ViUg4MjEpsW8feCT8VspT7oD2viz 1n9Ki9GTCbZpUJCXS2uzav9JUX0kNXAs1MD7JA5Xq0514DFIcWp2oSzBzuLLV2R6Wz5F GYZvr21HNnKjZpbE4GHiyMiaelhGgnHkNSelzuS0mUJwbhlaiqUUqytUO6lsb2vGwez5 xfHIE0HLcDO7u2aE+o+V9ALETDHdICN80EkguZDn8Duav5kH5s0HENj/aCbNQpyBqnG2 Es+g== X-Gm-Message-State: AFuF++mKBYYDWcigqcqkt41V8d97KcNPnbmHfdGwnGrN+aBRQu0zZpNm KhlCmo+m7uwL7CufTiEplir3Eu5kKVf2EjVyvc511D9TOZhL9+F2ZaDG X-Gm-Gg: AYBFou2Rh1mNoYMtsegqqmY8HlLHcKMEWjiQ6lgV6tnH6zeTx5IK0CjR+yXjUyBk+4L D4LxP99HIyU6ldwnpLWQJCLve9zlwGlO2dGyKF/Ub7wWSUAILGo9NpPFEUjdJCNlV/rLMKb7Rvu KzdHsENutrnjUZ7WppKmcJ/p0GP2GtjPTnQJl4RWnPaVbUgnU3h9C+0DyHGGytPhvMeDfaF9BQ7 zzoYQnxFlBBSfqi+45/TdbBoaNphvOM4rjJcRekDTVSIbH7XO8YZL9HOO8SvFRmsumIycGk6s5c vLC4MmuOhUnaGf3fSHSU5J6syfjiyX+EZeqz4GictM+rpfTfOJqyK3puo1DF/FpUZqY7cz8rrkK 8lKgkxNfhC7AXbHKJmCUnTz7FK4lcpVWXfYgBLN6Flh7Y5BtWkDDkbBH/6CAxusVHdCfUC/+sgJ CHPBuheaZWNALhFLhFgraZ1YLKwTAWOrKReNNtff7INw4t1FztiETbq1jha59/wzyEFKxc51rDH pc2GybPbgLDyi3HurTojpvdoT8eEEkDkwWPj1A= X-Received: by 2002:a17:90b:50c6:b0:38e:97f0:aa4b with SMTP id 98e67ed59e1d1-396d1017316mr23143867a91.13.1787989636494; Sat, 29 Aug 2026 00:47:16 -0700 (PDT) Received: from localhost.localdomain (vmi2317720.contaboserver.net. [84.247.152.65]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-396b186803asm10574548a91.11.2026.08.29.00.47.12 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Sat, 29 Aug 2026 00:47:16 -0700 (PDT) From: "Lian Wang (ProcessMission)" To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, Andrew Morton , Chris Li , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , Kemeng Shi , Kairui Song , Jihan LIN , Kunwu Chan Subject: [RESEND RFC PATCH v2 03/13] mm/swap: cleanup and document swap device availability flag usage Date: Sat, 29 Aug 2026 15:47:08 +0800 Message-ID: <20260829-swap-pcp-priq-v2-resend-3-68d3d925578c@gmail.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260829-swap-pcp-priq-v2-resend-0-68d3d925578c@gmail.com> References: <20260829-swap-pcp-priq-v2-resend-0-68d3d925578c@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Kairui Song Rework swap device flag and metadata handling and swapon/swapoff to be cleaner and better documented, in preparation for locking cleanup. Consolidate existing routines and introduce swap_device_enable() and swap_device_disable() as the clean boundary of exposing or isolating a swap device. swap_device_enable() sets the proper flags and exposes the device as an allocation candidate, while swap_device_disable() clears related flags and ensures no more allocations will happen. Keep the mapping lookup and disable transition in the same swap_lock critical section. Otherwise, a concurrent swapoff can release the selected slot and swapon can reuse it for a different device before the first syscall disables the swap_info_struct it found. Drain the cluster allocators in the same helper after dropping swap_lock, so the identity transition remains atomic without holding the lock over all cluster locks. Add comment blocks documenting the lifetime and locking rules for swap device flags and their locking conventions. Apart from closing that lifecycle race, the remaining changes are code rearrangement and documentation. Signed-off-by: Kairui Song Co-developed-by: Lian Wang (ProcessMission) Signed-off-by: Lian Wang (ProcessMission) Tested-by: Kunwu Chan Tested-by: Kunwu Chan --- include/linux/swap.h | 27 ++++-- mm/swapfile.c | 207 ++++++++++++++++++++----------------------- 2 files changed, 115 insertions(+), 119 deletions(-) diff --git a/include/linux/swap.h b/include/linux/swap.h index 5658a1634b85..d9e535cd07c5 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -194,6 +194,22 @@ struct swap_extent { ((offsetof(union swap_header, magic.magic) - \ offsetof(union swap_header, info.badpages)) / sizeof(int)) =20 +/* + * Swap device flags, except the ones documented below, all are immutable + * after exposed by swap_device_enable, and until the device is freed again + * (SWP_USED unset). The exceptions: + * - SWP_USED: Protected by swap_lock. Indicates the device is inuse. Once + * set, won't be cleared unless all reference to this device is freed and + * swapoff finished. + * - SWP_WRITEOK: Protected by both swap_lock and swap_avail_lock, clearing + * this flag also waits for all current cluster lock users to exit so + * checking this flag while holding any of these locks ensures the device + * is safe to use at the moment. Note: clearing this flag doesn't affect + * pending IO or async requests, it only prevents further entry allocati= on + * or new async request (e.g. discard) from initiating. + * - SWP_HIBERNATION: Protected by swap_lock. Indicates if the device + * is pinned for hibernation. + */ enum { SWP_USED =3D (1 << 0), /* is slot in swap_info[] used? */ SWP_WRITEOK =3D (1 << 1), /* ok to write to this swap? */ @@ -262,14 +278,9 @@ struct swap_info_struct { struct file *swap_file; /* seldom referenced */ struct completion comp; /* seldom referenced */ spinlock_t lock; /* - * protect map scan related fields like - * inuse_pages and all cluster lists. - * Other fields are only changed - * at swapon/swapoff, so are protected - * by swap_lock. changing flags need - * hold this lock and swap_lock. If - * both locks need hold, hold swap_lock - * first. + * Protect cluster lists. Other fields + * are only changed at swapon/swapoff, + * so are protected by swap_lock. */ struct work_struct discard_work; /* discard worker */ struct work_struct reclaim_work; /* reclaim worker */ diff --git a/mm/swapfile.c b/mm/swapfile.c index 46772d0e3e68..d7115b9195a6 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -1199,28 +1199,20 @@ static void del_from_avail_list(struct swap_info_st= ruct *si, bool swapoff) =20 spin_lock(&swap_avail_lock); =20 - if (swapoff) { - /* - * Forcefully remove it. Clear the SWP_WRITEOK flags for - * swapoff here so it's synchronized by both si->lock and - * swap_avail_lock, to ensure the result can be seen by - * add_to_avail_list. - */ - lockdep_assert_held(&si->lock); - si->flags &=3D ~SWP_WRITEOK; - atomic_long_or(SWAP_USAGE_OFFLIST_BIT, &si->inuse_pages); - } else { - /* - * If not called by swapoff, take it off-list only if it's - * full and SWAP_USAGE_OFFLIST_BIT is not set (strictly - * si->inuse_pages =3D=3D pages), any concurrent slot freeing, - * or device already removed from plist by someone else - * will make this return false. - */ + /* + * Force remove it only for swapoff. Else, take it off-list only if + * it's full and SWAP_USAGE_OFFLIST_BIT is not set (strictly + * si->inuse_pages =3D=3D pages), so concurrent slot freeing, or + * concurrent list removal will make the cmpxchg fail and skip + * the removal. + */ + if (!swapoff) { pages =3D si->pages; if (!atomic_long_try_cmpxchg(&si->inuse_pages, &pages, pages | SWAP_USAGE_OFFLIST_BIT)) goto skip; + } else { + atomic_long_or(SWAP_USAGE_OFFLIST_BIT, &si->inuse_pages); } =20 plist_del(&si->avail_list, &swap_avail_head); @@ -1230,21 +1222,21 @@ static void del_from_avail_list(struct swap_info_st= ruct *si, bool swapoff) } =20 /* SWAP_USAGE_OFFLIST_BIT can only be cleared by this helper. */ -static void add_to_avail_list(struct swap_info_struct *si, bool swapon) +static void add_to_avail_list(struct swap_info_struct *si) { long val; unsigned long pages; =20 spin_lock(&swap_avail_lock); =20 - /* Corresponding to SWP_WRITEOK clearing in del_from_avail_list */ - if (swapon) { - lockdep_assert_held(&si->lock); - si->flags |=3D SWP_WRITEOK; - } else { - if (!(READ_ONCE(si->flags) & SWP_WRITEOK)) - goto skip; - } + /* + * Add the device to the avail list if SWP_WRITEOK is set and + * SWAP_USAGE_OFFLIST_BIT is still set. Swapoff clears + * SWP_WRITEOK first, so the device won't be re-added after + * swapoff starts unless swap_device_enable resurrects it. + */ + if (!(si->flags & SWP_WRITEOK)) + goto skip; =20 if (!(atomic_long_read(&si->inuse_pages) & SWAP_USAGE_OFFLIST_BIT)) goto skip; @@ -1300,7 +1292,7 @@ static void swap_usage_sub(struct swap_info_struct *s= i, unsigned int nr_entries) * add it to the plist. */ if (unlikely(val & SWAP_USAGE_OFFLIST_BIT)) - add_to_avail_list(si, false); + add_to_avail_list(si); } =20 static void swap_range_alloc(struct swap_info_struct *si, @@ -1354,7 +1346,7 @@ static bool get_swap_device_info(struct swap_info_str= uct *si) * up to dated. * * Paired with the spin_unlock() after setup_swap_info() in - * enable_swap_info(), and smp_wmb() in swapoff. + * swap_device_enable(), and smp_wmb() in swapoff. */ smp_rmb(); return true; @@ -2977,58 +2969,87 @@ static int setup_swap_extents(struct swap_info_stru= ct *sis, return generic_swapfile_activate(sis, swap_file, span); } =20 -static void _enable_swap_info(struct swap_info_struct *si) +/* + * Mark a fully initialized swap device writable and expose it to the + * allocator. The caller must have resurrected its percpu ref first. + */ +static void swap_device_enable(struct swap_info_struct *si) { - atomic_long_add(si->pages, &nr_swap_pages); - total_swap_pages +=3D si->pages; + spin_lock(&swap_lock); =20 - assert_spin_locked(&swap_lock); + spin_lock(&swap_avail_lock); + si->flags |=3D SWP_WRITEOK; + spin_unlock(&swap_avail_lock); =20 + atomic_long_add(si->pages, &nr_swap_pages); + total_swap_pages +=3D si->pages; plist_add(&si->list, &swap_active_head); + spin_unlock(&swap_lock); =20 - /* Add back to available list */ - add_to_avail_list(si, true); + add_to_avail_list(si); } =20 -/* - * Called after the swap device is ready, resurrect its percpu ref, it's n= ow - * safe to reference it. Add it to the list to expose it to the allocator. - */ -static void enable_swap_info(struct swap_info_struct *si) +static int swap_device_disable(struct address_space *mapping, + struct swap_info_struct **swap_info) { - percpu_ref_resurrect(&si->users); - spin_lock(&swap_lock); - spin_lock(&si->lock); - _enable_swap_info(si); - spin_unlock(&si->lock); - spin_unlock(&swap_lock); -} + struct swap_info_struct *si; + struct swap_cluster_info *ci; + unsigned long offset, end; + int err =3D -EINVAL; =20 -static void reinsert_swap_info(struct swap_info_struct *si) -{ spin_lock(&swap_lock); - spin_lock(&si->lock); - _enable_swap_info(si); - spin_unlock(&si->lock); - spin_unlock(&swap_lock); -} + plist_for_each_entry(si, &swap_active_head, list) { + if ((si->flags & SWP_WRITEOK) && + si->swap_file->f_mapping =3D=3D mapping) { + err =3D 0; + break; + } + } + if (err) + goto unlock; =20 -/* - * Called after clearing SWP_WRITEOK, ensures cluster_alloc_range - * see the updated flags, so there will be no more allocations. - */ -static void wait_for_allocation(struct swap_info_struct *si) -{ - unsigned long offset; - unsigned long end =3D ALIGN(si->max, SWAPFILE_CLUSTER); - struct swap_cluster_info *ci; + /* + * Refuse swapoff while the device is pinned for hibernation. + */ + if (si->flags & SWP_HIBERNATION) { + err =3D -EBUSY; + goto unlock; + } =20 - BUG_ON(si->flags & SWP_WRITEOK); + if (security_vm_enough_memory_mm(current->mm, si->pages)) { + err =3D -ENOMEM; + goto unlock; + } + vm_unacct_memory(si->pages); =20 + spin_lock(&swap_avail_lock); + si->flags &=3D ~SWP_WRITEOK; + spin_unlock(&swap_avail_lock); + + plist_del(&si->list, &swap_active_head); + total_swap_pages -=3D si->pages; + atomic_long_sub(si->pages, &nr_swap_pages); + + end =3D ALIGN(si->max, SWAPFILE_CLUSTER); +unlock: + spin_unlock(&swap_lock); + if (err) + return err; + + del_from_avail_list(si, true); + + /* + * The swap allocator doesn't take swap_lock. Looping through every + * cluster lock after clearing SWP_WRITEOK ensures that allocators see + * the updated flag and that no allocation remains in flight. + */ for (offset =3D 0; offset < end; offset +=3D SWAPFILE_CLUSTER) { ci =3D swap_cluster_lock(si, offset); swap_cluster_unlock(ci); } + + *swap_info =3D si; + return 0; } =20 static void free_swap_cluster_info(struct swap_cluster_info *cluster_info, @@ -3073,7 +3094,6 @@ static void flush_percpu_swap_cluster(struct swap_inf= o_struct *si) } } =20 - SYSCALL_DEFINE1(swapoff, const char __user *, specialfile) { struct swap_info_struct *p =3D NULL; @@ -3082,7 +3102,7 @@ SYSCALL_DEFINE1(swapoff, const char __user *, special= file) struct address_space *mapping; struct inode *inode; unsigned int maxpages; - int err, found =3D 0; + int err; =20 if (!capable(CAP_SYS_ADMIN)) return -EPERM; @@ -3095,44 +3115,11 @@ SYSCALL_DEFINE1(swapoff, const char __user *, speci= alfile) return PTR_ERR(victim); =20 mapping =3D victim->f_mapping; - spin_lock(&swap_lock); - plist_for_each_entry(p, &swap_active_head, list) { - if (p->flags & SWP_WRITEOK) { - if (p->swap_file->f_mapping =3D=3D mapping) { - found =3D 1; - break; - } - } - } - if (!found) { - err =3D -EINVAL; - spin_unlock(&swap_lock); - goto out_dput; - } - - /* Refuse swapoff while the device is pinned for hibernation */ - if (p->flags & SWP_HIBERNATION) { - err =3D -EBUSY; - spin_unlock(&swap_lock); - goto out_dput; - } - - if (!security_vm_enough_memory_mm(current->mm, p->pages)) - vm_unacct_memory(p->pages); - else { - err =3D -ENOMEM; - spin_unlock(&swap_lock); - goto out_dput; - } - spin_lock(&p->lock); - del_from_avail_list(p, true); - plist_del(&p->list, &swap_active_head); - atomic_long_sub(p->pages, &nr_swap_pages); - total_swap_pages -=3D p->pages; - spin_unlock(&p->lock); - spin_unlock(&swap_lock); + err =3D swap_device_disable(mapping, &p); + filp_close(victim, NULL); =20 - wait_for_allocation(p); + if (err) + return err; =20 set_current_oom_origin(); err =3D try_to_unuse(p->type); @@ -3140,8 +3127,8 @@ SYSCALL_DEFINE1(swapoff, const char __user *, special= file) =20 if (err) { /* re-insert swap space back into swap_list */ - reinsert_swap_info(p); - goto out_dput; + swap_device_enable(p); + return err; } =20 /* @@ -3201,13 +3188,10 @@ SYSCALL_DEFINE1(swapoff, const char __user *, speci= alfile) p->flags =3D 0; spin_unlock(&swap_lock); =20 - err =3D 0; atomic_inc(&proc_poll_event); wake_up_interruptible(&proc_poll_wait); =20 -out_dput: - filp_close(victim, NULL); - return err; + return 0; } =20 #ifdef CONFIG_PROC_FS @@ -3629,7 +3613,7 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialf= ile, int, swap_flags) /* * Allocate or reuse existing !SWP_USED swap_info. The returned * si will stay in a dying status, so nothing will access its content - * until enable_swap_info resurrects its percpu ref and expose it. + * until swap_device_enable resurrects its percpu ref and expose it. */ si =3D alloc_swap_info(); if (IS_ERR(si)) @@ -3794,7 +3778,8 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialf= ile, int, swap_flags) si->swap_file =3D swap_file; =20 /* Sets SWP_WRITEOK, resurrect the percpu ref, expose the swap device */ - enable_swap_info(si); + percpu_ref_resurrect(&si->users); + swap_device_enable(si); =20 pr_info("Adding %uk swap on %s. Priority:%d extents:%d across:%lluk %s%s= %s%s\n", K(si->pages), name->name, si->prio, nr_extents, --=20 2.55.0 From nobody Sat Sep 26 22:02:19 2026 Received: from mail-pj1-f54.google.com (mail-pj1-f54.google.com [209.85.216.54]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BE52535C69E for ; Sat, 29 Aug 2026 07:47:26 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.54 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787989648; cv=none; b=XMrzMr8q1rE7qKau/Vul1GFuTxLkZ6mIE38SVJFlFLwByQwTwJ6tObPaejXakSkDLwsWWS3ogvmPchD6LimmpZYmwGi1860Air5CsG+LNjmSxwe46sTyBN0w6HFroq5BMqopop8e40HGuQ4IlIh6eTNUWMxVhTqVvc5sQXVF1m8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787989648; c=relaxed/simple; bh=CM6cPd6FOwbZ3CIDYweTMGQ1LQbkEFKNsimGOKG1Mjk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=FtjgKfgOWQP+m2odyuM8xCdrUz72il8R0Sud4t1gMI8IAmtdOyFnKINof2DTvCNj+Bq5CVKfa6PKLh3hDmf2FZ1SF5MD/4iAM5LaHCeCYjQVzhOtVXJX5ktTch5GHXwTllxKoqmheO4JuvGB+HSXgOFPyQYujvwsCIEr11eBDf0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=gZBULyKY; arc=none smtp.client-ip=209.85.216.54 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="gZBULyKY" Received: by mail-pj1-f54.google.com with SMTP id 98e67ed59e1d1-38dfe910e9dso2081727a91.3 for ; Sat, 29 Aug 2026 00:47:26 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787989646; x=1788594446; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=ajNLhdeOa3qTi7x0+Rqx535kN8WB+xlcS2id0wkaitw=; b=gZBULyKYzH3QwnMzHgGvj6R2Pj8jE3qBUN1lSnJZPShUvYMJXJ/RsJKmxqj9FQqA/8 dfWGNUDR9TINH+GTSSOi/cXy3jhWUttB/kvqlklCEaGwR/kwGtKFzm7ix1l9Cv5r2lh+ PAjVGmCbR3sZisbFPXny4iFdSG4Sny1cyBibYcKFFb1qhH/5qjmNvH3QrjeoojSts2jl W7V4bLOwpgxuf+CVEtAnoKa8aj9umafLxz2FEUBf4B+bVxQh+gG+buz0RyCSV1KCiMAo RpPNsj0MMF2F03NfNw5UZXDUn+XfyAJROMa74sjzxaBzq6hU7/njeRGnRpRJpgoWl3xc r4Ug== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787989646; x=1788594446; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=ajNLhdeOa3qTi7x0+Rqx535kN8WB+xlcS2id0wkaitw=; b=Vra7N4I9IpwAx+tTmeegn0TgrVPYzllAE/zsQm/bm4rQB9/swyq+SU6wIbQPWSbOX+ 0AsPn66geM6CvFnU/8KJ71qNQAH8+hnz49CutKqL7UzeYJHZpp3ZBm8H6RcLspDZtzX+ rgGaNr3ikbIvqXcUfeSwBjA+RB9c0W1s3SlGgQeG7RzZYjsUnWmRbMBHpAqEslpu7DSl 6boJvmXJcscSRrLSg/6pki9q7VfhJEN6L3SSRetXV2ftjDAOuWOrkJYh2LWXTjCYbjpf MYu8YjnbfgmBxs1djKyNDTiysxQKTRqG5AXJ2+vCbwk+qzl4BZMhMkVIJk+vixuPOjBd dUIQ== X-Gm-Message-State: AFuF++kwFWwYC4YjAvmt5p42Tk7V1edQYvZB5Um1pBlwE/ORSxsub4yg uZzLbkdUAhsTIjT/3nX0w5mCwj8QibRRfrlCFxWNswzKltBXTQE1hfc+ X-Gm-Gg: AYBFou3rtc8TJ2UFz2PfXQzg8ikwCCtRy9F/hEzv9X1Grw/v+AeAr0dt5b69d0wt3B+ rvqfEXI/kCZce0QFAw6KFRpoBTDTqCw8r+0fazMClEtivPggNH2c33QRrGNCoKd6/TVSBvwMxCC DinEWADro+xsPUnfwr8rohWeFq05ue9UebNKSxCyzTCSa103Z3O42gtotgqWlgJ4kZBeRTkfLGS bNfpeGfPZV7YDrMpgDFu3zaRs/6NFlXv6ers4IsG0+C1H2o6FlukyKY7AV3o4QSlzUh636cYbIh Ax9rIER3uqYHjSdl6IBGseVRVddYlP7TCK0F/jhdVDvmdhcwpLw8wtCUXdyBWN9VW5vDTJqTX8+ OyAGzjLVWxivcAbBFATybwO3O3WDoNoDJ51rCKXVCFuyC7HYJ9etzhZ8aDCEvLRpp2XAOIRn0go dRXu2nnfnKZ26wUaUK3DaJVVCfRjDoH9Lb8iKDrWnG34XQBmB7EIP4b81F0w68O8pRa4yuHLrG7 l4aUNkC5QkkQRl8EnrxvEktYgQE X-Received: by 2002:a17:90b:1806:b0:398:9be8:ea68 with SMTP id 98e67ed59e1d1-3989be8eaafmr5958988a91.21.1787989646053; Sat, 29 Aug 2026 00:47:26 -0700 (PDT) Received: from localhost.localdomain (vmi2317720.contaboserver.net. [84.247.152.65]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-398a14768dcsm853443a91.4.2026.08.29.00.47.21 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Sat, 29 Aug 2026 00:47:25 -0700 (PDT) From: "Lian Wang (ProcessMission)" To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, Andrew Morton , Chris Li , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , Kemeng Shi , Kairui Song , Jihan LIN , Kunwu Chan Subject: [RESEND RFC PATCH v2 04/13] mm/swap: introduce swap device iteration helper Date: Sat, 29 Aug 2026 15:47:17 +0800 Message-ID: <20260829-swap-pcp-priq-v2-resend-4-68d3d925578c@gmail.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260829-swap-pcp-priq-v2-resend-0-68d3d925578c@gmail.com> References: <20260829-swap-pcp-priq-v2-resend-0-68d3d925578c@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Kairui Song There are many users that access the swap info array just to iterate through the swap devices to find usable ones. Introduce a generic helper for this to prepare for dropping the lock convention. Also slightly adjust __swap_type_to_info to allow access of swapped off devices by moving the sanity check into its only current caller, and this way makes more sense too: the only caller is __swap_entry_to_info, where we should never see an entry referring to a dead device. Signed-off-by: Kairui Song Signed-off-by: Lian Wang (ProcessMission) Tested-by: Kunwu Chan Tested-by: Kunwu Chan --- mm/swap.h | 12 ++++---- mm/swapfile.c | 77 +++++++++++++++++++++++++++++++-------------------- 2 files changed, 53 insertions(+), 36 deletions(-) diff --git a/mm/swap.h b/mm/swap.h index 90a551a88df6..403574a7fa8b 100644 --- a/mm/swap.h +++ b/mm/swap.h @@ -107,16 +107,16 @@ static inline unsigned int swp_cluster_offset(swp_ent= ry_t entry) */ static inline struct swap_info_struct *__swap_type_to_info(int type) { - struct swap_info_struct *si; - - si =3D READ_ONCE(swap_info[type]); /* rcu_dereference() */ - VM_WARN_ON_ONCE(percpu_ref_is_zero(&si->users)); /* race with swapoff */ - return si; + return READ_ONCE(swap_info[type]); /* rcu_dereference() */ } =20 static inline struct swap_info_struct *__swap_entry_to_info(swp_entry_t en= try) { - return __swap_type_to_info(swp_type(entry)); + struct swap_info_struct *si; + + si =3D __swap_type_to_info(swp_type(entry)); /* rcu_dereference() */ + VM_WARN_ON_ONCE(percpu_ref_is_zero(&si->users)); /* race with swapoff */ + return si; } =20 static inline struct swap_cluster_info *__swap_offset_to_cluster( diff --git a/mm/swapfile.c b/mm/swapfile.c index d7115b9195a6..3847717f4d1d 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -106,6 +106,34 @@ static DEFINE_SPINLOCK(swap_avail_lock); =20 struct swap_info_struct *swap_info[MAX_SWAPFILES]; =20 +static inline struct swap_info_struct *__swap_iter(int *i, unsigned long f= lag) +{ + lockdep_assert_held(&swap_lock); + while (*i < nr_swapfiles) { + struct swap_info_struct *si =3D __swap_type_to_info(*i); + + VM_WARN_ON(!si); + (*i)++; + if (flag && !((si->flags & flag) =3D=3D flag)) + continue; + return si; + } + return NULL; +} + +#define __for_each_swap(si, flag) \ + for (int __i =3D 0; ((si) =3D __swap_iter(&__i, flag));) + +/* + * for_each_swap - iterate through all allocated and inuse swap devices + * @si: the iterator + * + * Context: The caller must hold swap_lock. The lock may be dropped during + * the loop body but must be re-acquired before the next iteration. + */ +#define for_each_swap(si) __for_each_swap(si, SWP_USED) +#define for_each_avail_swap(si) __for_each_swap(si, SWP_USED | SWP_WRITEOK) + static struct kmem_cache *swap_table_cachep; =20 /* Protects si->swap_file for /proc/swaps usage */ @@ -2210,26 +2238,18 @@ void swap_free_hibernation_slot(swp_entry_t entry) =20 static int __find_hibernation_swap_type(dev_t device, sector_t offset) { - int type; - - lockdep_assert_held(&swap_lock); + struct swap_info_struct *si; =20 if (!device) return -EINVAL; =20 - for (type =3D 0; type < nr_swapfiles; type++) { - struct swap_info_struct *sis =3D swap_info[type]; - - if (!(sis->flags & SWP_WRITEOK)) - continue; - - if (device =3D=3D sis->bdev->bd_dev) { - struct swap_extent *se =3D first_se(sis); - - if (se->start_block =3D=3D offset) - return type; + for_each_avail_swap(si) { + if (device =3D=3D si->bdev->bd_dev) { + if (first_se(si)->start_block =3D=3D offset) + return si->type; } } + return -ENODEV; } =20 @@ -2335,16 +2355,13 @@ int find_hibernation_swap_type(dev_t device, sector= _t offset) =20 int find_first_swap(dev_t *device) { - int type, ret =3D -ENODEV; + int ret =3D -ENODEV; + struct swap_info_struct *si; =20 spin_lock(&swap_lock); - for (type =3D 0; type < nr_swapfiles; type++) { - struct swap_info_struct *sis =3D swap_info[type]; - - if (!(sis->flags & SWP_WRITEOK)) - continue; - *device =3D sis->bdev->bd_dev; - ret =3D type; + for_each_avail_swap(si) { + *device =3D si->bdev->bd_dev; + ret =3D si->type; break; } spin_unlock(&swap_lock); @@ -2830,12 +2847,14 @@ static int try_to_unuse(unsigned int type) */ static void drain_mmlist(void) { + struct swap_info_struct *si; struct list_head *p, *next; - unsigned int type; =20 - for (type =3D 0; type < nr_swapfiles; type++) - if (swap_usage_in_pages(swap_info[type])) + for_each_swap(si) { + if (swap_usage_in_pages(si)) return; + } + spin_lock(&mmlist_lock); list_for_each_safe(p, next, &init_mm.mmlist) list_del_init(p); @@ -3827,14 +3846,12 @@ SYSCALL_DEFINE2(swapon, const char __user *, specia= lfile, int, swap_flags) =20 void si_swapinfo(struct sysinfo *val) { - unsigned int type; + struct swap_info_struct *si; unsigned long nr_to_be_unused =3D 0; =20 spin_lock(&swap_lock); - for (type =3D 0; type < nr_swapfiles; type++) { - struct swap_info_struct *si =3D swap_info[type]; - - if ((si->flags & SWP_USED) && !(si->flags & SWP_WRITEOK)) + for_each_swap(si) { + if (!(si->flags & SWP_WRITEOK)) nr_to_be_unused +=3D swap_usage_in_pages(si); } val->freeswap =3D atomic_long_read(&nr_swap_pages) + nr_to_be_unused; --=20 2.55.0 From nobody Sat Sep 26 22:02:19 2026 Received: from mail-pj1-f51.google.com (mail-pj1-f51.google.com [209.85.216.51]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1428E3264C1 for ; Sat, 29 Aug 2026 07:47:34 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.51 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787989656; cv=none; b=Yca0meW7hlIurVfwEXXapbd7OrN+WGPjInwJQLmYQByk2jta8FSd1jg4rq2fKhBtThZQaQ7+A3SiFBs6YfCehL4xkAhbBUSv03BWFhqIdun79yZC7hRS05mbjngGM66+ZSMUhB5eQLhy/xtdJgDAWam5BtBmG8mZZcU56eybJSc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787989656; c=relaxed/simple; bh=eXlV30C1xX92vS5aeXmw7j6LRd3WlS+9wBkvu258iW8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=d4RviCyYDj6/kOtow06Gsf3b0sZ+hJL9hz3FWU5DyP7X480H1dJspVgvE7yBvde//+ReYSLptVfRzK71Ky1dF2SH4M6uC7uoTJNxF8KZYHCJoSwZ7ZPDc0ulhcdAzhdnngitCgXPbRtnhHjSNGhFo0wx0gTK46wrEDjHslBCal0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=esBo4pRg; arc=none smtp.client-ip=209.85.216.51 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="esBo4pRg" Received: by mail-pj1-f51.google.com with SMTP id 98e67ed59e1d1-3964dfb5b9aso1826479a91.1 for ; Sat, 29 Aug 2026 00:47:34 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787989654; x=1788594454; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=cHvxeW8lAUpfJ18lPgBl3HLzf+AeoDV2A4HcxkONfAI=; b=esBo4pRgM5i7XCkbnQApFqJyENuZoOOBTWDC9bBjf82/XZYTTV+r1kN7h647WD11Q0 runBZIUNLM2grs0L0tUoebY2Vys9MHCJQwkJsLD41+dMW8gY9WA98nJM4dVz7tgWQI43 UIEoZr01v3/j+eT74pbPp1Xb5VcWpYfotyoRO3SZ/YQhfRbMxpBO96nynnq3l7v+t5gy wHA+CEbH5aQtMQkiRJYngGai4pejfP026/1fU7I+swv+WIozWPjNIWvwxVGbrDySC50P DfnD6am1JAHHXfNdpy/DasVRMG49LUJjEPXnNXZTQKbS6xeHyA/Wl3iFGRkG2G/wflQw odcA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787989654; x=1788594454; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=cHvxeW8lAUpfJ18lPgBl3HLzf+AeoDV2A4HcxkONfAI=; b=ngiZDOUSMAhn6wnjZ5Jpr1fKm4uvNlj1pT5AnJ/DfOliidVmmCVSnySTmqUmaPb5vE QmLNtlCZ4stg53dGBGN0ojaRceWAGoVhsd08/vtqjuItn65VECiJd0AdaKZIkGmDNevg 2oUvMnRN5pZ9YBLDEISpCJQyVDo1+Cg2Zhf6Hz5fb0nPX7+XEzrpQDuNMwBGGWwCnV5h aBq68z+k6yHLbGLY3K9jTEWCzdTDWXj7amf9qD4LJs3YwGDvLyIzDD/fnVYGY8SgJhmZ KbcTVAZn2lIfwccZ2TN1OyJw91KIH/3i4UWLFHmrl5YXr4kLbwOv2ZAS0pQ19B/qPyLd 3TMw== X-Gm-Message-State: AFuF++lG/Ma6G/oERaL8vDevpfY/6xV+QyRQNam8Zhl3vSwFiI4YdWN+ +4Z1MnwcH/J1TadEpR3m4EbnvdPKi28oufN624J/yOfD7qZsav/xX+II X-Gm-Gg: AYBFou1pDFH+PE9mkyZT7n5EVyHdIxJGzl/svcUx2RmdZq+uK9SrG6RWoFpm2VjuZBY dGTCl69/zJhNtSLN5H3kJD8eTxAktFUdyZtBSjMyPO5a1DZ5dLttaLmgyDllhAWoSdyCMR8uewY g5WlVl9/yUQx+pLfIet4I9lDz4R2j9lJuY0Tqqv2K+Nx4fnN3vmhcpoLXRwfndkSTGdNjSYwY3J riu5hkd9dOCsos4D/+VDRMPqU90B2x9uYqhfw0AqGi0JktFz/FFkf/CHZ8zNK8ULdJTjYDjwPzv x7zJXjtipvk8ttattNv+ZLaSadXhD4ox/h+0szxo8m18zAmI0ybImioMaxZbiKVNOZq7+FpWk3e mw6ElkK2d6AT7xZsmJCKKW8ZTelRo9mW1fwaW2wCrjDWE8spclL89RzEVQp0dZaJoRVPJv9bPhh T6dpyEBVMDX3dmzJ9U1mE27Ei6zYxN23PKw/QWxX1KHcTZ/RVMQWQabYQQAVUStS4hdztlWGznB 1VEEmxTJc9VTowrCrE65d6+u7ii X-Received: by 2002:a17:90b:5108:b0:398:9c00:29ea with SMTP id 98e67ed59e1d1-3989c003c23mr7250051a91.18.1787989654371; Sat, 29 Aug 2026 00:47:34 -0700 (PDT) Received: from localhost.localdomain (vmi2317720.contaboserver.net. [84.247.152.65]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-398bd356f80sm172493a91.9.2026.08.29.00.47.29 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Sat, 29 Aug 2026 00:47:34 -0700 (PDT) From: "Lian Wang (ProcessMission)" To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, Andrew Morton , Chris Li , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , Kemeng Shi , Kairui Song , Jihan LIN , Kunwu Chan Subject: [RESEND RFC PATCH v2 05/13] mm/swap: change the swapon lock into a percpu rwsem Date: Sat, 29 Aug 2026 15:47:26 +0800 Message-ID: <20260829-swap-pcp-priq-v2-resend-5-68d3d925578c@gmail.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260829-swap-pcp-priq-v2-resend-0-68d3d925578c@gmail.com> References: <20260829-swap-pcp-priq-v2-resend-0-68d3d925578c@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Kairui Song Reading swap metadata is a common hot path, while swapon and swapoff are costly and very rare. Convert the existing swap_lock spinlock into a percpu rwsem so that readers can run concurrently without contention. This is a first step toward cleaning up the swap lock model and we might already see a performance gain for existing readers. Also update the comments in swap_info_struct to reflect the lock name change. Signed-off-by: Kairui Song Signed-off-by: Lian Wang (ProcessMission) Tested-by: Kunwu Chan Tested-by: Kunwu Chan --- include/linux/swap.h | 10 ++-- mm/swapfile.c | 115 +++++++++++++++++++++++-------------------- 2 files changed, 66 insertions(+), 59 deletions(-) diff --git a/include/linux/swap.h b/include/linux/swap.h index d9e535cd07c5..7b8ba0d1903f 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -197,17 +197,17 @@ struct swap_extent { /* * Swap device flags, except the ones documented below, all are immutable * after exposed by swap_device_enable, and until the device is freed again - * (SWP_USED unset). The exceptions: - * - SWP_USED: Protected by swap_lock. Indicates the device is inuse. Once + * (SWP_USED unset). The exceptions are all protected by swapon_rwsem: + * - SWP_USED: Protected by swapon_rwsem. Indicates the device is inuse. O= nce * set, won't be cleared unless all reference to this device is freed and * swapoff finished. - * - SWP_WRITEOK: Protected by both swap_lock and swap_avail_lock, clearing + * - SWP_WRITEOK: Protected by both swapon_rwsem and swap_avail_lock, clea= ring * this flag also waits for all current cluster lock users to exit so * checking this flag while holding any of these locks ensures the device * is safe to use at the moment. Note: clearing this flag doesn't affect * pending IO or async requests, it only prevents further entry allocati= on * or new async request (e.g. discard) from initiating. - * - SWP_HIBERNATION: Protected by swap_lock. Indicates if the device + * - SWP_HIBERNATION: Protected by swapon_rwsem. Indicates if the device * is pinned for hibernation. */ enum { @@ -280,7 +280,7 @@ struct swap_info_struct { spinlock_t lock; /* * Protect cluster lists. Other fields * are only changed at swapon/swapoff, - * so are protected by swap_lock. + * so are protected by swapon_rwsem. */ struct work_struct discard_work; /* discard worker */ struct work_struct reclaim_work; /* reclaim worker */ diff --git a/mm/swapfile.c b/mm/swapfile.c index 3847717f4d1d..173375a7a215 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -57,23 +57,29 @@ static void move_cluster(struct swap_info_struct *si, enum swap_cluster_flags new_flags); =20 /* - * Protects the swap_info array, and the SWP_USED flag. swap_info contains - * lazily allocated & freed swap device info struts, and SWP_USED indicates - * which device is used, ~SWP_USED devices and can be reused. - * - * Also protects swap_active_head total_swap_pages, and the SWP_WRITEOK fl= ag. + * Serializes swapon/swapoff (writers) and protects the swap_info + * array, nr_swapfiles, total_swap_pages, and part of swap device + * info content (see comment of swap_info_struct). Readers + * (allocation, /proc/swaps, etc.) take percpu_down_read() which + * is cheap as the hot path. + */ +DEFINE_STATIC_PERCPU_RWSEM(swapon_rwsem); +/* + * swap_info contains lazily allocated swap device info structs, and + * SWP_USED indicates which device is used, ~SWP_USED devices can be + * reused. Protected by swapon_rwsem, but reading could be lockless. */ -static DEFINE_SPINLOCK(swap_lock); +struct swap_info_struct *swap_info[MAX_SWAPFILES]; static unsigned int nr_swapfiles; -atomic_long_t nr_swap_pages; +long total_swap_pages; + /* * Some modules use swappable objects and may try to swap them out under * memory pressure (via the shrinker). Before doing so, they may wish to * check to see if any swap space is available. */ +atomic_long_t nr_swap_pages; EXPORT_SYMBOL_GPL(nr_swap_pages); -/* protected with swap_lock. reading in vm_swap_full() doesn't need lock */ -long total_swap_pages; #define DEF_SWAP_PRIO -1 unsigned long swapfile_maximum_size; #ifdef CONFIG_MIGRATION @@ -85,7 +91,7 @@ static const char Bad_offset[] =3D "Bad swap offset entry= "; =20 /* * all active swap_info_structs - * protected with swap_lock, and ordered by priority. + * protected with swapon_rwsem, and ordered by priority. */ static PLIST_HEAD(swap_active_head); =20 @@ -95,20 +101,18 @@ static PLIST_HEAD(swap_active_head); * This is used by folio_alloc_swap() instead of swap_active_head * because swap_active_head includes all swap_info_structs, * but folio_alloc_swap() doesn't need to look at full ones. - * This uses its own lock instead of swap_lock because when a + * This uses its own lock instead of swapon_rwsem because when a * swap_info_struct changes between not-full/full, it needs to * add/remove itself to/from this list, but the swap_info_struct->lock - * is held and the locking order requires swap_lock to be taken + * is held and the locking order requires swapon_rwsem to be taken * before any swap_info_struct->lock. */ static PLIST_HEAD(swap_avail_head); static DEFINE_SPINLOCK(swap_avail_lock); =20 -struct swap_info_struct *swap_info[MAX_SWAPFILES]; - static inline struct swap_info_struct *__swap_iter(int *i, unsigned long f= lag) { - lockdep_assert_held(&swap_lock); + lockdep_assert_held(&swapon_rwsem); while (*i < nr_swapfiles) { struct swap_info_struct *si =3D __swap_type_to_info(*i); =20 @@ -128,7 +132,7 @@ static inline struct swap_info_struct *__swap_iter(int = *i, unsigned long flag) * for_each_swap - iterate through all allocated and inuse swap devices * @si: the iterator * - * Context: The caller must hold swap_lock. The lock may be dropped during + * Context: The caller must hold swapon_rwsem. The lock may be dropped dur= ing * the loop body but must be re-acquired before the next iteration. */ #define for_each_swap(si) __for_each_swap(si, SWP_USED) @@ -1371,10 +1375,10 @@ static bool get_swap_device_info(struct swap_info_s= truct *si) /* * Guarantee the si->users are checked before accessing other * fields of swap_info_struct, and si->flags (SWP_WRITEOK) is - * up to dated. + * up to date. * - * Paired with the spin_unlock() after setup_swap_info() in - * swap_device_enable(), and smp_wmb() in swapoff. + * Paired with percpu_up_write() in swap_device_enable(), and + * smp_wmb() after clearing SWP_WRITEOK in swapoff. */ smp_rmb(); return true; @@ -1459,10 +1463,10 @@ static bool swap_sync_discard(void) bool ret =3D false; struct swap_info_struct *si, *next; =20 - spin_lock(&swap_lock); + percpu_down_read(&swapon_rwsem); start_over: plist_for_each_entry_safe(si, next, &swap_active_head, list) { - spin_unlock(&swap_lock); + percpu_up_read(&swapon_rwsem); if (get_swap_device_info(si)) { if (si->flags & SWP_PAGE_DISCARD) ret =3D swap_do_scheduled_discard(si); @@ -1471,11 +1475,11 @@ static bool swap_sync_discard(void) if (ret) return true; =20 - spin_lock(&swap_lock); + percpu_down_read(&swapon_rwsem); if (plist_node_empty(&next->list)) goto start_over; } - spin_unlock(&swap_lock); + percpu_up_read(&swapon_rwsem); =20 return false; } @@ -2275,7 +2279,7 @@ int pin_hibernation_swap_type(dev_t device, sector_t = offset) int ret; struct swap_info_struct *si; =20 - spin_lock(&swap_lock); + percpu_down_write(&swapon_rwsem); ret =3D __find_hibernation_swap_type(device, offset); if (ret < 0) goto out; @@ -2299,7 +2303,7 @@ int pin_hibernation_swap_type(dev_t device, sector_t = offset) si->flags |=3D SWP_HIBERNATION; =20 out: - spin_unlock(&swap_lock); + percpu_up_write(&swapon_rwsem); return ret; } =20 @@ -2317,11 +2321,11 @@ void unpin_hibernation_swap_type(int type) { struct swap_info_struct *si; =20 - spin_lock(&swap_lock); + percpu_down_write(&swapon_rwsem); si =3D swap_type_to_info(type); if (si) si->flags &=3D ~SWP_HIBERNATION; - spin_unlock(&swap_lock); + percpu_up_write(&swapon_rwsem); } =20 /** @@ -2346,9 +2350,9 @@ int find_hibernation_swap_type(dev_t device, sector_t= offset) { int type; =20 - spin_lock(&swap_lock); + percpu_down_read(&swapon_rwsem); type =3D __find_hibernation_swap_type(device, offset); - spin_unlock(&swap_lock); + percpu_up_read(&swapon_rwsem); =20 return type; } @@ -2358,13 +2362,13 @@ int find_first_swap(dev_t *device) int ret =3D -ENODEV; struct swap_info_struct *si; =20 - spin_lock(&swap_lock); + percpu_down_read(&swapon_rwsem); for_each_avail_swap(si) { *device =3D si->bdev->bd_dev; ret =3D si->type; break; } - spin_unlock(&swap_lock); + percpu_up_read(&swapon_rwsem); return ret; } =20 @@ -2393,7 +2397,7 @@ unsigned int count_swap_pages(int type, int free) { unsigned int n =3D 0; =20 - spin_lock(&swap_lock); + percpu_down_read(&swapon_rwsem); if ((unsigned int)type < nr_swapfiles) { struct swap_info_struct *sis =3D swap_info[type]; =20 @@ -2405,7 +2409,7 @@ unsigned int count_swap_pages(int type, int free) } spin_unlock(&sis->lock); } - spin_unlock(&swap_lock); + percpu_up_read(&swapon_rwsem); return n; } #endif /* CONFIG_HIBERNATION */ @@ -2717,10 +2721,10 @@ static unsigned int find_next_to_unuse(struct swap_= info_struct *si, unsigned long swp_tb; =20 /* - * No need for swap_lock here: we're just looking + * No need for swapon_rwsem here: we're just looking * for whether an entry is in use, not modifying it; false * hits are okay, and sys_swapoff() has already prevented new - * allocations from this area (while holding swap_lock). + * allocations from this area (while holding swapon_rwsem). */ for (i =3D prev + 1; i < si->max; i++) { swp_tb =3D swap_table_get(__swap_offset_to_cluster(si, i), @@ -2841,8 +2845,8 @@ static int try_to_unuse(unsigned int type) =20 /* * After a successful try_to_unuse, if no swap is now in use, we know - * we can empty the mmlist. swap_lock must be held on entry and exit. - * Note that mmlist_lock nests inside swap_lock, and an mm must be + * we can empty the mmlist. swapon_rwsem must be held on entry and exit. + * Note that mmlist_lock nests inside swapon_rwsem, and an mm must be * added to the mmlist just after page_duplicate - before would be racy. */ static void drain_mmlist(void) @@ -2994,8 +2998,7 @@ static int setup_swap_extents(struct swap_info_struct= *sis, */ static void swap_device_enable(struct swap_info_struct *si) { - spin_lock(&swap_lock); - + percpu_down_write(&swapon_rwsem); spin_lock(&swap_avail_lock); si->flags |=3D SWP_WRITEOK; spin_unlock(&swap_avail_lock); @@ -3003,7 +3006,7 @@ static void swap_device_enable(struct swap_info_struc= t *si) atomic_long_add(si->pages, &nr_swap_pages); total_swap_pages +=3D si->pages; plist_add(&si->list, &swap_active_head); - spin_unlock(&swap_lock); + percpu_up_write(&swapon_rwsem); =20 add_to_avail_list(si); } @@ -3016,7 +3019,7 @@ static int swap_device_disable(struct address_space *= mapping, unsigned long offset, end; int err =3D -EINVAL; =20 - spin_lock(&swap_lock); + percpu_down_write(&swapon_rwsem); plist_for_each_entry(si, &swap_active_head, list) { if ((si->flags & SWP_WRITEOK) && si->swap_file->f_mapping =3D=3D mapping) { @@ -3051,14 +3054,14 @@ static int swap_device_disable(struct address_space= *mapping, =20 end =3D ALIGN(si->max, SWAPFILE_CLUSTER); unlock: - spin_unlock(&swap_lock); + percpu_up_write(&swapon_rwsem); if (err) return err; =20 del_from_avail_list(si, true); =20 /* - * The swap allocator doesn't take swap_lock. Looping through every + * The swap allocator doesn't take swapon_rwsem. Looping through every * cluster lock after clearing SWP_WRITEOK ensures that allocators see * the updated flag and that no allocation remains in flight. */ @@ -3172,7 +3175,7 @@ SYSCALL_DEFINE1(swapoff, const char __user *, special= file) atomic_dec(&nr_rotate_swap); =20 mutex_lock(&swapon_mutex); - spin_lock(&swap_lock); + percpu_down_write(&swapon_rwsem); spin_lock(&p->lock); drain_mmlist(); =20 @@ -3183,7 +3186,7 @@ SYSCALL_DEFINE1(swapoff, const char __user *, special= file) p->max =3D 0; p->cluster_info =3D NULL; spin_unlock(&p->lock); - spin_unlock(&swap_lock); + percpu_up_write(&swapon_rwsem); arch_swap_invalidate_area(p->type); zswap_swapoff(p->type); mutex_unlock(&swapon_mutex); @@ -3202,10 +3205,14 @@ SYSCALL_DEFINE1(swapoff, const char __user *, speci= alfile) * Clear the SWP_USED flag after all resources are freed so that swapon * can reuse this swap_info in alloc_swap_info() safely. It is ok to * not hold p->lock after we cleared its SWP_WRITEOK. + * + * The write lock ensures the flag clear is visible to lockless + * readers of swap_type_to_info() before alloc_swap_info() reuses + * this slot. */ - spin_lock(&swap_lock); + percpu_down_write(&swapon_rwsem); p->flags =3D 0; - spin_unlock(&swap_lock); + percpu_up_write(&swapon_rwsem); =20 atomic_inc(&proc_poll_event); wake_up_interruptible(&proc_poll_wait); @@ -3365,13 +3372,13 @@ static struct swap_info_struct *alloc_swap_info(voi= d) return ERR_PTR(-ENOMEM); } =20 - spin_lock(&swap_lock); + percpu_down_write(&swapon_rwsem); for (type =3D 0; type < nr_swapfiles; type++) { if (!(swap_info[type]->flags & SWP_USED)) break; } if (type >=3D MAX_SWAPFILES) { - spin_unlock(&swap_lock); + percpu_up_write(&swapon_rwsem); percpu_ref_exit(&p->users); kvfree(p); return ERR_PTR(-EPERM); @@ -3396,7 +3403,7 @@ static struct swap_info_struct *alloc_swap_info(void) plist_node_init(&p->list, 0); plist_node_init(&p->avail_list, 0); p->flags =3D SWP_USED; - spin_unlock(&swap_lock); + percpu_up_write(&swapon_rwsem); if (defer) { percpu_ref_exit(&defer->users); kvfree(defer); @@ -3829,9 +3836,9 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialf= ile, int, swap_flags) * Clear the SWP_USED flag after all resources are freed so * alloc_swap_info can reuse this si safely. */ - spin_lock(&swap_lock); + percpu_down_write(&swapon_rwsem); si->flags =3D 0; - spin_unlock(&swap_lock); + percpu_up_write(&swapon_rwsem); if (inced_nr_rotate_swap) atomic_dec(&nr_rotate_swap); if (swap_file) @@ -3849,14 +3856,14 @@ void si_swapinfo(struct sysinfo *val) struct swap_info_struct *si; unsigned long nr_to_be_unused =3D 0; =20 - spin_lock(&swap_lock); + percpu_down_read(&swapon_rwsem); for_each_swap(si) { if (!(si->flags & SWP_WRITEOK)) nr_to_be_unused +=3D swap_usage_in_pages(si); } val->freeswap =3D atomic_long_read(&nr_swap_pages) + nr_to_be_unused; val->totalswap =3D total_swap_pages + nr_to_be_unused; - spin_unlock(&swap_lock); + percpu_up_read(&swapon_rwsem); } =20 /* --=20 2.55.0 From nobody Sat Sep 26 22:02:19 2026 Received: from mail-pl1-f182.google.com (mail-pl1-f182.google.com [209.85.214.182]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7FAD33264C1 for ; Sat, 29 Aug 2026 07:47:44 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.182 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787989665; cv=none; b=nk+aaKdic1enumAWihN0SUAPEfDTaStrnDR/v1kU9uyfAqCt3KWp55TwVRjEdeWojGfcLXtsvwDt5rSmqnb6tpssGuWmuNJJVtamKtyHp5LYiXxs4I/KpNptk5FnkesH0yRQbIGHJFTLe05hM1ZvBDC3wx7K+XFNQoo1Hq66e24= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787989665; c=relaxed/simple; bh=rrUrHgsNZcNvQrQ4/aoZC1ACLWxdnMwtaUSsbxGf3M4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=AQ4sGE5n1+SYsjTUtL7veukoUG3CpxqYvJdoLvi/xWzMYzhyuK3BNQmpZ1Np6j33HjJFjZErHowmBBgNZFPfIIs5H89D57uUrmOYkdX0DWC8CitJDCzKOkzkr2GtxKiGDKaHyriIzVDj3AfcYPrJZaMCQY0XF15N/9I9TeAVf9s= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=rh/vAyP6; arc=none smtp.client-ip=209.85.214.182 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="rh/vAyP6" Received: by mail-pl1-f182.google.com with SMTP id d9443c01a7336-2d58efc7356so20786395ad.1 for ; Sat, 29 Aug 2026 00:47:44 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787989664; x=1788594464; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=iWn6EGMddmTMon7V+5jiUnHk1hdkEvtDIqVtgSXnKtY=; b=rh/vAyP6ZxvlAtc/8E/bSpuiekJ3tke2Ht2Xa76Zy11EMc8LMyok9wOotb4HpnK6xH siMF3wZ/9jGWM2zT22tveHm7DNUuTgH6xZY1uDfHcCn1M0quR+vep2ZW6mz/JRiPcXio 2AZy1UuIV0Z9Gemzuait7eY1kdSMUVrDtbYEZPCuJWQ79lTVV6Owfg0v21Q9xONJNorM lcsd1aBKkScopYJasY04qEkoxwVSZNFI3M2Y7RZUYfxK0Vwz2Df2io6NAJs8TLtNG/+C aq6e0jxSlQozJMBavscnE0yBxU1X3y/tPcMqCIgLy9uLsn1z2MSsirbxusUACBveR7yn YtCQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787989664; x=1788594464; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=iWn6EGMddmTMon7V+5jiUnHk1hdkEvtDIqVtgSXnKtY=; b=UI9DNCkaHiEAl5dE0qF2SfHwGwthI6lN7AsdWCYJp4z/VlIfSi/i9EA77BTT6DTt9D S5k5XXNC5fiVxfIF0LswkyJDmMBx0WdOZZa+VDQdDJvZT3O5+BCd09mOGtJBrAQl7mCU 2dfDVLpghr83/PIvmZLxYr3mDq52hs3x3J/HTrawTA4KMLlnJmwRPYKvmREADJh3MscU 5V7xnr+D84SalaGvHEdSIWczDECue9tSR2xdoVdgeLfJd+pD2liOJJNmAh/FW78nyuVu yDzTBPBIS0QIXtoItYLfBnZwKS11KRTspa9yDeYHsHSa3JoI82A+EBxzUkBlyxTcOYgA cRhw== X-Gm-Message-State: AFuF++mAy6rg8450z3uwZ+lNytQ1BgmPsSFa+YChD6x5yz30mhEV2pnd ywZ6OUYIL5ecENlDFTfuR47c0pMhQSn8HTjHMbLOuO7ChZjVdsDwQK1N X-Gm-Gg: AYBFou09lvuof8jf+zKGApwKI+8tNqvfIrOANUzqv4jTfHa+X26Pns8m8huTyHMRFVs 5mVYHCy9Jjy6poIkNNamYk8wGypTCGkHsckmKBLdSiKO6oV/Ov8grAdydOSw1KDZwRWuLyYFJ1C IBIxbJaaVtWHw2B20Q09pgPfHBfgBemnNr3lzhQ3PcNkQRHRAQH8/tbMhQBX7rBXtSPRM1WgUvf itisZL4pDGpP1duIulcvVvPtm1c83yoHY237xh2nJv2L2dnxo8XUG9tbe/m3AKizTk4Dr/ziVHF /FQyIM0lqotVBWWB6a0NlthIgVu/Vqp+wXbu9QQncSrRbM8h0iEp4R9Tn8Ct8HOZ4jO5kzWTGdu bBCMk6z0i/jEPurIu2HcyN0Wv7Kcu6xskg/Wujx+d6JSlEuWPl/B+oPp2T6hUokbRAjGwrnWlbG 4Nqm1gw6u19lQultIkXOF7Hrp5rHoo7bYAR1jWG2btkrFSUbqBhrWzzgKCzL7kRzgXQL5NV0R3J EEInaQRP3t6+6qu1FAUZWOuEZC0 X-Received: by 2002:a17:902:e804:b0:2ce:9c48:22d3 with SMTP id d9443c01a7336-2d74df05f9cmr199834015ad.11.1787989663724; Sat, 29 Aug 2026 00:47:43 -0700 (PDT) Received: from localhost.localdomain (vmi2317720.contaboserver.net. [84.247.152.65]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2d75986aea3sm12686395ad.52.2026.08.29.00.47.38 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Sat, 29 Aug 2026 00:47:43 -0700 (PDT) From: "Lian Wang (ProcessMission)" To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, Andrew Morton , Chris Li , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , Kemeng Shi , Kairui Song , Jihan LIN , Kunwu Chan Subject: [RESEND RFC PATCH v2 06/13] mm/swap: remove swapon mutex and update proc reader Date: Sat, 29 Aug 2026 15:47:35 +0800 Message-ID: <20260829-swap-pcp-priq-v2-resend-6-68d3d925578c@gmail.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260829-swap-pcp-priq-v2-resend-0-68d3d925578c@gmail.com> References: <20260829-swap-pcp-priq-v2-resend-0-68d3d925578c@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Kairui Song Now that swapon_rwsem protects all swapon/swapoff-sensitive data, the old swapon_mutex is redundant. Remove it and convert the /proc/swaps seq_file reader to use the rwsem instead. Also convert the swap_start iterator to use for_each_swap() for consistency. Also adapt swap_next, it needs an explicit SWP_USED check. swap_file is not reliably protected by swapon_rwsem, so checking SWP_USED is more accurate and stable for filtering in-use devices. Signed-off-by: Kairui Song Signed-off-by: Lian Wang (ProcessMission) Tested-by: Kunwu Chan Tested-by: Kunwu Chan --- mm/swapfile.c | 19 +++++-------------- 1 file changed, 5 insertions(+), 14 deletions(-) diff --git a/mm/swapfile.c b/mm/swapfile.c index 173375a7a215..6809c099eecc 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -140,9 +140,6 @@ static inline struct swap_info_struct *__swap_iter(int = *i, unsigned long flag) =20 static struct kmem_cache *swap_table_cachep; =20 -/* Protects si->swap_file for /proc/swaps usage */ -static DEFINE_MUTEX(swapon_mutex); - static DECLARE_WAIT_QUEUE_HEAD(proc_poll_wait); /* Activity counter to indicate that a swapon or swapoff has occurred */ static atomic_t proc_poll_event =3D ATOMIC_INIT(0); @@ -3174,7 +3171,6 @@ SYSCALL_DEFINE1(swapoff, const char __user *, special= file) if (!(p->flags & SWP_SOLIDSTATE)) atomic_dec(&nr_rotate_swap); =20 - mutex_lock(&swapon_mutex); percpu_down_write(&swapon_rwsem); spin_lock(&p->lock); drain_mmlist(); @@ -3189,7 +3185,6 @@ SYSCALL_DEFINE1(swapoff, const char __user *, special= file) percpu_up_write(&swapon_rwsem); arch_swap_invalidate_area(p->type); zswap_swapoff(p->type); - mutex_unlock(&swapon_mutex); kfree(p->global_cluster); p->global_cluster =3D NULL; free_swap_cluster_info(cluster_info, maxpages); @@ -3235,20 +3230,18 @@ static __poll_t swaps_poll(struct file *file, poll_= table *wait) return EPOLLIN | EPOLLRDNORM; } =20 -/* iterator */ static void *swap_start(struct seq_file *swap, loff_t *pos) { struct swap_info_struct *si; - int type; loff_t l =3D *pos; =20 - mutex_lock(&swapon_mutex); + percpu_down_read(&swapon_rwsem); =20 if (!l) return SEQ_START_TOKEN; =20 - for (type =3D 0; (si =3D swap_type_to_info(type)); type++) { - if (!(si->swap_file)) + for_each_swap(si) { + if (!si->swap_file) continue; if (!--l) return si; @@ -3269,7 +3262,7 @@ static void *swap_next(struct seq_file *swap, void *v= , loff_t *pos) =20 ++(*pos); for (; (si =3D swap_type_to_info(type)); type++) { - if (!(si->swap_file)) + if (!(si->flags & SWP_USED) || !(si->swap_file)) continue; return si; } @@ -3279,7 +3272,7 @@ static void *swap_next(struct seq_file *swap, void *v= , loff_t *pos) =20 static void swap_stop(struct seq_file *swap, void *v) { - mutex_unlock(&swapon_mutex); + percpu_up_read(&swapon_rwsem); } =20 static int swap_show(struct seq_file *swap, void *v) @@ -3789,7 +3782,6 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialf= ile, int, swap_flags) goto free_swap_zswap; } =20 - mutex_lock(&swapon_mutex); prio =3D DEF_SWAP_PRIO; if (swap_flags & SWAP_FLAG_PREFER) prio =3D swap_flags & SWAP_FLAG_PRIO_MASK; @@ -3815,7 +3807,6 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialf= ile, int, swap_flags) (si->flags & SWP_AREA_DISCARD) ? "s" : "", (si->flags & SWP_PAGE_DISCARD) ? "c" : ""); =20 - mutex_unlock(&swapon_mutex); atomic_inc(&proc_poll_event); wake_up_interruptible(&proc_poll_wait); =20 --=20 2.55.0 From nobody Sat Sep 26 22:02:19 2026 Received: from mail-pf1-f182.google.com (mail-pf1-f182.google.com [209.85.210.182]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7A758399036 for ; Sat, 29 Aug 2026 07:47:54 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.182 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787989676; cv=none; b=qJ9uzH28rOa5vrzmSHyCclU9tQd7rce9rozxr2inS3l6ArantWa3ikPIVql9eGKuJCWRAHk3KtZat2wDjpqNLCNforQGUQDfFIqOJ5VQBs6qhL6zOZZfXLD2zZ+zDtHlFeWUy21kq9eqd+SpJz2HKeSB0HffCrWfmVO81regW2E= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787989676; c=relaxed/simple; bh=dMc3mNca30GEB/rfUITzpe5+BKZBQefCWx5qUlrHu5U=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=eG0GxCY2BzlkjvJGmFXeXdtyweHb5OOL4o7bDuSF6XUtVw3psZkmxw4qaoSrJLALKn3Aglt9eWJ0JJWIRbT9eWvgnFqqSQzTkvPBuR/qVvkSG4LnYWLPPXQnhauxWpUtfVif2PEPXMKGNlIK0Hr0CiRN6KJPJOpQXAzQCzjz8z0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=c/Pc5yVI; arc=none smtp.client-ip=209.85.210.182 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="c/Pc5yVI" Received: by mail-pf1-f182.google.com with SMTP id d2e1a72fcca58-853e2610bb4so2043603b3a.0 for ; Sat, 29 Aug 2026 00:47:54 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787989674; x=1788594474; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=7vGEGo1ju5lk19uFkoLxptYJ6h4M8dZ1YyQasiBTx1o=; b=c/Pc5yVIlFfj5+aRIiVot2qFlUco28PM0nOqxQa9jAL//t9YB9GJlfdvSWb+fr97LH urtReq5ELi4/Kuv8//ct+TqSW5CnX0HoyqySO05ldaWT+0QQt+Kt+oWNr3mmwbCkqtvx fnydlNkso0mdIuzLOKhXH4swzi676+lORK0TlrxE7E3iHD5Qud/bcgRT1liRiEZX17D4 Wd/5Hliwxb3JSDSlVm31jL/+t6ZGBncXB1Nytr3ljuMfwLVO4nwt4qMBsH9NRbP9yNw8 1+Hm71VV8cbyqnWMqC6xvXOpXqgCr4ddsxh8IGlz0lZNQr3le4HmTTleZx2zWvssCH5o VUvA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787989674; x=1788594474; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=7vGEGo1ju5lk19uFkoLxptYJ6h4M8dZ1YyQasiBTx1o=; b=awcTt/dB5kn2CBd1ptFyf7hxCju/yI/jAZKPFYRIlRouOpW32q/9HVEyNKsyxjCB4B w91FKy6CstSM2dbegJOeU2IMIUK09KPehiRwYz38I9v/J6aSra+tWOwE8mJ+XSektcmR cZLsugdkruWZJROwBY2T7IPxkYVNsw+fcRvbmNsXlTq8AKW9NdDw9m+UtHIs7SYGbyr1 26S92aZt42afx05Ee2dCBLWDQZ0WrgFw8eDfWM2FD4K5+6UZ3PGArKlEoM4dBH0thvWV e/OimII+7eyoD5VmmfB4qT2vANzKkULvy4dKqxKssR9yXTpHUb1SiVNg+vLKEpPwazao kH0A== X-Gm-Message-State: AFuF++m9cQyCCzOoziKrzIbRvjSY4WOnrgSr/lbZ+RYcUmiqKDWs0gvW BaXMcKz3vCtcS+agq2IcM/tcaQTv5svZoFTd1+pGoWNB8yeJUn7iT9FK X-Gm-Gg: AR+sD10Y5oOfAFee+MXbaqL4pwmcyAAGFkYGUGZ/GoHu5K2edB3fp5vX49zrtVjt57H InH9zQyfNiWp7DwUVX/VmcW5NAPejQiE6TfxLGjM5Af+snkVPifQuXbxUBMIczDyRpABNqCUcuO IEJ5d1KcER0rC9Wz/ow0l2yKNm6r++ieWADFnICYwJwfHlpVmoR06qvPsVdw66rYvY74fP7yvdR 5giWVT647Eed0gjY4LHhc3+V9DYjy4PDFMctL/Qm837KufSmWodYYOSKnwKJ3zc7SBAmHQbDAWg PhnITHx2FDU1b4ZJYHFE6hiTtwLPrfXmOFro1sE1cmoA2i9iTPSCvKmMYOXxUN4LINFAFrQq/0V 4PT8xUPw2S8HpNz9AhGPF+26180EGY9oW3CArtiFTQiSf+H0WkKHrNgVrZS0JhHFddpB7knQER4 gzIz3ZinaAN4ILpkFGPUR6F/r1+qIVkLaAX1BbSxjy2UaayxXL3brdl9Q80F9L0uuF4JqY7tZbj uO9BHphThBLiHlRzuGW9ROFx8Xn X-Received: by 2002:a05:6a00:ae07:b0:842:63f5:d097 with SMTP id d2e1a72fcca58-854c6485ba9mr21043098b3a.3.1787989673723; Sat, 29 Aug 2026 00:47:53 -0700 (PDT) Received: from localhost.localdomain (vmi2317720.contaboserver.net. [84.247.152.65]) by smtp.gmail.com with ESMTPSA id d2e1a72fcca58-8569f49ed7asm1340256b3a.8.2026.08.29.00.47.48 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Sat, 29 Aug 2026 00:47:53 -0700 (PDT) From: "Lian Wang (ProcessMission)" To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, Andrew Morton , Chris Li , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , Kemeng Shi , Kairui Song , Jihan LIN , Kunwu Chan Subject: [RESEND RFC PATCH v2 07/13] mm/swap: consolidate swap inuse accounting helpers Date: Sat, 29 Aug 2026 15:47:44 +0800 Message-ID: <20260829-swap-pcp-priq-v2-resend-7-68d3d925578c@gmail.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260829-swap-pcp-priq-v2-resend-0-68d3d925578c@gmail.com> References: <20260829-swap-pcp-priq-v2-resend-0-68d3d925578c@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Kairui Song Inline the swap_usage_add/sub helpers into their only callers and rename the callers to better reflect what they actually do: track per-device inuse page counts. No functional change. Signed-off-by: Kairui Song Signed-off-by: Lian Wang (ProcessMission) Tested-by: Kunwu Chan Tested-by: Kunwu Chan --- mm/swapfile.c | 64 ++++++++++++++++++++++----------------------------- 1 file changed, 27 insertions(+), 37 deletions(-) diff --git a/mm/swapfile.c b/mm/swapfile.c index 6809c099eecc..1fa6945d6415 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -49,8 +49,8 @@ #include "internal.h" #include "swap.h" =20 -static void swap_range_alloc(struct swap_info_struct *si, - unsigned int nr_entries); +static void swap_device_inuse_add(struct swap_info_struct *si, + unsigned int nr_entries); static bool folio_swapcache_freeable(struct folio *folio); static void move_cluster(struct swap_info_struct *si, struct swap_cluster_info *ci, struct list_head *list, @@ -984,7 +984,7 @@ static bool __swap_cluster_alloc_entries(struct swap_in= fo_struct *si, if (cluster_is_empty(ci)) ci->order =3D order; ci->count +=3D nr_pages; - swap_range_alloc(si, nr_pages); + swap_device_inuse_add(si, nr_pages); =20 return true; } @@ -1292,53 +1292,36 @@ static void add_to_avail_list(struct swap_info_stru= ct *si) } =20 /* - * swap_usage_add / swap_usage_sub of each slot are serialized by ci->lock - * within each cluster, so the total contribution to the global counter sh= ould - * always be positive and cannot exceed the total number of usable slots. + * swap_device_inuse_add / swap_device_inuse_sub are called within cluster + * update critical sections and serialized by ci->lock within each cluster, + * so the total contribution to the global counter should always be positi= ve + * and cannot exceed the total number of usable slots. */ -static bool swap_usage_add(struct swap_info_struct *si, unsigned int nr_en= tries) +static void swap_device_inuse_add(struct swap_info_struct *si, + unsigned int nr_entries) { - long val =3D atomic_long_add_return_relaxed(nr_entries, &si->inuse_pages); + long inuse_pages; =20 /* * If device is full, and SWAP_USAGE_OFFLIST_BIT is not set, * remove it from the plist. */ - if (unlikely(val =3D=3D si->pages)) { - del_from_avail_list(si, false); - return true; - } - - return false; -} - -static void swap_usage_sub(struct swap_info_struct *si, unsigned int nr_en= tries) -{ - long val =3D atomic_long_sub_return_relaxed(nr_entries, &si->inuse_pages); - - /* - * If device is not full, and SWAP_USAGE_OFFLIST_BIT is set, - * add it to the plist. - */ - if (unlikely(val & SWAP_USAGE_OFFLIST_BIT)) - add_to_avail_list(si); -} - -static void swap_range_alloc(struct swap_info_struct *si, - unsigned int nr_entries) -{ - if (swap_usage_add(si, nr_entries)) { + inuse_pages =3D atomic_long_add_return_relaxed(nr_entries, &si->inuse_pag= es); + if (unlikely(inuse_pages =3D=3D si->pages)) { if (vm_swap_full()) schedule_work(&si->reclaim_work); + del_from_avail_list(si, false); } + atomic_long_sub(nr_entries, &nr_swap_pages); } =20 -static void swap_range_free(struct swap_info_struct *si, unsigned long off= set, - unsigned int nr_entries) +static void swap_device_inuse_sub(struct swap_info_struct *si, unsigned lo= ng offset, + unsigned int nr_entries) { unsigned long end =3D offset + nr_entries - 1; void (*swap_slot_free_notify)(struct block_device *, unsigned long); + long inuse_pages; unsigned int i; =20 for (i =3D 0; i < nr_entries; i++) @@ -1362,7 +1345,14 @@ static void swap_range_free(struct swap_info_struct = *si, unsigned long offset, */ smp_wmb(); atomic_long_add(nr_entries, &nr_swap_pages); - swap_usage_sub(si, nr_entries); + + /* + * If device is not full, and SWAP_USAGE_OFFLIST_BIT is set, + * add it back to the plist. + */ + inuse_pages =3D atomic_long_sub_return_relaxed(nr_entries, &si->inuse_pag= es); + if (unlikely(inuse_pages & SWAP_USAGE_OFFLIST_BIT)) + add_to_avail_list(si); } =20 static bool get_swap_device_info(struct swap_info_struct *si) @@ -1975,7 +1965,7 @@ void __swap_cluster_free_entries(struct swap_info_str= uct *si, if (batch_id) mem_cgroup_uncharge_swap(batch_id, ci_off - batch_off); =20 - swap_range_free(si, ci_head + ci_start, nr_pages); + swap_device_inuse_sub(si, ci_head + ci_start, nr_pages); swap_cluster_assert_empty(ci, ci_start, nr_pages, false); =20 if (!ci->count) @@ -2834,7 +2824,7 @@ static int try_to_unuse(unsigned int type) success: /* * Make sure that further cleanups after try_to_unuse() returns happen - * after swap_range_free() reduces si->inuse_pages to 0. + * after swap_device_inuse_sub() reduces si->inuse_pages to 0. */ smp_mb(); return 0; --=20 2.55.0 From nobody Sat Sep 26 22:02:19 2026 Received: from mail-pl1-f169.google.com (mail-pl1-f169.google.com [209.85.214.169]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CB0C2311C2A for ; Sat, 29 Aug 2026 07:48:03 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.169 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787989685; cv=none; b=B7fEvx86QuIHTcBbL8ZzsYEDOERc0uvGK8mtapd+gQs4gVcc5uL2BmJT+xFkfojHMq3EOtwzq+dTlNFZYK5oRyizH8JDHfoQSsvoTBeGuqyhic7yKrvOxq78ovvA0YHr0SyAN6A1cRUu7TtCws6PvC+tKjOVYk+j8Gl3ABF7ZTY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787989685; c=relaxed/simple; bh=SfeUrH+0uWX7Hs71vNWmXYEDpQydR2VFcPQxTnmr1Io=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=sAtxhSBBQXe0/iTT2SZsDuIW4VifWPxbNGscRx6al4nkkSYoJ3+a8/ditoIHwWQSdNx4lKrMKbU5YpCJj5Ww3A2dw1Hgs+oc2JKfF1E4av9cEzGNqaQZVWUsOC/dc+bkGY1tDsRjA/NGc5lsNnLiF+2MZl8swN1M0NFN3JCBykw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=mzhnENaI; arc=none smtp.client-ip=209.85.214.169 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="mzhnENaI" Received: by mail-pl1-f169.google.com with SMTP id d9443c01a7336-2cedda2ce6fso11969685ad.1 for ; Sat, 29 Aug 2026 00:48:03 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787989683; x=1788594483; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=e3pO0sfpG3DfTZP/kIIBTuTwDNPHKJvFu4PoPkWSsN0=; b=mzhnENaIOL4WZjcFI1/8I4D+Wj1SBatIT23y6UBiIf/++nSXzvmmfBhS/eWoOivBds W78tP9UUuXoLSG6tcLG3XDDz9ghFV1ebqEB1QuKQuccNOXBrkEdsfGeWzwhLebtzprrC cY0yzM0w50Fl58R3Gh6rU98mAQzMRoP5jf4DiIgVG6J++cAVHKf/DNNZc4Wd8dK8YvzI iWdRg5mTothM3eKfvHleGXgQblqGqWNuC4UaLbsNpv4Rrmd1wY07lJMR52T6RwSyVc2m xXKXNB/twoi3JscPkSUcJZDGK5rpjaN8rKDJecdX74yjIQ/bUuHKfeKPMeU4nXL3xuR4 umzA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787989683; x=1788594483; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=e3pO0sfpG3DfTZP/kIIBTuTwDNPHKJvFu4PoPkWSsN0=; b=DfTikelrXOK8PvDApFk+Vp/3R0WpPWXLmwETUbrKTFP8dlUJXwavQm7DH6Wn1Gta04 bRI2hVMFIlR2gGDnOBBicHkp0yop+hAEz7tffyVlQT+edjxYTt07XZCIljJdzvRh2Rqq wIy8qdh8tGI1eLmH2ZIPEMcBRamDSqTxMVndcYnPW+w6dzdiSv0nhvrjjDdLwbBFT1Hw uNDSFNWaygCprKgOZX77sEaN5QuRQzgOM/8f6w34a7rCX3qustAKNEZtSJAABLIZcC42 u5ZhjiP9J0EJSOZ43ch8zyexgQ/sAjIu46orlEWACrqcoVUWH0A0l3NWMtlYVOHVuF2b KyLA== X-Gm-Message-State: AFuF++ntH/pSB2YhVVb9BKXUKdVUCe9q6yiJ2AHT1+pUFK86o2uQnGAz liqtfq4+VlnOonse9gf+ZPyyzq/WkMgUPpzvGSFGWbsBtZb9u0fzbzB/ X-Gm-Gg: AYBFou187wSWZNA/UW5z+7vIs2bysjV2kSHrDmTIyusvAnF/921M0OWAUgSQp69/nN+ sSLxrRTvR9FXiKIIGl1TkTPiApDcmhWQoALsfS07YoyWl13G4o2Mz7JutxC4Tsz8hg8TS2T8LCt WcOXzhIDJuHujDwR1ywNBxSlVsDuGtBs+JVjrv/Z7QyPNw7C5xW8IVuNc8qySA/PoCLjPOSqb4U qvuzBsSfLwuztsejl4EXJ9c+4DjPmCS5uNyqwMl/s85ockTJGTfj1e0dTpYyCAhHB8myNRwwBCh 50twq+CWX+g6peId2pn4AFvbJmgE7vdDNJMpLkdWEq3VNaf1cRYWIWSMXlNZG4M0foREYr87Wz+ E3GBrEeTqFlQGw2EplFJSdiYxAHjUH4UODpkcgU+NwXBtXhJiDLy89sM7KSAdUSKLEdsjOGl6w8 fbpeMFcL0k8+TJdBDDgBLR02eKFFcRTGPPb7Wg13ze7J2zz7JuTRl/ufcV8wyntBgU+Y4+NuClU CYkWcG1egUuAv3NY0Ls35G0q2GZ X-Received: by 2002:a17:902:da89:b0:2d8:d4d3:da52 with SMTP id d9443c01a7336-2d8d4d3db4cmr65289365ad.22.1787989682717; Sat, 29 Aug 2026 00:48:02 -0700 (PDT) Received: from localhost.localdomain (vmi2317720.contaboserver.net. [84.247.152.65]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2d7594fb1c3sm12506505ad.1.2026.08.29.00.47.57 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Sat, 29 Aug 2026 00:48:02 -0700 (PDT) From: "Lian Wang (ProcessMission)" To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, Andrew Morton , Chris Li , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , Kemeng Shi , Kairui Song , Jihan LIN , Kunwu Chan , Baoquan He Subject: [RESEND RFC PATCH v2 08/13] mm/swap: change back to use each swap device's percpu cluster Date: Sat, 29 Aug 2026 15:47:54 +0800 Message-ID: <20260829-swap-pcp-priq-v2-resend-8-68d3d925578c@gmail.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260829-swap-pcp-priq-v2-resend-0-68d3d925578c@gmail.com> References: <20260829-swap-pcp-priq-v2-resend-0-68d3d925578c@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Youngjun Park This reverts commit 1b7e90020eb7 ("mm, swap: use percpu cluster as allocation fast path"). The global percpu cluster causes several issues: 1) It prevents efficient swap allocation for users that do not want the default swap device policy. Because the cluster cache sits above device selection, all allocation users must stick with the default allocation policy. This fundamentally conflicts with incoming ideas like swap tiering. 2) It can cause priority inversion. Consider two memcgs: memcg1 can access devices A and B, where A has higher priority; memcg2 can only access B. Memcg2 can write the global percpu cluster with device B, then memcg1 takes B in the fast path even though higher-priority device A is not exhausted. 3) The fast-path / slow-path design bundled with the global per-CPU cache is problematic: only the fast path uses the global cache, and the slow path is forced to rotate the allocation list. This complicates consumers like discard. Revert to per-device percpu clusters first. A later patch will introduce new infrastructure for a high-performance global allocation path. Suggested-by: Kairui Song Co-developed-by: Baoquan He Signed-off-by: Baoquan He Signed-off-by: Youngjun Park Signed-off-by: Kairui Song Signed-off-by: Lian Wang (ProcessMission) Tested-by: Kunwu Chan Tested-by: Kunwu Chan --- include/linux/swap.h | 13 ++- mm/swapfile.c | 184 +++++++++++++++---------------------------- 2 files changed, 72 insertions(+), 125 deletions(-) diff --git a/include/linux/swap.h b/include/linux/swap.h index 7b8ba0d1903f..aa66d8454186 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -245,10 +245,17 @@ enum { #endif =20 /* - * We keep using same cluster for rotational device so IO will be sequenti= al. - * The purpose is to optimize SWAP throughput on these device. + * We assign a cluster to each CPU, so each CPU can allocate swap entry fr= om + * its own cluster and swapout sequentially. The purpose is to optimize sw= apout + * throughput. */ +struct percpu_cluster { + local_lock_t lock; /* Protect the percpu_cluster above */ + unsigned int next[SWAP_NR_ORDERS]; /* Likely next allocation offset */ +}; + struct swap_sequential_cluster { + spinlock_t lock; /* Serialize usage of global cluster */ unsigned int next[SWAP_NR_ORDERS]; /* Likely next allocation offset */ }; =20 @@ -271,8 +278,8 @@ struct swap_info_struct { /* list of cluster that are fragmented or contented */ unsigned int pages; /* total of usable pages of swap */ atomic_long_t inuse_pages; /* number of those currently in use */ + struct percpu_cluster __percpu *percpu_cluster; /* per cpu's swap locatio= n */ struct swap_sequential_cluster *global_cluster; /* Use one global cluster= for rotating device */ - spinlock_t global_cluster_lock; /* Serialize usage of global cluster */ struct rb_root swap_extent_root;/* root of the swap extent rbtree */ struct block_device *bdev; /* swap device or bdev of swap file */ struct file *swap_file; /* seldom referenced */ diff --git a/mm/swapfile.c b/mm/swapfile.c index 1fa6945d6415..79ecff2d0bd3 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -146,18 +146,6 @@ static atomic_t proc_poll_event =3D ATOMIC_INIT(0); =20 atomic_t nr_rotate_swap =3D ATOMIC_INIT(0); =20 -struct percpu_swap_cluster { - struct swap_info_struct *si[SWAP_NR_ORDERS]; - unsigned long offset[SWAP_NR_ORDERS]; - local_lock_t lock; -}; - -static DEFINE_PER_CPU(struct percpu_swap_cluster, percpu_swap_cluster) =3D= { - .si =3D { NULL }, - .offset =3D { SWAP_ENTRY_INVALID }, - .lock =3D INIT_LOCAL_LOCK(), -}; - /* May return NULL on invalid type, caller must check for NULL return */ static struct swap_info_struct *swap_type_to_info(int type) { @@ -562,9 +550,10 @@ swap_cluster_populate(struct swap_info_struct *si, * Only cluster isolation from the allocator does table allocation. * Swap allocator uses percpu clusters and holds the local lock. */ - lockdep_assert_held(&this_cpu_ptr(&percpu_swap_cluster)->lock); - if (!(si->flags & SWP_SOLIDSTATE)) - lockdep_assert_held(&si->global_cluster_lock); + if (si->flags & SWP_SOLIDSTATE) + lockdep_assert_held(this_cpu_ptr(&si->percpu_cluster->lock)); + else + lockdep_assert_held(&si->global_cluster->lock); lockdep_assert_held(&ci->lock); =20 if (!swap_cluster_alloc_table(ci, __GFP_HIGH | __GFP_NOMEMALLOC | @@ -577,9 +566,10 @@ swap_cluster_populate(struct swap_info_struct *si, * the potential recursive allocation is limited. */ spin_unlock(&ci->lock); - if (!(si->flags & SWP_SOLIDSTATE)) - spin_unlock(&si->global_cluster_lock); - local_unlock(&percpu_swap_cluster.lock); + if (si->flags & SWP_SOLIDSTATE) + local_unlock(&si->percpu_cluster->lock); + else + spin_unlock(&si->global_cluster->lock); =20 ret =3D swap_cluster_alloc_table(ci, __GFP_HIGH | __GFP_NOMEMALLOC | GFP_KERNEL); @@ -592,9 +582,10 @@ swap_cluster_populate(struct swap_info_struct *si, * could happen with ignoring the percpu cluster is fragmentation, * which is acceptable since this fallback and race is rare. */ - local_lock(&percpu_swap_cluster.lock); - if (!(si->flags & SWP_SOLIDSTATE)) - spin_lock(&si->global_cluster_lock); + if (si->flags & SWP_SOLIDSTATE) + local_lock(&si->percpu_cluster->lock); + else + spin_lock(&si->global_cluster->lock); spin_lock(&ci->lock); =20 if (ret) { @@ -700,7 +691,7 @@ static bool swap_do_scheduled_discard(struct swap_info_= struct *si) ci =3D list_first_entry(&si->discard_clusters, struct swap_cluster_info,= list); /* * Delete the cluster from list to prepare for discard, but keep - * the CLUSTER_FLAG_DISCARD flag, percpu_swap_cluster could be + * the CLUSTER_FLAG_DISCARD flag, there could be percpu_cluster * pointing to it, or ran into by relocate_cluster. */ list_del(&ci->list); @@ -1032,12 +1023,10 @@ static unsigned int alloc_swap_scan_cluster(struct = swap_info_struct *si, out: relocate_cluster(si, ci); swap_cluster_unlock(ci); - if (si->flags & SWP_SOLIDSTATE) { - this_cpu_write(percpu_swap_cluster.offset[order], next); - this_cpu_write(percpu_swap_cluster.si[order], si); - } else { + if (si->flags & SWP_SOLIDSTATE) + this_cpu_write(si->percpu_cluster->next[order], next); + else si->global_cluster->next[order] =3D next; - } return found; } =20 @@ -1138,13 +1127,17 @@ static unsigned long cluster_alloc_swap_entry(struc= t swap_info_struct *si, if (order && !(si->flags & SWP_BLKDEV)) return 0; =20 - if (!(si->flags & SWP_SOLIDSTATE)) { + if (si->flags & SWP_SOLIDSTATE) { + /* Fast path using per CPU cluster */ + local_lock(&si->percpu_cluster->lock); + offset =3D __this_cpu_read(si->percpu_cluster->next[order]); + } else { /* Serialize HDD SWAP allocation for each device. */ - spin_lock(&si->global_cluster_lock); + spin_lock(&si->global_cluster->lock); offset =3D si->global_cluster->next[order]; - if (offset =3D=3D SWAP_ENTRY_INVALID) - goto new_cluster; + } =20 + if (offset !=3D SWAP_ENTRY_INVALID) { ci =3D swap_cluster_lock(si, offset); /* Cluster could have been used by another order */ if (cluster_is_usable(ci, order)) { @@ -1158,7 +1151,6 @@ static unsigned long cluster_alloc_swap_entry(struct = swap_info_struct *si, goto done; } =20 -new_cluster: /* * If the device need discard, prefer new cluster over nonfull * to spread out the writes. @@ -1215,8 +1207,10 @@ static unsigned long cluster_alloc_swap_entry(struct= swap_info_struct *si, goto done; } done: - if (!(si->flags & SWP_SOLIDSTATE)) - spin_unlock(&si->global_cluster_lock); + if (si->flags & SWP_SOLIDSTATE) + local_unlock(&si->percpu_cluster->lock); + else + spin_unlock(&si->global_cluster->lock); =20 return found; } @@ -1371,41 +1365,8 @@ static bool get_swap_device_info(struct swap_info_st= ruct *si) return true; } =20 -/* - * Fast path try to get swap entries with specified order from current - * CPU's swap entry pool (a cluster). - */ -static bool swap_alloc_fast(struct folio *folio) -{ - unsigned int order =3D folio_order(folio); - struct swap_cluster_info *ci; - struct swap_info_struct *si; - unsigned int offset; - - /* - * Once allocated, swap_info_struct will never be completely freed, - * so checking it's liveness by get_swap_device_info is enough. - */ - si =3D this_cpu_read(percpu_swap_cluster.si[order]); - offset =3D this_cpu_read(percpu_swap_cluster.offset[order]); - if (!si || !offset || !get_swap_device_info(si)) - return false; - - ci =3D swap_cluster_lock(si, offset); - if (cluster_is_usable(ci, order)) { - if (cluster_is_empty(ci)) - offset =3D cluster_offset(si, ci); - alloc_swap_scan_cluster(si, ci, folio, offset); - } else { - swap_cluster_unlock(ci); - } - - put_swap_device(si); - return folio_test_swapcache(folio); -} - /* Rotate the device and switch to a new cluster */ -static void swap_alloc_slow(struct folio *folio) +static void swap_alloc_entry(struct folio *folio) { struct swap_info_struct *si, *next; =20 @@ -1775,10 +1736,7 @@ int folio_alloc_swap(struct folio *folio) } =20 again: - local_lock(&percpu_swap_cluster.lock); - if (!swap_alloc_fast(folio)) - swap_alloc_slow(folio); - local_unlock(&percpu_swap_cluster.lock); + swap_alloc_entry(folio); =20 if (!order && unlikely(!folio_test_swapcache(folio))) { if (swap_sync_discard()) @@ -2167,31 +2125,14 @@ void swap_put_entries_direct(swp_entry_t entry, int= nr) */ swp_entry_t swap_alloc_hibernation_slot(int type) { - struct swap_info_struct *pcp_si, *si =3D swap_type_to_info(type); - unsigned long pcp_offset, offset =3D SWAP_ENTRY_INVALID; - struct swap_cluster_info *ci; + struct swap_info_struct *si =3D swap_type_to_info(type); + unsigned long offset =3D SWAP_ENTRY_INVALID; swp_entry_t entry =3D {0}; =20 if (!si) goto fail; =20 - /* - * Try the local cluster first if it matches the device. If - * not, try grab a new cluster and override local cluster. - */ - local_lock(&percpu_swap_cluster.lock); - pcp_si =3D this_cpu_read(percpu_swap_cluster.si[0]); - pcp_offset =3D this_cpu_read(percpu_swap_cluster.offset[0]); - if (pcp_si =3D=3D si && pcp_offset) { - ci =3D swap_cluster_lock(si, pcp_offset); - if (cluster_is_usable(ci, 0)) - offset =3D alloc_swap_scan_cluster(si, ci, NULL, pcp_offset); - else - swap_cluster_unlock(ci); - } - if (!offset) - offset =3D cluster_alloc_swap_entry(si, NULL); - local_unlock(&percpu_swap_cluster.lock); + offset =3D cluster_alloc_swap_entry(si, NULL); if (offset) entry =3D swp_entry(si->type, offset); =20 @@ -3082,27 +3023,6 @@ static void free_swap_cluster_info(struct swap_clust= er_info *cluster_info, kvfree(cluster_info); } =20 -/* - * Called after swap device's reference count is dead, so - * neither scan nor allocation will use it. - */ -static void flush_percpu_swap_cluster(struct swap_info_struct *si) -{ - int cpu, i; - struct swap_info_struct **pcp_si; - - for_each_possible_cpu(cpu) { - pcp_si =3D per_cpu_ptr(percpu_swap_cluster.si, cpu); - /* - * Invalidate the percpu swap cluster cache, si->users - * is dead, so no new user will point to it, just flush - * any existing user. - */ - for (i =3D 0; i < SWAP_NR_ORDERS; i++) - cmpxchg(&pcp_si[i], si, NULL); - } -} - SYSCALL_DEFINE1(swapoff, const char __user *, specialfile) { struct swap_info_struct *p =3D NULL; @@ -3154,7 +3074,6 @@ SYSCALL_DEFINE1(swapoff, const char __user *, special= file) =20 flush_work(&p->discard_work); flush_work(&p->reclaim_work); - flush_percpu_swap_cluster(p); =20 destroy_swap_extents(p, p->swap_file); =20 @@ -3175,6 +3094,8 @@ SYSCALL_DEFINE1(swapoff, const char __user *, special= file) percpu_up_write(&swapon_rwsem); arch_swap_invalidate_area(p->type); zswap_swapoff(p->type); + free_percpu(p->percpu_cluster); + p->percpu_cluster =3D NULL; kfree(p->global_cluster); p->global_cluster =3D NULL; free_swap_cluster_info(cluster_info, maxpages); @@ -3523,7 +3444,7 @@ static int setup_swap_clusters_info(struct swap_info_= struct *si, { unsigned long nr_clusters =3D DIV_ROUND_UP(maxpages, SWAPFILE_CLUSTER); struct swap_cluster_info *cluster_info; - int err =3D -ENOMEM; + int cpu, err =3D -ENOMEM; unsigned long i; =20 cluster_info =3D kvzalloc_objs(*cluster_info, nr_clusters); @@ -3533,13 +3454,26 @@ static int setup_swap_clusters_info(struct swap_inf= o_struct *si, for (i =3D 0; i < nr_clusters; i++) spin_lock_init(&cluster_info[i].lock); =20 - if (!(si->flags & SWP_SOLIDSTATE)) { + if (si->flags & SWP_SOLIDSTATE) { + si->percpu_cluster =3D alloc_percpu(struct percpu_cluster); + if (!si->percpu_cluster) + goto err; + + for_each_possible_cpu(cpu) { + struct percpu_cluster *cluster; + + cluster =3D per_cpu_ptr(si->percpu_cluster, cpu); + for (i =3D 0; i < SWAP_NR_ORDERS; i++) + cluster->next[i] =3D SWAP_ENTRY_INVALID; + local_lock_init(&cluster->lock); + } + } else { si->global_cluster =3D kmalloc_obj(*si->global_cluster); if (!si->global_cluster) goto err; for (i =3D 0; i < SWAP_NR_ORDERS; i++) si->global_cluster->next[i] =3D SWAP_ENTRY_INVALID; - spin_lock_init(&si->global_cluster_lock); + spin_lock_init(&si->global_cluster->lock); } =20 /* @@ -3708,11 +3642,6 @@ SYSCALL_DEFINE2(swapon, const char __user *, special= file, int, swap_flags) =20 maxpages =3D si->max; =20 - /* Set up the swap cluster info */ - error =3D setup_swap_clusters_info(si, swap_header, maxpages); - if (error) - goto bad_swap_unlock_inode; - if (si->bdev && bdev_stable_writes(si->bdev)) si->flags |=3D SWP_STABLE_WRITES; =20 @@ -3726,6 +3655,15 @@ SYSCALL_DEFINE2(swapon, const char __user *, special= file, int, swap_flags) inced_nr_rotate_swap =3D true; } =20 + /* + * Set up the swap cluster info. This must run after SWP_SOLIDSTATE + * is determined above, as it decides whether to allocate the per-CPU + * cluster (solid state) or the global cluster (rotational). + */ + error =3D setup_swap_clusters_info(si, swap_header, maxpages); + if (error) + goto bad_swap_unlock_inode; + if ((swap_flags & SWAP_FLAG_DISCARD) && si->bdev && bdev_max_discard_sectors(si->bdev)) { /* @@ -3807,6 +3745,8 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialf= ile, int, swap_flags) bad_swap_unlock_inode: inode_unlock(inode); bad_swap: + free_percpu(si->percpu_cluster); + si->percpu_cluster =3D NULL; kfree(si->global_cluster); si->global_cluster =3D NULL; inode =3D NULL; --=20 2.55.0 From nobody Sat Sep 26 22:02:19 2026 Received: from mail-pj1-f54.google.com (mail-pj1-f54.google.com [209.85.216.54]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8F3FF35E1CE for ; Sat, 29 Aug 2026 07:48:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.54 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787989694; cv=none; b=qo2kMHBHs6iCEy+ExIJBLENWVb6XxQZUB1s5mEeSgs9OTetlDiwblNjxl1UEFFFgCekoKEHW57z0bFP4YoyH5ZteHzH0RUy0m8ghdnQFPNe1PuC/aqE0qY5WD10V4RWUGrJ0RQySrTSmTNll9F8o5zZrbipi3sFC2sLjvuvH0qk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787989694; c=relaxed/simple; bh=tkQ+mPYpv33KOBzZsKiJtl42QmUJg8qcwq+mqggj8uc=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=TEPk8mOaO0wSfazgQmACOn4PyqYR9fytbuHmBUv4KqcdbV7rGWR1vgk/3BSlh3XVGeinawXFYaFlQRhFZ4zix8Aq8etPEpu+w/Fb4DCw6ZR5NQRq9cxb9dJ+snc8fBxUrLsGwrcKh+eYSMm45uAEjSyDYjWO9EdOwwoFBIEcPy0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=irXDkGOB; arc=none smtp.client-ip=209.85.216.54 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="irXDkGOB" Received: by mail-pj1-f54.google.com with SMTP id 98e67ed59e1d1-3965d3d9ab8so1341082a91.3 for ; Sat, 29 Aug 2026 00:48:12 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787989692; x=1788594492; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=LNuXN9S7i6TqJ8ICcOb8D2g40X31e66OWo2XPTKLXSs=; b=irXDkGOBe2bTsZQ4soTTYQPphfnKAb4NsTOyOMDZ4+psn4CYZI+ec/lPRgY/oSxukU vTniRW4JI3CcS4BVyEyd6v4mvCWcmyHqT95l+LKGJC0y+JUGsExhp7nvkzvYYzVMdaHZ szMqOgo4iIfrJB0nvZK8eAh8J1tfRwyi8QMW7/2gmyoM10pb5ypAUreD18HzqGJaH9IG 4oByHwZFennJA9BGaBLeSuMLFPM09RG/C6ap734GPkmtocaYVwuHaVVi49D+VIpa4I7s ijRC7Ghiz89EH+/9URQ+m99Vpy9Ee/GoJ9kAuwY8Kj8PuQszRYYudsRVEesBkNx/zrXL F6MA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787989692; x=1788594492; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=LNuXN9S7i6TqJ8ICcOb8D2g40X31e66OWo2XPTKLXSs=; b=TWos8oLVgtMGvmSGIKNgoYX9/FaUrhzaACoBfz/awHCl9WIUfLtV8e+MbymP/a/MUJ DZlqaU0Xls6RP1u6kF/sjCKYdPUammgVf+PhFWItPX498PvYtXLcnhmjDjFwn4+pqLDe 9l9FtfttzkB/mzymAv6nWXn7mL3FRqgLRLCeEFRN6nTjfBqCQ0DLu0gb531wvMjl6PKU F79oNz9RY1X7gExM/cF7H3ABWFlrSd8G/Cd/PP2szVkwfe5EhUdYGcS13HxZLtMyGdnx gbzcczxCsH+kAiNzkqVJ5UjX6ooJ4dKfpsW+pJDWRtJDoAszmi0Ho3CylkzjNblyY0XQ oS7A== X-Gm-Message-State: AFuF++mh9NJO0piJeqHU+raMdjr7ywYQd7VIMehNqCftoUy/rLD5TZ3e STS1UfwW7LF3XIgTnKZpA+9R6Ofh/tj+KzHcHl0hKM9Ak7wHPPhz3zyg X-Gm-Gg: AYBFou0dC2027Q02AMItM0xdXiNvE3sGCPnbSyNyfdrwCePYfK06VnFmYLzCMCSplTn riXjaK1bU8a1ItLo30g+PpkVCuBR+rlQesnYI+tzfwukdwBqarymriuO/om7ygMue6pHI5SD2c6 yku2MtGkv3sP6c1YPUbCDxdGYkp99R50pex6DdAJ7PXuSGmE/0PrCT+7kRMwdSaIzAmPG8nITzQ sd8DoUTYTBpx1gX2Cb4iRaGhmYE6AACPr5ZXUhsrC2SaeAe3e+fYwDrI3eONkNipfBZL4rDjk0w P1+a3eYhA0Cjmk9EcE6Zdb5d4tRxkynLLNSNb8Jh/PHC7Ekd5TbVs15SXgy1PumkCHQzqLUtEfE LTkKQGNjyOaBgM5XXdkTkoc7w5fDjU59IcPWQZJo9O+KfktrM6p/mA6/dpVx0iXmEUOtFy255Fw MWCyTDhikJ9BitDoT60OmUnav9G3pacSyXlNJrIiqco7zFkmJjBYEzmCxSugzIizSx6MATtDSfC 1tYPXbBd8suPwCIN1aIpf+KEgNf X-Received: by 2002:a17:90b:4c0b:b0:398:9c0c:7c72 with SMTP id 98e67ed59e1d1-3989c0c7ffemr6904606a91.25.1787989691567; Sat, 29 Aug 2026 00:48:11 -0700 (PDT) Received: from localhost.localdomain (vmi2317720.contaboserver.net. [84.247.152.65]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-396ddca5bfbsm6255487a91.17.2026.08.29.00.48.06 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Sat, 29 Aug 2026 00:48:11 -0700 (PDT) From: "Lian Wang (ProcessMission)" To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, Andrew Morton , Chris Li , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , Kemeng Shi , Kairui Song , Jihan LIN , Kunwu Chan , Kees Cook , "Gustavo A. R. Silva" , linux-hardening@vger.kernel.org Subject: [RESEND RFC PATCH v2 09/13] mm/swap: add priority queue for swap device allocation Date: Sat, 29 Aug 2026 15:48:03 +0800 Message-ID: <20260829-swap-pcp-priq-v2-resend-9-68d3d925578c@gmail.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260829-swap-pcp-priq-v2-resend-0-68d3d925578c@gmail.com> References: <20260829-swap-pcp-priq-v2-resend-0-68d3d925578c@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Kairui Song The swap allocator uses swap_avail_head, a plist ordered by priority, to select devices. Devices at the same priority are rotated with plist_requeue(), so every cluster transition serializes allocators on swap_avail_lock and repeatedly drops and reacquires that global lock. Replace device selection with a priority queue made of immutable rings. Each ring contains devices at one priority, and each CPU has a local reader that rotates through a ring after a fixed allocation quota. A task-local cursor keeps a retry walk stable even if another task updates the shared reader between attempts, so each peer is visited exactly once before the allocator considers a lower priority. Disable task migration across the retry walk so queue selection, quota accounting and per-CPU cluster allocation stay on the same CPU. This preserves per-CPU pacing without making the sleepable allocation loop an atomic context. Keep the queue structure stable under swapon_rwsem. Full or disabled devices remain in their ring with a tag in the low bit of the stored pointer. Serialize tag writers with swap_queue_update_lock and pair the lockless full-pointer reads and writes with READ_ONCE() and WRITE_ONCE(). The in-use counter carries a separate off-list bit so full-to-available transitions update the counter and pointer tag consistently. Publish swap_file, the live percpu reference, queue membership and SWP_WRITEOK in one swapon writer section. Preserve the writer-serialized swapoff lookup and disable invariant established earlier in the series. For large folios, try every device in the selected priority ring before returning -E2BIG to request a split. Do not fall back to a lower priority device merely because the first same-priority device is fragmented. The old available plist is still maintained in parallel in this commit so the transition remains bisectable. It is removed by the next patch. Link: https://lore.kernel.org/20260714-swap-pcp-priq-v1-9-de9b164ed419@tenc= ent.com Link: https://lore.kernel.org/alZ7UBXweuuOX4qz@yjaykim-PowerEdge-T330 Link: https://lore.kernel.org/77d6da3d-10af-49a1-a356-72aa8b462e85@gmail.com Signed-off-by: Kairui Song Co-developed-by: Lian Wang (ProcessMission) Signed-off-by: Lian Wang (ProcessMission) Tested-by: Kunwu Chan Tested-by: Kunwu Chan --- include/linux/swap.h | 5 +- mm/swapfile.c | 526 +++++++++++++++++++++++++++++++++++++------ 2 files changed, 464 insertions(+), 67 deletions(-) diff --git a/include/linux/swap.h b/include/linux/swap.h index aa66d8454186..37fe2e4d2774 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -201,8 +201,9 @@ struct swap_extent { * - SWP_USED: Protected by swapon_rwsem. Indicates the device is inuse. O= nce * set, won't be cleared unless all reference to this device is freed and * swapoff finished. - * - SWP_WRITEOK: Protected by both swapon_rwsem and swap_avail_lock, clea= ring - * this flag also waits for all current cluster lock users to exit so + * - SWP_WRITEOK: Protected by both swapon_rwsem and swap_queue_update_loc= k. + * Clearing this flag is followed by waiting for all current cluster lock + * users to exit, so * checking this flag while holding any of these locks ensures the device * is safe to use at the moment. Note: clearing this flag doesn't affect * pending IO or async requests, it only prevents further entry allocati= on diff --git a/mm/swapfile.c b/mm/swapfile.c index 79ecff2d0bd3..1b7bc968b5f7 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -55,6 +55,7 @@ static bool folio_swapcache_freeable(struct folio *folio); static void move_cluster(struct swap_info_struct *si, struct swap_cluster_info *ci, struct list_head *list, enum swap_cluster_flags new_flags); +static bool get_swap_device_info(struct swap_info_struct *si); =20 /* * Serializes swapon/swapoff (writers) and protects the swap_info @@ -160,20 +161,382 @@ static struct swap_info_struct *swap_entry_to_info(s= wp_entry_t entry) return swap_type_to_info(swp_type(entry)); } =20 +/* + * All available swap_info_structs are grouped by priority rings, the rings + * are ordered in a queue by priority (higher prio value =3D higher priori= ty). + * The allocator iterates and rotates devices within each priority ring. + * When all devices in a ring are iterated, it goes to the next lower + * priority ring. + */ +struct swap_prio_ring { + int prio; + unsigned int size; + struct swap_info_struct *dev[] __counted_by(size); +}; + +/* + * The ring is protected by swapon_rwsem so updating it is costly. To make + * the allocator and other users skip full devices faster, the lowest bit = of + * a device pointer is used to mark it disabled (temporarily unavailable). + * This relies on the natural alignment of struct swap_info_struct. + */ +#define SWAP_DEVICE_MASKED_SHIFT 0 +#define SWAP_DEVICE_MASKED_BIT BIT(SWAP_DEVICE_MASKED_SHIFT) +static_assert(__alignof__(struct swap_info_struct) >=3D 2); + +/* + * Serializes queue content mutations and keeps SWAP_USAGE_OFFLIST_BIT + * consistent with the masked state of each device pointer. + */ +static DEFINE_SPINLOCK(swap_queue_update_lock); + +/* + * Swap queue is protected by both swap_queue_update_lock and swapon_rwsem. + * Only swapon/swapoff will take the write lock, and modify the queue leng= th + * or any ring's length. swap_queue_update_lock protects the content so + * devices can be masked easily without taking the writelock, which is hea= vy. + */ +static struct swap_prio_ring **swap_queue; +static unsigned int swap_queue_len; + +/* + * Each CPU has its read iterator, so the queue itself will remain read + * only and the CPU side reader rotates by iterating the devices + * periodically using the counter. + */ +#define SWAP_ROUND_ROBIN_QUOTA SWAPFILE_CLUSTER +struct swap_ring_iterator { + int offset; + long rr_counter; +}; + +struct swap_queue_reader { + local_lock_t lock; + struct swap_ring_iterator ri[]; +}; + +struct swap_queue_cursor { + bool valid; + bool ring_only; + unsigned int ring_idx; + unsigned int offset; +}; + +static __percpu struct swap_queue_reader *swap_queue_readers; + +static inline bool swap_device_masked(struct swap_info_struct *si) +{ + return (unsigned long)si & SWAP_DEVICE_MASKED_BIT; +} + +static inline struct swap_info_struct *swap_device_mask_ptr(struct swap_in= fo_struct *si) +{ + return (struct swap_info_struct *)((unsigned long)si | + SWAP_DEVICE_MASKED_BIT); +} + +static inline struct swap_info_struct *swap_device_unmask_ptr(struct swap_= info_struct *si) +{ + return (struct swap_info_struct *)((unsigned long)si & ~SWAP_DEVICE_MASKE= D_BIT); +} + +static struct swap_queue_reader __percpu *swap_queue_prealloc_readers(int = nr_rings, gfp_t gfp) +{ + struct swap_queue_reader __percpu *readers; + + if (!nr_rings) + return NULL; + + readers =3D __alloc_percpu_gfp(struct_size(readers, ri, nr_rings), + __alignof__(*readers), gfp); + return readers; +} + +static void swap_queue_install_readers(struct swap_queue_reader __percpu *= readers) +{ + int ring_idx, cpu; + struct swap_prio_ring *ring; + + free_percpu(swap_queue_readers); + swap_queue_readers =3D readers; + if (!readers) + return; + + /* Distribute each CPU's swap IO fairly across devices. */ + for_each_possible_cpu(cpu) { + local_lock_init(&per_cpu_ptr(readers, cpu)->lock); + + for (ring_idx =3D 0; ring_idx < swap_queue_len; ring_idx++) { + ring =3D swap_queue[ring_idx]; + per_cpu_ptr(readers, cpu)->ri[ring_idx].offset =3D + cpu % ring->size; + per_cpu_ptr(readers, cpu)->ri[ring_idx].rr_counter =3D + SWAP_ROUND_ROBIN_QUOTA; + } + } +} + +static struct swap_info_struct *swap_queue_get_device(long nr_alloc, int n= r_iter, + struct swap_queue_cursor *cursor) +{ + bool rotate =3D false; + struct swap_info_struct *si; + struct swap_ring_iterator *ri; + struct swap_prio_ring *ring; + unsigned int dev_idx, queue_idx; + + if (!swap_queue_len) + return ERR_PTR(-ENOENT); + + queue_idx =3D 0; + while (nr_iter >=3D swap_queue[queue_idx]->size) { + nr_iter -=3D swap_queue[queue_idx]->size; + if (++queue_idx >=3D swap_queue_len) + return ERR_PTR(cursor->ring_only ? -E2BIG : -ENOENT); + } + if (cursor->ring_only && cursor->ring_idx !=3D queue_idx) + return ERR_PTR(-E2BIG); + + ring =3D swap_queue[queue_idx]; + local_lock(&swap_queue_readers->lock); + ri =3D this_cpu_ptr(&swap_queue_readers->ri[queue_idx]); + /* + * Snapshot the starting offset for this allocation's walk. The shared + * iterator can move between retries, but the cursor must visit every + * device in the ring exactly once before falling through. + */ + if (!cursor->valid || cursor->ring_idx !=3D queue_idx) { + cursor->valid =3D true; + cursor->ring_idx =3D queue_idx; + if (ri->rr_counter < nr_alloc) + rotate =3D true; + else if (ri->offset >=3D ring->size) + rotate =3D true; + if (rotate) { + ri->offset++; + ri->offset %=3D ring->size; + ri->rr_counter =3D SWAP_ROUND_ROBIN_QUOTA; + } + cursor->offset =3D ri->offset; + } + + dev_idx =3D (cursor->offset + nr_iter) % ring->size; + if (nr_iter) { + ri->offset =3D dev_idx; + ri->rr_counter =3D SWAP_ROUND_ROBIN_QUOTA; + } + ri->rr_counter -=3D nr_alloc; + si =3D READ_ONCE(ring->dev[dev_idx]); + local_unlock(&swap_queue_readers->lock); + + if (swap_device_masked(si)) + return ERR_PTR(-EBUSY); + + si =3D swap_device_unmask_ptr(si); + return si; +} + +static bool swap_queue_find(struct swap_info_struct *si, + unsigned int *ring_idx, unsigned int *dev_idx) +{ + unsigned int i, j; + struct swap_prio_ring *ring; + + lockdep_assert(lockdep_is_held(&swapon_rwsem) || + lockdep_is_held(&swap_queue_update_lock)); + + for (i =3D 0; i < swap_queue_len; i++) { + ring =3D swap_queue[i]; + if (ring->prio !=3D si->prio) + continue; + for (j =3D 0; j < ring->size; j++) { + if (swap_device_unmask_ptr(READ_ONCE(ring->dev[j])) !=3D si) + continue; + *ring_idx =3D i; + *dev_idx =3D j; + return true; + } + } + return false; +} + +static void swap_queue_mask(struct swap_info_struct *si) +{ + unsigned int ring_idx, dev_idx; + + lockdep_assert_held(&swap_queue_update_lock); + if (swap_queue_find(si, &ring_idx, &dev_idx)) + WRITE_ONCE(swap_queue[ring_idx]->dev[dev_idx], + swap_device_mask_ptr(si)); +} + +static void swap_queue_unmask(struct swap_info_struct *si) +{ + unsigned int ring_idx, dev_idx; + + lockdep_assert_held(&swap_queue_update_lock); + if (swap_queue_find(si, &ring_idx, &dev_idx)) + WRITE_ONCE(swap_queue[ring_idx]->dev[dev_idx], si); +} + +static int swap_queue_add(struct swap_info_struct *si) +{ + struct swap_prio_ring **new_queue =3D NULL, **old_queue =3D NULL; + struct swap_queue_reader __percpu *new_readers =3D NULL; + struct swap_prio_ring *ring, *new_ring =3D NULL, *old_ring =3D NULL; + int prio =3D si->prio; + int i, pos, err =3D -ENOMEM; + gfp_t gfp; + + /* Swap not usable here because this is swap, just reclaim cache. */ + gfp =3D GFP_NOIO | __GFP_HIGH; + lockdep_assert_held_write(&swapon_rwsem); + + for (pos =3D 0; pos < swap_queue_len; pos++) { + if (swap_queue[pos]->prio =3D=3D prio) + goto add_to_ring; + if (swap_queue[pos]->prio < prio) + break; + } + + /* No ring at this priority: insert a new one at pos. */ + new_readers =3D swap_queue_prealloc_readers(swap_queue_len + 1, gfp); + if (!new_readers) + goto failed; + new_queue =3D kmalloc_array(swap_queue_len + 1, sizeof(*swap_queue), gfp); + if (!new_queue) + goto failed; + new_ring =3D kmalloc(struct_size(new_ring, dev, 1), gfp); + if (!new_ring) + goto failed; + if (!get_swap_device_info(si)) + goto failed; + + new_ring->prio =3D prio; + new_ring->size =3D 1; + new_ring->dev[0] =3D si; + for (i =3D 0; i < pos; i++) + new_queue[i] =3D swap_queue[i]; + new_queue[pos] =3D new_ring; + for (i =3D pos; i < swap_queue_len; i++) + new_queue[i + 1] =3D swap_queue[i]; + + spin_lock(&swap_queue_update_lock); + old_queue =3D swap_queue; + swap_queue =3D new_queue; + swap_queue_len++; + spin_unlock(&swap_queue_update_lock); + kfree(old_queue); + + swap_queue_install_readers(new_readers); + return 0; + +add_to_ring: + ring =3D swap_queue[pos]; + new_ring =3D kmalloc(struct_size(ring, dev, ring->size + 1), gfp); + if (!new_ring) + goto failed; + if (!get_swap_device_info(si)) + goto failed; + spin_lock(&swap_queue_update_lock); + memcpy(new_ring, ring, struct_size(ring, dev, ring->size)); + new_ring->size++; + new_ring->dev[new_ring->size - 1] =3D si; + old_ring =3D swap_queue[pos]; + swap_queue[pos] =3D new_ring; + spin_unlock(&swap_queue_update_lock); + kfree(old_ring); + return 0; + +failed: + free_percpu(new_readers); + kfree(new_queue); + kfree(new_ring); + return err; +} + +static void swap_queue_del(struct swap_info_struct *si) +{ + gfp_t gfp; + unsigned int ring_idx, dev_idx; + struct swap_queue_reader __percpu *new_readers =3D NULL; + struct swap_prio_ring *ring, *new_ring =3D NULL, *old_ring =3D NULL; + struct swap_prio_ring **new_queue =3D NULL, **old_queue =3D NULL; + + lockdep_assert_held_write(&swapon_rwsem); + if (!swap_queue_find(si, &ring_idx, &dev_idx)) { + WARN_ON(1); + return; + } + + /* + * To shrink memory usage, pre-allocate new smaller data before + * locking. Failure is fine, swapoff will release them anyway. + */ + gfp =3D GFP_NOIO | __GFP_HIGH; + ring =3D swap_queue[ring_idx]; + if (ring->size > 1) + new_ring =3D kmalloc(struct_size(ring, dev, ring->size - 1), gfp); + if (ring->size =3D=3D 1 && swap_queue_len > 1) { + new_readers =3D swap_queue_prealloc_readers(swap_queue_len - 1, + gfp); + new_queue =3D kmalloc(sizeof(*swap_queue) * + (swap_queue_len - 1), gfp); + } + + spin_lock(&swap_queue_update_lock); + if (ring->size > 1) { + /* Shift trailing devices left to fill the gap. */ + while (++dev_idx < ring->size) + ring->dev[dev_idx - 1] =3D + ring->dev[dev_idx]; + ring->size--; + if (new_ring) { + memcpy(new_ring, ring, + struct_size(ring, dev, ring->size)); + old_ring =3D ring; + swap_queue[ring_idx] =3D new_ring; + } + } else { + /* Last device in this ring: remove the ring. */ + old_ring =3D ring; + swap_queue_len--; + while (++ring_idx <=3D swap_queue_len) + swap_queue[ring_idx - 1] =3D + swap_queue[ring_idx]; + if (new_queue) { + memcpy(new_queue, swap_queue, + sizeof(*swap_queue) * swap_queue_len); + old_queue =3D swap_queue; + swap_queue =3D new_queue; + } else if (!swap_queue_len) { + old_queue =3D swap_queue; + swap_queue =3D NULL; + } + if (new_readers || !swap_queue_len) + swap_queue_install_readers(new_readers); + } + spin_unlock(&swap_queue_update_lock); + + kfree(old_ring); + kfree(old_queue); + put_swap_device(si); +} + /* * Use the second highest bit of inuse_pages counter as the indicator - * if one swap device is on the available plist, so the atomic can + * if one swap device is unavailable for allocation, so the atomic can * still be updated arithmetically while having special data embedded. * * inuse_pages counter is the only thing indicating if a device should - * be on avail_lists or not (except swapon / swapoff). By embedding the - * off-list bit in the atomic counter, updates no longer need any lock - * to check the list status. + * be in the available queue or not (except swapon / swapoff). By + * embedding the off-list bit in the atomic counter, updates no longer + * need any lock to check the list status. * - * This bit will be set if the device is not on the plist and not - * usable, will be cleared if the device is on the plist. + * This bit will be set if the device is not in the available queue + * and not usable, will be cleared if the device is in the queue. */ -#define SWAP_USAGE_OFFLIST_BIT (1UL << (BITS_PER_TYPE(atomic_t) - 2)) +#define SWAP_USAGE_OFFLIST_BIT BIT(BITS_PER_LONG - 2) #define SWAP_USAGE_COUNTER_MASK (~SWAP_USAGE_OFFLIST_BIT) static long swap_usage_in_pages(struct swap_info_struct *si) { @@ -1221,6 +1584,7 @@ static void del_from_avail_list(struct swap_info_stru= ct *si, bool swapoff) unsigned long pages; =20 spin_lock(&swap_avail_lock); + spin_lock(&swap_queue_update_lock); =20 /* * Force remove it only for swapoff. Else, take it off-list only if @@ -1238,9 +1602,10 @@ static void del_from_avail_list(struct swap_info_str= uct *si, bool swapoff) atomic_long_or(SWAP_USAGE_OFFLIST_BIT, &si->inuse_pages); } =20 + swap_queue_mask(si); plist_del(&si->avail_list, &swap_avail_head); - skip: + spin_unlock(&swap_queue_update_lock); spin_unlock(&swap_avail_lock); } =20 @@ -1251,12 +1616,12 @@ static void add_to_avail_list(struct swap_info_stru= ct *si) unsigned long pages; =20 spin_lock(&swap_avail_lock); + spin_lock(&swap_queue_update_lock); =20 /* - * Add the device to the avail list if SWP_WRITEOK is set and - * SWAP_USAGE_OFFLIST_BIT is still set. Swapoff clears - * SWP_WRITEOK first, so the device won't be re-added after - * swapoff starts unless swap_device_enable resurrects it. + * Mark the device as avail if SWP_WRITEOK is set. Swapoff clears + * SWP_WRITEOK first, so check that first so the device won't be + * re-added after swapoff started. */ if (!(si->flags & SWP_WRITEOK)) goto skip; @@ -1267,21 +1632,23 @@ static void add_to_avail_list(struct swap_info_stru= ct *si) val =3D atomic_long_fetch_and_relaxed(~SWAP_USAGE_OFFLIST_BIT, &si->inuse= _pages); =20 /* - * When device is full and device is on the plist, only one updater will - * see (inuse_pages =3D=3D si->pages) and will call del_from_avail_list. = If - * that updater happen to be here, just skip adding. + * When device is full and marked as available, one reader will see + * (inuse_pages =3D=3D si->pages) and should mark it as unavailable and + * set SWAP_USAGE_OFFLIST_BIT. If that updater happens to be here, just + * skip the rest. */ pages =3D si->pages; - if (val =3D=3D pages) { + if ((val & SWAP_USAGE_COUNTER_MASK) =3D=3D pages) { /* Just like the cmpxchg in del_from_avail_list */ if (atomic_long_try_cmpxchg(&si->inuse_pages, &pages, pages | SWAP_USAGE_OFFLIST_BIT)) goto skip; } =20 + swap_queue_unmask(si); plist_add(&si->avail_list, &swap_avail_head); - skip: + spin_unlock(&swap_queue_update_lock); spin_unlock(&swap_avail_lock); } =20 @@ -1298,7 +1665,7 @@ static void swap_device_inuse_add(struct swap_info_st= ruct *si, =20 /* * If device is full, and SWAP_USAGE_OFFLIST_BIT is not set, - * remove it from the plist. + * mark it unavailable. */ inuse_pages =3D atomic_long_add_return_relaxed(nr_entries, &si->inuse_pag= es); if (unlikely(inuse_pages =3D=3D si->pages)) { @@ -1342,7 +1709,7 @@ static void swap_device_inuse_sub(struct swap_info_st= ruct *si, unsigned long off =20 /* * If device is not full, and SWAP_USAGE_OFFLIST_BIT is set, - * add it back to the plist. + * add it back to the available queue. */ inuse_pages =3D atomic_long_sub_return_relaxed(nr_entries, &si->inuse_pag= es); if (unlikely(inuse_pages & SWAP_USAGE_OFFLIST_BIT)) @@ -1365,41 +1732,44 @@ static bool get_swap_device_info(struct swap_info_s= truct *si) return true; } =20 -/* Rotate the device and switch to a new cluster */ -static void swap_alloc_entry(struct folio *folio) +static int swap_alloc_entry(struct folio *folio) { - struct swap_info_struct *si, *next; + struct swap_queue_cursor cursor =3D {}; + long nr_pages =3D folio_nr_pages(folio); + struct swap_info_struct *si; + int nr_iter, ret; =20 - spin_lock(&swap_avail_lock); -start_over: - plist_for_each_entry_safe(si, next, &swap_avail_head, avail_list) { - /* Rotate the device and switch to a new cluster */ - plist_requeue(&si->avail_list, &swap_avail_head); - spin_unlock(&swap_avail_lock); - if (get_swap_device_info(si)) { - cluster_alloc_swap_entry(si, folio); - put_swap_device(si); - if (folio_test_swapcache(folio)) - return; - if (folio_test_large(folio)) - return; + percpu_down_read(&swapon_rwsem); + migrate_disable(); + for (nr_iter =3D 0;; nr_iter++) { + si =3D swap_queue_get_device(nr_pages, nr_iter, &cursor); + if (IS_ERR(si)) { + ret =3D PTR_ERR(si); + if (ret =3D=3D -EBUSY) + continue; + break; + } + cluster_alloc_swap_entry(si, folio); + + if (folio_test_swapcache(folio)) { + ret =3D 0; + break; } =20 - spin_lock(&swap_avail_lock); /* - * if we got here, it's likely that si was almost full before, - * multiple callers probably all tried to get a page from the - * same si and it filled up before we could get one; or, the si - * filled up between us dropping swap_avail_lock. - * Since we dropped the swap_avail_lock, the swap_avail_list - * may have been modified; so if next is still in the - * swap_avail_head list then try it, otherwise start over if we - * have not gotten any slots. + * For large allocations, try every device at the same priority, + * but ask the caller to split instead of falling back to a lower + * priority ring. */ - if (plist_node_empty(&next->avail_list)) - goto start_over; + if (folio_test_large(folio)) { + cursor.ring_only =3D true; + continue; + } } - spin_unlock(&swap_avail_lock); + + migrate_enable(); + percpu_up_read(&swapon_rwsem); + return ret; } =20 /* @@ -1713,6 +2083,7 @@ int folio_alloc_swap(struct folio *folio) { unsigned int order =3D folio_order(folio); unsigned int size =3D 1 << order; + int ret; =20 VM_BUG_ON_FOLIO(!folio_test_locked(folio), folio); VM_BUG_ON_FOLIO(!folio_test_uptodate(folio), folio); @@ -1736,7 +2107,7 @@ int folio_alloc_swap(struct folio *folio) } =20 again: - swap_alloc_entry(folio); + ret =3D swap_alloc_entry(folio); =20 if (!order && unlikely(!folio_test_swapcache(folio))) { if (swap_sync_discard()) @@ -1748,7 +2119,7 @@ int folio_alloc_swap(struct folio *folio) swap_cache_del_folio(folio); =20 if (unlikely(!folio_test_swapcache(folio))) - return -ENOMEM; + return ret ? ret : -ENOMEM; =20 return 0; } @@ -2921,24 +3292,29 @@ static int setup_swap_extents(struct swap_info_stru= ct *sis, } =20 /* - * Mark a fully initialized swap device writable and expose it to the - * allocator. The caller must have resurrected its percpu ref first. + * Mark a fully initialized swap device writable and expose it to the allo= cator. + * The caller must have resurrected its percpu ref before entering this he= lper. */ -static void swap_device_enable(struct swap_info_struct *si) +static void __swap_device_enable(struct swap_info_struct *si) { - percpu_down_write(&swapon_rwsem); - spin_lock(&swap_avail_lock); - si->flags |=3D SWP_WRITEOK; - spin_unlock(&swap_avail_lock); + lockdep_assert_held_write(&swapon_rwsem); =20 + spin_lock(&swap_queue_update_lock); + si->flags |=3D SWP_WRITEOK; + spin_unlock(&swap_queue_update_lock); atomic_long_add(si->pages, &nr_swap_pages); total_swap_pages +=3D si->pages; plist_add(&si->list, &swap_active_head); - percpu_up_write(&swapon_rwsem); - add_to_avail_list(si); } =20 +static void swap_device_enable(struct swap_info_struct *si) +{ + percpu_down_write(&swapon_rwsem); + __swap_device_enable(si); + percpu_up_write(&swapon_rwsem); +} + static int swap_device_disable(struct address_space *mapping, struct swap_info_struct **swap_info) { @@ -2972,10 +3348,9 @@ static int swap_device_disable(struct address_space = *mapping, } vm_unacct_memory(si->pages); =20 - spin_lock(&swap_avail_lock); + spin_lock(&swap_queue_update_lock); si->flags &=3D ~SWP_WRITEOK; - spin_unlock(&swap_avail_lock); - + spin_unlock(&swap_queue_update_lock); plist_del(&si->list, &swap_active_head); total_swap_pages -=3D si->pages; atomic_long_sub(si->pages, &nr_swap_pages); @@ -3060,6 +3435,11 @@ SYSCALL_DEFINE1(swapoff, const char __user *, specia= lfile) return err; } =20 + percpu_down_write(&swapon_rwsem); + swap_queue_del(p); + percpu_ref_kill(&p->users); + percpu_up_write(&swapon_rwsem); + /* * Wait for swap operations protected by get/put_swap_device() * to complete. Because of synchronize_rcu() here, all swap @@ -3068,7 +3448,6 @@ SYSCALL_DEFINE1(swapoff, const char __user *, special= file) * prevent folio_test_swapcache() and the following swap cache * operations from racing with swapoff. */ - percpu_ref_kill(&p->users); synchronize_rcu(); wait_for_completion(&p->comp); =20 @@ -3721,11 +4100,28 @@ SYSCALL_DEFINE2(swapon, const char __user *, specia= lfile, int, swap_flags) si->prio =3D prio; si->list.prio =3D -si->prio; si->avail_list.prio =3D -si->prio; - si->swap_file =3D swap_file; =20 - /* Sets SWP_WRITEOK, resurrect the percpu ref, expose the swap device */ + /* + * Publish swap_file before making the percpu ref live, then add the devi= ce + * to the queue and make it writable under the same write-side lock. This + * keeps lockless ref users and /proc/swaps from observing partial state. + */ + percpu_down_write(&swapon_rwsem); + si->swap_file =3D swap_file; percpu_ref_resurrect(&si->users); - swap_device_enable(si); + error =3D swap_queue_add(si); + if (error) { + si->swap_file =3D NULL; + percpu_ref_kill(&si->users); + } else { + __swap_device_enable(si); + } + percpu_up_write(&swapon_rwsem); + if (error) { + wait_for_completion(&si->comp); + inode->i_flags &=3D ~S_SWAPFILE; + goto free_swap_zswap; + } =20 pr_info("Adding %uk swap on %s. Priority:%d extents:%d across:%lluk %s%s= %s%s\n", K(si->pages), name->name, si->prio, nr_extents, --=20 2.55.0 From nobody Sat Sep 26 22:02:19 2026 Received: from mail-pj1-f44.google.com (mail-pj1-f44.google.com [209.85.216.44]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 924D935CB81 for ; Sat, 29 Aug 2026 07:48:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.44 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787989702; cv=none; b=uM+sbgsruOIJXbrsLKlzd4jOGAeOgUu2GjFE+nMu6bOX50NWyotV1k0abpcD/m0/vAwKOvQ9a3yQRQvUsS8pRXnXSXzq52BycrbfrhKoSsMWMteeLm8SVTEjoRb8U/kwOXoaTCQFZyXh3sNdyfHRzUjE8aQe2uieRGMk6BFsU0E= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787989702; c=relaxed/simple; bh=9bX8t1k0MccBKhH+qGjWHCG8x1C8CXkWD1zFec0CM5E=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=gQtfQjQfQxnYOXv4OHDIlWwVMadBhMf14Zhuxqvp9RMOVaxxQREYkZ9jyx3pAv1eqIkJ+qZIlZoHWFoRBhNroktlVfkcn2F/esUwFSFm7zauUG72/v6yMvihSCKsINvNVL+HF179UKZpP2pZXs/4P2hkNX+f9IEBJIp6ZM+vkMY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=iL/YPfrd; arc=none smtp.client-ip=209.85.216.44 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="iL/YPfrd" Received: by mail-pj1-f44.google.com with SMTP id 98e67ed59e1d1-382ef647e20so1761520a91.1 for ; Sat, 29 Aug 2026 00:48:20 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787989700; x=1788594500; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=fQpy6R10A7XgwYtm9Mv+zsa+EQTNg+ZbkFt80LhBXZk=; b=iL/YPfrdq4nCZJThP/NWDNEheMAj3KwRASi847msh4gVj90vmPcbn+lxwWlz1e32bn f0kCZhaAW06/sVw8pmRKTjRYLAYUrP3hLLd09T4YNEIB3gRuXtC43sIWf59fmQoyERZY YwRfm72TposZgD5sHK7OZfYKmozu448R00AvnhtYwqRWI25PX++g7v7/Bpd6kHSvZoei S77a8RD7LGasRxiGuSEZ9T17eqght0GbtBoOVXrn0azBtTgve0YonlGy5pRB2d96wlEz osYNc/NNf+06PoN3QUezKSTwH8wnsbDJyThXQ59eRTozmjmZpEBMx1GERwGu9wFVt9MZ AZGg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787989700; x=1788594500; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=fQpy6R10A7XgwYtm9Mv+zsa+EQTNg+ZbkFt80LhBXZk=; b=OM9d4Yqq0YdO5MEr0wfZNzJKekRE5SmK1cMFHed8f0gzM0vElyR8jr+LmYKUCvTiFU faRenDfkS0JqZ0Ciigx2p5GAnvoB5zOsAtYzfbI3TGt83+SWFdatoZXmMYm5dXAHwLba LKeyIA3u13LRwGGCbmBZLgA22+kRV3IclhF61Z0UmmCA6hzmde98By1mNhQ7VuIzOMXf D0iNAx/ZXZ2KJe7FLNHodg6Iz1dzkss7l1agMqm8Xaq2V+XvAkJKEW38XLvzBUdkUS13 HSjIXfYptDCWdsGf36b/gcp+YnErB0h3Khfn2rKgT2LpIng+BwA60P6hL0T6LHHWp5za 1W2w== X-Gm-Message-State: AFuF++lyJILHJrvuTyQK0lbEPTYtxOxUE6loWvQERsGqOgkF4PdW5sal bfeDE/lUQnsvEt3198mgHw3hlaaK0ZzctLYArUHGjljiACI63J13qHmr X-Gm-Gg: AYBFou2qX8nxVxeMHvQQti1saUIjnnOSrv7lM9rhw0R8TOoFdKcvmT6Ro2nrKJ9BRh4 JJ9VkLXoyXqAYIdVN5EEvnyYYsHGOMZMxG7W9tPqjLcQuNybVn3/YOTKB7HJ9/uqjXtSMZ+WpSD 2gOEhja4bEnqGpX0geNmj7f+fyeT6HmWVs6r7odwR+CSje4itZaI2nmNWdagritDsmXBRImqjVx UNBmnG81Lh+zNx+Ntomnts/x2z778Gv/NQafMoHqvl7ziZmTXya2lmKc5MSWmlrJfI/S5EYxmpF gGBmquBTAJIDMUGiTc/prvveMAG0Fx5ngPncZiifqHqEvAJdBJYDG/QdegqjSe1a8fa1swLU5Cm VfRBfRJgzrmY/QPV/VFODugdKpqkCZ2Zpo2eBXemxM3K1qRnwAtBokMaV/Dfk+389HOO+TCkEN4 ezl/S3/ZP+VlTBFtWUq+9FfOgK9C5/NjVQtNvIwZ8kYTTyulL178dfzML/EATB6t9fXiBTPCEtE /jXkfKk4tfoMNsqUHd+c1POHXfICA== X-Received: by 2002:a17:90a:4ca6:b0:398:bd66:35f5 with SMTP id 98e67ed59e1d1-398bd66372fmr151113a91.25.1787989699843; Sat, 29 Aug 2026 00:48:19 -0700 (PDT) Received: from localhost.localdomain (vmi2317720.contaboserver.net. [84.247.152.65]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-396dda7c922sm6117862a91.7.2026.08.29.00.48.15 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Sat, 29 Aug 2026 00:48:19 -0700 (PDT) From: "Lian Wang (ProcessMission)" To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, Andrew Morton , Chris Li , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , Kemeng Shi , Kairui Song , Jihan LIN , Kunwu Chan Subject: [RESEND RFC PATCH v2 10/13] mm/swap: remove available list Date: Sat, 29 Aug 2026 15:48:12 +0800 Message-ID: <20260829-swap-pcp-priq-v2-resend-10-68d3d925578c@gmail.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260829-swap-pcp-priq-v2-resend-0-68d3d925578c@gmail.com> References: <20260829-swap-pcp-priq-v2-resend-0-68d3d925578c@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Kairui Song Pure cleanup, no functional change. After the priority queue replaced the allocator's device selection, swap_avail_head and swap_avail_lock are no longer used. Remove them. The throttle path now walks swap_active_head under swapon_rwsem. Check SWP_WRITEOK and the existing off-list state. This preserves the old priority ordering and does not select a full swap device. Signed-off-by: Kairui Song Signed-off-by: Lian Wang (ProcessMission) Tested-by: Kunwu Chan Tested-by: Kunwu Chan --- include/linux/swap.h | 1 - mm/swapfile.c | 33 ++++++--------------------------- 2 files changed, 6 insertions(+), 28 deletions(-) diff --git a/include/linux/swap.h b/include/linux/swap.h index 37fe2e4d2774..0c8c1ebcab9c 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -293,7 +293,6 @@ struct swap_info_struct { struct work_struct discard_work; /* discard worker */ struct work_struct reclaim_work; /* reclaim worker */ struct list_head discard_clusters; /* discard clusters list */ - struct plist_node avail_list; /* entry in swap_avail_head */ const struct swap_ops *ops; }; =20 diff --git a/mm/swapfile.c b/mm/swapfile.c index 1b7bc968b5f7..e183dfad264e 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -96,21 +96,6 @@ static const char Bad_offset[] =3D "Bad swap offset entr= y "; */ static PLIST_HEAD(swap_active_head); =20 -/* - * all available (active, not full) swap_info_structs - * protected with swap_avail_lock, ordered by priority. - * This is used by folio_alloc_swap() instead of swap_active_head - * because swap_active_head includes all swap_info_structs, - * but folio_alloc_swap() doesn't need to look at full ones. - * This uses its own lock instead of swapon_rwsem because when a - * swap_info_struct changes between not-full/full, it needs to - * add/remove itself to/from this list, but the swap_info_struct->lock - * is held and the locking order requires swapon_rwsem to be taken - * before any swap_info_struct->lock. - */ -static PLIST_HEAD(swap_avail_head); -static DEFINE_SPINLOCK(swap_avail_lock); - static inline struct swap_info_struct *__swap_iter(int *i, unsigned long f= lag) { lockdep_assert_held(&swapon_rwsem); @@ -1583,7 +1568,6 @@ static void del_from_avail_list(struct swap_info_stru= ct *si, bool swapoff) { unsigned long pages; =20 - spin_lock(&swap_avail_lock); spin_lock(&swap_queue_update_lock); =20 /* @@ -1603,10 +1587,8 @@ static void del_from_avail_list(struct swap_info_str= uct *si, bool swapoff) } =20 swap_queue_mask(si); - plist_del(&si->avail_list, &swap_avail_head); skip: spin_unlock(&swap_queue_update_lock); - spin_unlock(&swap_avail_lock); } =20 /* SWAP_USAGE_OFFLIST_BIT can only be cleared by this helper. */ @@ -1615,7 +1597,6 @@ static void add_to_avail_list(struct swap_info_struct= *si) long val; unsigned long pages; =20 - spin_lock(&swap_avail_lock); spin_lock(&swap_queue_update_lock); =20 /* @@ -1646,10 +1627,8 @@ static void add_to_avail_list(struct swap_info_struc= t *si) } =20 swap_queue_unmask(si); - plist_add(&si->avail_list, &swap_avail_head); skip: spin_unlock(&swap_queue_update_lock); - spin_unlock(&swap_avail_lock); } =20 /* @@ -3684,7 +3663,6 @@ static struct swap_info_struct *alloc_swap_info(void) } p->swap_extent_root =3D RB_ROOT; plist_node_init(&p->list, 0); - plist_node_init(&p->avail_list, 0); p->flags =3D SWP_USED; percpu_up_write(&swapon_rwsem); if (defer) { @@ -4099,7 +4077,6 @@ SYSCALL_DEFINE2(swapon, const char __user *, specialf= ile, int, swap_flags) */ si->prio =3D prio; si->list.prio =3D -si->prio; - si->avail_list.prio =3D -si->prio; =20 /* * Publish swap_file before making the percpu ref live, then add the devi= ce @@ -4243,14 +4220,16 @@ void __folio_throttle_swaprate(struct folio *folio,= gfp_t gfp) if (current->throttle_disk) return; =20 - spin_lock(&swap_avail_lock); - plist_for_each_entry(si, &swap_avail_head, avail_list) { - if (si->bdev) { + percpu_down_read(&swapon_rwsem); + plist_for_each_entry(si, &swap_active_head, list) { + if ((si->flags & SWP_WRITEOK) && + !(atomic_long_read(&si->inuse_pages) & + SWAP_USAGE_OFFLIST_BIT) && si->bdev) { blkcg_schedule_throttle(si->bdev->bd_disk, true); break; } } - spin_unlock(&swap_avail_lock); + percpu_up_read(&swapon_rwsem); } #endif =20 --=20 2.55.0 From nobody Sat Sep 26 22:02:19 2026 Received: from mail-pj1-f53.google.com (mail-pj1-f53.google.com [209.85.216.53]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E3DFB334C1D for ; Sat, 29 Aug 2026 07:48:29 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.53 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787989713; cv=none; b=LCZgX4jSLIaq4VOu8IjKYD/r/2FiyqOiDXmCT53jDyX27VGIHkGKbHCzoC+2gtHhK//6zfHg1eGNv2FauCvUGjSmuJotOYtu0lJlg627d1p5/NdgNeAULSkeDJrSVhvQJXWkeRNlJ2kiGqx4HTTJydRoGGgwqP69pc4//sQRoDQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787989713; c=relaxed/simple; bh=4MUIpHCzYF63DZ48ekKA4MO/TtB0skrAysUyZkikHr0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=CXao5Zgmxc0aJekUHUv/SQCZhixUAziWXrEq72aRloyh050T5DrUcVNgq+WEfuvf8SDcFNCw6Uw+YTAX6nLppYXn+GeT4tSvEg6HJf0B23z4artHNL81L15hZfZCZnr1rUZJv+UwMVuGlRKNMNmN1e8bl+8O3IaaEttLAHCjkp4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=dFNoJVLU; arc=none smtp.client-ip=209.85.216.53 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="dFNoJVLU" Received: by mail-pj1-f53.google.com with SMTP id 98e67ed59e1d1-38e42560ebcso1634094a91.1 for ; Sat, 29 Aug 2026 00:48:29 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787989708; x=1788594508; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=Zmu3XpqP05NRue2fBxvw9Ta0ua8JaiA5G6dY6hEI5XE=; b=dFNoJVLU3s8irrEoZXUwQzalMb3yImH5Yy92+lgherv3GkkJkm+0ckl0CqMFukW8EU hj1FX3soCiIqSbIHbJPxf36umnGkOtM4sRs9m+sJ+dW40cALdlYplzxXlvRMk+t4Fz+R YZqdEmMXxUs8Q6q6Wlb1He4qq339rLvASZ7F9Aafh465ci2dWp1izR0iH84gWHBCUEew OWrRd1D/4AB2YRrbjBZqVQD7FLfW4EFfg11Zz3IfiYt+VEZW9FMy6fIQGyigXHNq0FA8 qUZcmjKl+gDKjfNbcy0XTtRxtUCJDS9Z7nCxqbgqZ+WaMp/Ssi4c0QzvuHIz79MB+qJH ZVug== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787989708; x=1788594508; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=Zmu3XpqP05NRue2fBxvw9Ta0ua8JaiA5G6dY6hEI5XE=; b=nxsw5VkbyBn3mbYHPbDZeDsZQisKT/Vap6fTOYS7kOMOIzgJZV0aRks9QJYkPOFVqa /hVxezaluvixId8bxiBeS9on3LmyR5/Qe+9jKjATEze3zc14YWkJa3LBjQoTYv4KdOGx +ccewdTu5EOpog75lbgaSd4ZSsB48GEpaKM8nr77heIhd5HtRx0mRGUCsW/NEzJihT+8 M5NiCcofh79oKvZOW13Gn/MkBAW7vBAKqr553ZliMl0KcIWqCqkz070wRlq7M4MCItkt qm4T/e58076MsOOhKOps3P5L2KkEa34CvsnJcpZMMLxEEmQmsas9u1L+nJ7nWq6Jcck1 ZZtg== X-Gm-Message-State: AFuF++l0TKdLf3w+xQ2sVROLJ0DxRXzqvPpa0d903u5Ako4nwmtOYOJb xdZG2Kh9sz3AfwBNwWeVmmKlcQdAbmVnHkWmDW/zg5BI5BQnXCNIpSD2 X-Gm-Gg: AYBFou0Ty62jYGlB05fjoP9JzUQ9Z88F0HXyHDU/NmkPP5woLQ+bIWWk8jAlpnCr8h+ B6P3V/w9n8TmnKJZoWKMTQBtjDKbCe5ARHsbPvKHwZ81Tcorht2ODlIhvx5yM27seUxuoA5e+kh XDNiti569hCQE7oxfFALfQ0WH7J/KEQMpnSFZ1v2YUR9kJswAur0AhcCU0ncm00dmuD15/hzD7t 2+ezmQuUffX4vPzMv6UH0oo4YtFr4SzLOj3H+DrAtOYnrX2CV7u/QCNFR4zjE494UP0VKAr44ES 319HZCVopBKQQy2CQKjH+gtq7YHZi8UbMFoUUVbVMkyT9olJ93i54hbhmN6D5epqhBl6qSLAE2Y L01tgqQZ00UUOqi+01Bo+heV08RnyK++j1702oKawgl8hM5jT8Q96cqheF3I1sWl4fMnj+JJpmw 3aNxeFcdhl4Kym/gXCk4XNdy0mFxDV85jWpf0stWZkSbDQt7v21jeBomB68QfaLamxy1aNAK30Y sFakJpvtvR+CM85BijDxEomhikXzVLlI3zwcKQ= X-Received: by 2002:a17:90b:52c8:b0:398:9bd1:3214 with SMTP id 98e67ed59e1d1-3989bd13874mr6809369a91.21.1787989707985; Sat, 29 Aug 2026 00:48:27 -0700 (PDT) Received: from localhost.localdomain (vmi2317720.contaboserver.net. [84.247.152.65]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-396b0fd5c34sm10702719a91.8.2026.08.29.00.48.24 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Sat, 29 Aug 2026 00:48:27 -0700 (PDT) From: "Lian Wang (ProcessMission)" To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, Andrew Morton , Chris Li , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , Kemeng Shi , Kairui Song , Jihan LIN , Kunwu Chan Subject: [RESEND RFC PATCH v2 11/13] mm/swap: bound synchronous discard during allocation Date: Sat, 29 Aug 2026 15:48:20 +0800 Message-ID: <20260829-swap-pcp-priq-v2-resend-11-68d3d925578c@gmail.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260829-swap-pcp-priq-v2-resend-0-68d3d925578c@gmail.com> References: <20260829-swap-pcp-priq-v2-resend-0-68d3d925578c@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Kairui Song The old global per-CPU cluster cache made it difficult to select a specific device before handling its pending discards. Now that cluster state is per-device, let an allocator issue discard for that device when its free-cluster list is drained instead of waiting for a global fallback scan. Do not drain the entire discard list from one allocation. The list is also updated by freeing paths, so an unlocked non-empty check races with those writers, and a full synchronous drain has no fixed latency bound when other CPUs continue to enqueue clusters. Factor out swap_discard_one_cluster(). Isolate one cluster under si->lock, retain CLUSTER_FLAG_DISCARD while I/O is in flight, and return it to the free list after the discard completes. The background worker continues looping until the list is empty, while an allocator handles at most one cluster before retrying allocation. This keeps the proactive per-device policy while giving each allocation a fixed synchronous discard budget. Link: https://lore.kernel.org/all/CAMgjq7CsYhEjvtN85XGkrONYAJxve7gG593TFeOG= V-oax++kWA@mail.gmail.com/ Signed-off-by: Kairui Song Co-developed-by: Lian Wang (ProcessMission) Signed-off-by: Lian Wang (ProcessMission) Tested-by: Kunwu Chan Tested-by: Kunwu Chan --- mm/swapfile.c | 149 ++++++++++++++++++++++---------------------------- 1 file changed, 64 insertions(+), 85 deletions(-) diff --git a/mm/swapfile.c b/mm/swapfile.c index e183dfad264e..cb27f2e246f0 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -528,6 +528,26 @@ static long swap_usage_in_pages(struct swap_info_struc= t *si) return atomic_long_read(&si->inuse_pages) & SWAP_USAGE_COUNTER_MASK; } =20 +/* + * Serialize the allocation on single CPU or globally to avoid + * fragmentation and make the workflow easier to follow. + */ +static void swap_alloc_lock_device(struct swap_info_struct *si) +{ + if (si->flags & SWP_SOLIDSTATE) + local_lock(&si->percpu_cluster->lock); + else + spin_lock(&si->global_cluster->lock); +} + +static void swap_alloc_unlock_device(struct swap_info_struct *si) +{ + if (si->flags & SWP_SOLIDSTATE) + local_unlock(&si->percpu_cluster->lock); + else + spin_unlock(&si->global_cluster->lock); +} + /* Reclaim the swap entry anyway if possible */ #define TTRS_ANYWAY 0x1 /* @@ -914,10 +934,7 @@ swap_cluster_populate(struct swap_info_struct *si, * the potential recursive allocation is limited. */ spin_unlock(&ci->lock); - if (si->flags & SWP_SOLIDSTATE) - local_unlock(&si->percpu_cluster->lock); - else - spin_unlock(&si->global_cluster->lock); + swap_alloc_unlock_device(si); =20 ret =3D swap_cluster_alloc_table(ci, __GFP_HIGH | __GFP_NOMEMALLOC | GFP_KERNEL); @@ -930,10 +947,7 @@ swap_cluster_populate(struct swap_info_struct *si, * could happen with ignoring the percpu cluster is fragmentation, * which is acceptable since this fallback and race is rare. */ - if (si->flags & SWP_SOLIDSTATE) - local_lock(&si->percpu_cluster->lock); - else - spin_lock(&si->global_cluster->lock); + swap_alloc_lock_device(si); spin_lock(&ci->lock); =20 if (ret) { @@ -1023,44 +1037,40 @@ static struct swap_cluster_info *isolate_lock_clust= er( } =20 /* - * Doing discard actually. After a cluster discard is finished, the cluster - * will be added to free cluster list. Discard cluster is a bit special as - * they don't participate in allocation or reclaim, so clusters marked as - * CLUSTER_FLAG_DISCARD must remain off-list or on discard list. + * Discard one cluster. After the discard is finished, the cluster will be + * added to the free cluster list. Discard clusters are special because th= ey + * don't participate in allocation or reclaim, so CLUSTER_FLAG_DISCARD must + * remain set while a cluster is either queued or being discarded. */ -static bool swap_do_scheduled_discard(struct swap_info_struct *si) +static bool swap_discard_one_cluster(struct swap_info_struct *si) { struct swap_cluster_info *ci; - bool ret =3D false; unsigned int idx; =20 spin_lock(&si->lock); - while (!list_empty(&si->discard_clusters)) { - ci =3D list_first_entry(&si->discard_clusters, struct swap_cluster_info,= list); - /* - * Delete the cluster from list to prepare for discard, but keep - * the CLUSTER_FLAG_DISCARD flag, there could be percpu_cluster - * pointing to it, or ran into by relocate_cluster. - */ - list_del(&ci->list); - idx =3D cluster_index(si, ci); + if (list_empty(&si->discard_clusters)) { spin_unlock(&si->lock); - discard_swap_cluster(si, idx * SWAPFILE_CLUSTER, - SWAPFILE_CLUSTER); - - spin_lock(&ci->lock); - /* - * Discard is done, clear its flags as it's off-list, then - * return the cluster to allocation list. - */ - ci->flags =3D CLUSTER_FLAG_NONE; - __free_cluster(si, ci); - spin_unlock(&ci->lock); - ret =3D true; - spin_lock(&si->lock); + return false; } + + ci =3D list_first_entry(&si->discard_clusters, + struct swap_cluster_info, list); + /* + * Delete the cluster from the list, but keep CLUSTER_FLAG_DISCARD + * set while the discard is in flight. A percpu cluster may still + * point to it, or relocate_cluster() may encounter it. + */ + list_del(&ci->list); + idx =3D cluster_index(si, ci); spin_unlock(&si->lock); - return ret; + + discard_swap_cluster(si, idx * SWAPFILE_CLUSTER, SWAPFILE_CLUSTER); + + spin_lock(&ci->lock); + ci->flags =3D CLUSTER_FLAG_NONE; + __free_cluster(si, ci); + spin_unlock(&ci->lock); + return true; } =20 static void swap_discard_work(struct work_struct *work) @@ -1069,7 +1079,8 @@ static void swap_discard_work(struct work_struct *wor= k) =20 si =3D container_of(work, struct swap_info_struct, discard_work); =20 - swap_do_scheduled_discard(si); + while (swap_discard_one_cluster(si)) + cond_resched(); } =20 static void swap_users_ref_free(struct percpu_ref *ref) @@ -1467,6 +1478,7 @@ static unsigned long cluster_alloc_swap_entry(struct = swap_info_struct *si, struct swap_cluster_info *ci; unsigned int order =3D likely(folio) ? folio_order(folio) : 0; unsigned int offset =3D SWAP_ENTRY_INVALID, found =3D SWAP_ENTRY_INVALID; + bool discarded =3D false; =20 /* * Swapfile is not block device so unable @@ -1475,15 +1487,12 @@ static unsigned long cluster_alloc_swap_entry(struc= t swap_info_struct *si, if (order && !(si->flags & SWP_BLKDEV)) return 0; =20 - if (si->flags & SWP_SOLIDSTATE) { - /* Fast path using per CPU cluster */ - local_lock(&si->percpu_cluster->lock); +restart: + swap_alloc_lock_device(si); + if (si->flags & SWP_SOLIDSTATE) offset =3D __this_cpu_read(si->percpu_cluster->next[order]); - } else { - /* Serialize HDD SWAP allocation for each device. */ - spin_lock(&si->global_cluster->lock); + else offset =3D si->global_cluster->next[order]; - } =20 if (offset !=3D SWAP_ENTRY_INVALID) { ci =3D swap_cluster_lock(si, offset); @@ -1507,6 +1516,15 @@ static unsigned long cluster_alloc_swap_entry(struct= swap_info_struct *si, found =3D alloc_swap_scan_list(si, &si->free_clusters, folio, false); if (found) goto done; + + if (!discarded) { + swap_alloc_unlock_device(si); + if (swap_discard_one_cluster(si)) { + discarded =3D true; + goto restart; + } + swap_alloc_lock_device(si); + } } =20 if (order < PMD_ORDER) { @@ -1555,10 +1573,7 @@ static unsigned long cluster_alloc_swap_entry(struct= swap_info_struct *si, goto done; } done: - if (si->flags & SWP_SOLIDSTATE) - local_unlock(&si->percpu_cluster->lock); - else - spin_unlock(&si->global_cluster->lock); + swap_alloc_unlock_device(si); =20 return found; } @@ -1751,36 +1766,6 @@ static int swap_alloc_entry(struct folio *folio) return ret; } =20 -/* - * Discard pending clusters in a synchronized way when under high pressure. - * Return: true if any cluster is discarded. - */ -static bool swap_sync_discard(void) -{ - bool ret =3D false; - struct swap_info_struct *si, *next; - - percpu_down_read(&swapon_rwsem); -start_over: - plist_for_each_entry_safe(si, next, &swap_active_head, list) { - percpu_up_read(&swapon_rwsem); - if (get_swap_device_info(si)) { - if (si->flags & SWP_PAGE_DISCARD) - ret =3D swap_do_scheduled_discard(si); - put_swap_device(si); - } - if (ret) - return true; - - percpu_down_read(&swapon_rwsem); - if (plist_node_empty(&next->list)) - goto start_over; - } - percpu_up_read(&swapon_rwsem); - - return false; -} - static int swap_extend_table_alloc(struct swap_info_struct *si, struct swap_cluster_info *ci, unsigned int ci_off, gfp_t gfp) @@ -2085,14 +2070,8 @@ int folio_alloc_swap(struct folio *folio) } } =20 -again: ret =3D swap_alloc_entry(folio); =20 - if (!order && unlikely(!folio_test_swapcache(folio))) { - if (swap_sync_discard()) - goto again; - } - /* Need to call this even if allocation failed, for MEMCG_SWAP_FAIL. */ if (unlikely(mem_cgroup_try_charge_swap(folio))) swap_cache_del_folio(folio); --=20 2.55.0 From nobody Sat Sep 26 22:02:19 2026 Received: from mail-pg1-f172.google.com (mail-pg1-f172.google.com [209.85.215.172]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5E68739AD5E for ; Sat, 29 Aug 2026 07:48:37 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.172 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787989720; cv=none; b=DZIUwocXY9lgWrcC6ISib89psxndOFlHR7R73OPNZQid5FrKwblfvA45wsH5TJOg+8LTOCBtsKu4tnBIFPDxqOdpan9ijqsMiSpC5riTwSPrhybjF3H4mxWUaTyTi7Q2VlohVXDQkUu31JSs1K6n+YBTpn1G/uCZrxRIIHIU8Yg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787989720; c=relaxed/simple; bh=DsLEHOj7RofQYi0x3G+lLhCzYH//rcaCKr1DD+f77zg=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=hHCMR0MCCZA2nEGiLFYCPrUhdR3Yi079kJXnQo+PcbW3s64nIdbPzXXm2HQXiWxtz0cBuvmE+jGzY2Oe9jJFQajp2ml3k7I1SPjYVAesryEK+mGle7nM+IipoHijIkjRflBnIV/TZwa+67LNLLVvwc6PND0eDl0KaAkcnW0MpmI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=UlYu00tP; arc=none smtp.client-ip=209.85.215.172 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="UlYu00tP" Received: by mail-pg1-f172.google.com with SMTP id 41be03b00d2f7-cc149372c14so1316652a12.1 for ; Sat, 29 Aug 2026 00:48:37 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787989717; x=1788594517; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=K0brm21TE9GssWqZ74Js94Ds6AVGqAHayV0XzsofN18=; b=UlYu00tPe9UvVsWNdEJfqqczLPaU+mucEghuvhzi6jigtAJS0at2Lww3C6rDgIy0QE +xtbCJ1x5tQ7vIs9VX13e7eb7w/gXkJdioY0a0qWtCiROs5rqHL2ALhCbpvRunkCHDpp DtHAj2SeuePOjyRoth2PCF2AsT0jQlnbV/QpMMuwsj6dlymsEpoDJgInTkPP5p36wDbn 7Z7jyqvAb3/5bNDJQLaBNq1sbgGYLvfpNQJFOauXUrpgscaINmjvpdB0qsXn5V0kOE25 Y9xtGT8bei8z8Py83pfeYapgMzPT1/J348/bf5S/VFGcNIlliebsbk8AViEAxnPvC+Un wZ2w== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787989717; x=1788594517; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=K0brm21TE9GssWqZ74Js94Ds6AVGqAHayV0XzsofN18=; b=QNUmUIh0pQvTm9aifpxQWRVTjB6EiF8rfeKYxd6459CnFf/T14TBWLwB96dWug3RLX U+dVvEvC1C8Y9IcIKcB5+NuNBWcJ3Lk9w7fe3rvwYzTicxVpwICIssZD+9X70HWrG+wH J6UJSQ4k2bA5YsiOQ2ywrHiZIzOwdCrv/SbrydOVwdEyvHo6I9OPJ5IWp4ilR2b261c1 91LexGVp4+rb9OudTLLbvb/wD13dIEwU8yGydWDkccpPmh/KBuELFjZSp1vPr4sG8puD nZec/BIZmmqa8Z69rDQjQHkYkyKgR8psV9S4IJQDruyKm4YVROmnkhdOAabvhERbOFOv fPig== X-Gm-Message-State: AFuF++nxZVNh5SEreND9wwhk/JJTGpk81kDxPcOojtrq7YAAQP6mUaHs 50eb+eogBCsTgF5a/LVgh28LLFaAs8wa6WqRV4v5sUkjXhu9yan0DeTD X-Gm-Gg: AYBFou3Y0I3tnZXa9luJ1X7CEKFqe1RirBO+J8HwanU18TGUiBLB2ZeIjnXpGPkZnvC YxigTW29C7H0uO9rg6nEdZpZjfkhPthK3XUm8ff6IfrT+1ZZKXctc0uH1ewdfliFmQe4QBMwdeC sRwHQ/vv3MdyV8lv6KMVFkPffzPUtEQ8wJ3xiwuC5ed5HU+b93bLY+LK+FNFO9pkeN/8694UIf2 RwTVysk0LnZ4KqoSTnxcqAba34JCE5rWYW5DJSalxRbrb2gGnXH/dAM/z2/DOhXRkf19Bvkmnwd mG1sNuo/rgHVlJZXQLIT0xnsVlojeKV3kesJV2gy1aRGv9y57Tf4aDW9dT/jaOclEpACDy6LV5s AkaIzyCm4YyoY/6Wmt/2rjRUjK5pb0/Y6P6LkXOI9zIRwhfHe/vyNwV//UAwXYxQEk2CQEgbFLr WorabrlIMZptikMYqXV1WRmR126v+LopAIdKprC8GBxrGclszVzEUKSkU5d/pyx9zsVtSpMrC1c N3RgOOig1MwJKPVHTn0JvAhnJR2 X-Received: by 2002:a17:90b:37cc:b0:366:10f1:3d91 with SMTP id 98e67ed59e1d1-396d0d92d0amr23218510a91.1.1787989716688; Sat, 29 Aug 2026 00:48:36 -0700 (PDT) Received: from localhost.localdomain (vmi2317720.contaboserver.net. [84.247.152.65]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-396ddca5bfbsm6256974a91.17.2026.08.29.00.48.32 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Sat, 29 Aug 2026 00:48:36 -0700 (PDT) From: "Lian Wang (ProcessMission)" To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, Andrew Morton , Chris Li , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , Kemeng Shi , Kairui Song , Jihan LIN , Kunwu Chan Subject: [RESEND RFC PATCH v2 12/13] mm/swap: drop swap active plist Date: Sat, 29 Aug 2026 15:48:28 +0800 Message-ID: <20260829-swap-pcp-priq-v2-resend-12-68d3d925578c@gmail.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260829-swap-pcp-priq-v2-resend-0-68d3d925578c@gmail.com> References: <20260829-swap-pcp-priq-v2-resend-0-68d3d925578c@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Kairui Song The swap_active_head plist tracked all writable swap devices, ordered by priority. Now that every consumer that needed priority ordering has moved to the priority queue, the remaining users only need to find any writable device. Remove swap_active_head, the associated plist_node from swap_info_struct, and migrate the remaining plist operations. Signed-off-by: Kairui Song Signed-off-by: Lian Wang (ProcessMission) Tested-by: Kunwu Chan Tested-by: Kunwu Chan --- include/linux/swap.h | 1 - mm/swapfile.c | 27 +++++---------------------- 2 files changed, 5 insertions(+), 23 deletions(-) diff --git a/include/linux/swap.h b/include/linux/swap.h index 0c8c1ebcab9c..9d86db144b94 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -267,7 +267,6 @@ struct swap_info_struct { struct percpu_ref users; /* indicate and keep swap device valid. */ unsigned long flags; /* SWP_USED etc: see above */ signed short prio; /* swap priority of this type */ - struct plist_node list; /* entry in swap_active_head */ signed char type; /* strange name for an index */ unsigned int max; /* size of this swap device */ struct swap_cluster_info *cluster_info; /* cluster info. Only for SSD */ diff --git a/mm/swapfile.c b/mm/swapfile.c index cb27f2e246f0..0acb1f31df8c 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -41,7 +41,6 @@ #include #include #include -#include =20 #include #include @@ -90,12 +89,6 @@ bool swap_migration_ad_supported; static const char Bad_file[] =3D "Bad swap file entry "; static const char Bad_offset[] =3D "Bad swap offset entry "; =20 -/* - * all active swap_info_structs - * protected with swapon_rwsem, and ordered by priority. - */ -static PLIST_HEAD(swap_active_head); - static inline struct swap_info_struct *__swap_iter(int *i, unsigned long f= lag) { lockdep_assert_held(&swapon_rwsem); @@ -3262,7 +3255,6 @@ static void __swap_device_enable(struct swap_info_str= uct *si) spin_unlock(&swap_queue_update_lock); atomic_long_add(si->pages, &nr_swap_pages); total_swap_pages +=3D si->pages; - plist_add(&si->list, &swap_active_head); add_to_avail_list(si); } =20 @@ -3282,9 +3274,8 @@ static int swap_device_disable(struct address_space *= mapping, int err =3D -EINVAL; =20 percpu_down_write(&swapon_rwsem); - plist_for_each_entry(si, &swap_active_head, list) { - if ((si->flags & SWP_WRITEOK) && - si->swap_file->f_mapping =3D=3D mapping) { + for_each_avail_swap(si) { + if (si->swap_file->f_mapping =3D=3D mapping) { err =3D 0; break; } @@ -3309,7 +3300,6 @@ static int swap_device_disable(struct address_space *= mapping, spin_lock(&swap_queue_update_lock); si->flags &=3D ~SWP_WRITEOK; spin_unlock(&swap_queue_update_lock); - plist_del(&si->list, &swap_active_head); total_swap_pages -=3D si->pages; atomic_long_sub(si->pages, &nr_swap_pages); =20 @@ -3641,7 +3631,6 @@ static struct swap_info_struct *alloc_swap_info(void) */ } p->swap_extent_root =3D RB_ROOT; - plist_node_init(&p->list, 0); p->flags =3D SWP_USED; percpu_up_write(&swapon_rwsem); if (defer) { @@ -4050,12 +4039,7 @@ SYSCALL_DEFINE2(swapon, const char __user *, special= file, int, swap_flags) if (swap_flags & SWAP_FLAG_PREFER) prio =3D swap_flags & SWAP_FLAG_PRIO_MASK; =20 - /* - * The plist prio is negated because plist ordering is - * low-to-high, while swap ordering is high-to-low - */ si->prio =3D prio; - si->list.prio =3D -si->prio; =20 /* * Publish swap_file before making the percpu ref live, then add the devi= ce @@ -4176,7 +4160,7 @@ int swap_dup_entry_direct(swp_entry_t entry) #if defined(CONFIG_MEMCG) && defined(CONFIG_BLK_CGROUP) static bool __has_usable_swap(void) { - return !plist_head_empty(&swap_active_head); + return READ_ONCE(total_swap_pages) > 0; } =20 void __folio_throttle_swaprate(struct folio *folio, gfp_t gfp) @@ -4200,9 +4184,8 @@ void __folio_throttle_swaprate(struct folio *folio, g= fp_t gfp) return; =20 percpu_down_read(&swapon_rwsem); - plist_for_each_entry(si, &swap_active_head, list) { - if ((si->flags & SWP_WRITEOK) && - !(atomic_long_read(&si->inuse_pages) & + for_each_avail_swap(si) { + if (!(atomic_long_read(&si->inuse_pages) & SWAP_USAGE_OFFLIST_BIT) && si->bdev) { blkcg_schedule_throttle(si->bdev->bd_disk, true); break; --=20 2.55.0 From nobody Sat Sep 26 22:02:19 2026 Received: from mail-pg1-f182.google.com (mail-pg1-f182.google.com [209.85.215.182]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 78E271F09AD for ; Sat, 29 Aug 2026 07:48:46 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.182 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787989727; cv=none; b=V3SCW1gq3ENPTBz+/cc7hgsG+pxkVZxDEf3u5TEw5E44lsToHRiSXGvwaZP0z3733dcLpAKy1W8ZaVAMW+v43U6lwd5vqwouEJDGYYpDqB24H0mPsTxIDt4Kp73bUKpLI2IxQLqOMr1RdK7oqnfzJ9Ns4lJ4YpZK9xlvzQ+8vF4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787989727; c=relaxed/simple; bh=2vq6Z7ZSe/CnfQnpo6YCDKeG1V4nRowKTs6jGCSXi80=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=muYDPQ39JqOkHnGICwmxkbvGITpfxTtKbTMVCf4A2r7nQ14hyTEyexqKVupSyox1zJaqRQPGxyqtHE0gNnJHBMuL3+HrM+2y4wJnlB5eyxsym4a8CVLYQB3uVPLkKpEQfl0EAIy0KQ+R2C3POfoS8ubgkdhbA4Vwdxtz+SM4SeQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=bX7hPMNS; arc=none smtp.client-ip=209.85.215.182 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="bX7hPMNS" Received: by mail-pg1-f182.google.com with SMTP id 41be03b00d2f7-cbee846deecso2390094a12.1 for ; Sat, 29 Aug 2026 00:48:46 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787989726; x=1788594526; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=V58LX9cQfAsn8GO5XE1VypHteINKmbZsjfFFi0DDS4w=; b=bX7hPMNS6+f4keV19vGL3GBAe7mrQWpZxBE39+w5cjOEplWg2tv0is0stXdmqdo/rY Et1ohiGC8KdPviWSAVltzYEuK3sPvaBzhNpFU/RsYT3ak0vdGqPsCexM2cTjK572AB4e 8wsX+P6Qn3h5nNrCOKmSts1LayQuuwbDxIdCYRzGlj+/rSHoJkyH0YqgtQ1BfhgG+XPw yqEWQ9Bb8CWkGPGOXeGpWiWY49nkLkPDiB25uiKV6LUsHiOdhLqyv3pGhuFPnd4nJ/wW FHoQfczvqtySzKIfnz7oSZTjc8B8syUE+BLkAFdDb3UiT5PxomdS0kGcvvxbHORhEDwa zrkQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787989726; x=1788594526; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=V58LX9cQfAsn8GO5XE1VypHteINKmbZsjfFFi0DDS4w=; b=bjhjVVS3PoDJ3PwabFnUNOGiOQKPNZ4zI5Taw3OWiOLmsnqPxF3qtI6ucHMb6KAnOf /QXd1y8P3Vkv/fzrGvdFYlh6A+Ihi5ORqYDjpe/TdOQIlg56VehROmSPumSoVLburkvj 8Tta5gDxiJeKMDdlHY8wDIRKj+3225yj44t009PQKk/Gj2pT8h2+ZJJJuN91vNYljdNP c4ckhf2HN2L4rynelBZNbnDbsuFpwcK32OKaPfqQuWApvEnOhX4gHKZNbFNE+y0OGdNI vAqWMWAPNcrsZtDsA0JFOby493FGtegMzQjCTERwRB7D0P8R3tMQk7EgdBELbqxPiEzR JUtQ== X-Gm-Message-State: AFuF++k9UamwktNMloKFj7clXZBrncy/t17nlV1D27MhnmAp4wVVOgEn /8BnNJ5qdfzuMWwmtMF1oej+dOzYuqQgL0RzRWYqwscSLvzLiiimyf7m X-Gm-Gg: AYBFou3vpiA/1f2I670gi9eXYuvddjmVoKE2ygAVBvFCLi1GKZ2rb30R8EfW1+dZCyh qngOmmq2Gq7UKq8nI0Egn9kWnVWbam2fHHkqc1vnImq1oA8zJiQpxIvMTJ5HdikmsHOdMLL+jG/ JJgdB82UYX9J/yoeepAah4ikzseC0wx5Uf6eAThR9cqBpJob3Pt+YlXJWCaIH9NMQZeyvVa5H8J brdpOiUOOgUV88RQs+PzbBj8H/5edPojpDgGtv/aMKibfcQGbXw/arbshamO9Ycyv1gAO0agxYt c8oH4Wf/un1WSz/0VvxTDZjYghu8+lcFXGRR6shYmMSb8yzV5bxV3TBJ3rotI2qCgb+GFX2qV4C ZGL1wkJ7aMka/rgHjLRlDuVPBd3RXx9SzwSg1nUIilvZvRZYvLbdB2iQybk2mbabsz0oilFDBzM zvADQWpgfvFs6PfUkOvtzu2MozX8d2cQ/WlSd4AWmv5j6U9/htX+LrbBCfrMmWmXtfWPwsZ2jaF 84Y4BkUlrpDZdjFGZgyCp0jeC12 X-Received: by 2002:a17:90a:fc4e:b0:395:5eec:b932 with SMTP id 98e67ed59e1d1-396d0f8fdbemr24035986a91.11.1787989725941; Sat, 29 Aug 2026 00:48:45 -0700 (PDT) Received: from localhost.localdomain (vmi2317720.contaboserver.net. [84.247.152.65]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-396b186809dsm10748234a91.10.2026.08.29.00.48.41 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Sat, 29 Aug 2026 00:48:45 -0700 (PDT) From: "Lian Wang (ProcessMission)" To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, Andrew Morton , Chris Li , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , Kemeng Shi , Kairui Song , Jihan LIN , Kunwu Chan , Thomas Gleixner Subject: [RESEND RFC PATCH v2 13/13] lib/plist.c: remove requeue function Date: Sat, 29 Aug 2026 15:48:37 +0800 Message-ID: <20260829-swap-pcp-priq-v2-resend-13-68d3d925578c@gmail.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260829-swap-pcp-priq-v2-resend-0-68d3d925578c@gmail.com> References: <20260829-swap-pcp-priq-v2-resend-0-68d3d925578c@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Kairui Song Now the last user of plist requeue is gone, this function can be removed. Signed-off-by: Kairui Song Signed-off-by: Lian Wang (ProcessMission) Tested-by: Kunwu Chan Tested-by: Kunwu Chan --- include/linux/plist.h | 2 -- lib/plist.c | 64 ------------------------------------------- 2 files changed, 66 deletions(-) diff --git a/include/linux/plist.h b/include/linux/plist.h index 16cf4355b5c1..73d359c1a053 100644 --- a/include/linux/plist.h +++ b/include/linux/plist.h @@ -132,8 +132,6 @@ static inline void plist_node_init(struct plist_node *n= ode, int prio) extern void plist_add(struct plist_node *node, struct plist_head *head); extern void plist_del(struct plist_node *node, struct plist_head *head); =20 -extern void plist_requeue(struct plist_node *node, struct plist_head *head= ); - /** * plist_for_each - iterate over the plist * @pos: the type * to use as a loop counter diff --git a/lib/plist.c b/lib/plist.c index a5bef38add43..b05ae7ffea87 100644 --- a/lib/plist.c +++ b/lib/plist.c @@ -142,58 +142,6 @@ void plist_del(struct plist_node *node, struct plist_h= ead *head) plist_check_head(head); } =20 -/** - * plist_requeue - Requeue @node at end of same-prio entries. - * - * This is essentially an optimized plist_del() followed by - * plist_add(). It moves an entry already in the plist to - * after any other same-priority entries. - * - * @node: &struct plist_node pointer - entry to be moved - * @head: &struct plist_head pointer - list head - */ -void plist_requeue(struct plist_node *node, struct plist_head *head) -{ - struct plist_node *iter; - struct list_head *node_next =3D &head->node_list; - - plist_check_head(head); - BUG_ON(plist_head_empty(head)); - BUG_ON(plist_node_empty(node)); - - if (node =3D=3D plist_last(head)) - return; - - iter =3D plist_next(node); - - if (node->prio !=3D iter->prio) - return; - - plist_del(node, head); - - /* - * After plist_del(), iter is the replacement of the node. If the node - * was on prio_list, take shortcut to find node_next instead of looping. - */ - if (!list_empty(&iter->prio_list)) { - iter =3D list_entry(iter->prio_list.next, struct plist_node, - prio_list); - node_next =3D &iter->node_list; - goto queue; - } - - plist_for_each_continue(iter, head) { - if (node->prio !=3D iter->prio) { - node_next =3D &iter->node_list; - break; - } - } -queue: - list_add_tail(&node->node_list, node_next); - - plist_check_head(head); -} - #ifdef CONFIG_DEBUG_PLIST #include #include @@ -231,14 +179,6 @@ static void __init plist_test_check(int nr_expect) BUG_ON(prio_pos->prio_list.next !=3D &first->prio_list); } =20 -static void __init plist_test_requeue(struct plist_node *node) -{ - plist_requeue(node, &test_head); - - if (node !=3D plist_last(&test_head)) - BUG_ON(node->prio =3D=3D plist_next(node)->prio); -} - static int __init plist_test(void) { int nr_expect =3D 0, i, loop; @@ -262,10 +202,6 @@ static int __init plist_test(void) nr_expect--; } plist_test_check(nr_expect); - if (!plist_node_empty(test_node + i)) { - plist_test_requeue(test_node + i); - plist_test_check(nr_expect); - } } =20 for (i =3D 0; i < ARRAY_SIZE(test_node); i++) { --=20 2.55.0