From nobody Tue Sep 29 02:02:39 2026 Received: from mail-wm1-f100.google.com (mail-wm1-f100.google.com [209.85.128.100]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4E7CE353A68 for ; Thu, 13 Aug 2026 13:58:39 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.100 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786629522; cv=none; b=PlBeKzG85HgeOb3lRFI5/UPFZKNzAdLyaco10P9dJ/vsW7yWIm2q6SqhWmojBcYxW3IO4hy6FpeCLfvqPRcMDs9bHn7zDK2UwwZatbOFVg6wnPbH9feEj/5MrKyyW5Aevh5XkoYS/sEVV/bymIHdLDhdevlVRW5sI4fkYC5YY+8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786629522; c=relaxed/simple; bh=InMdX+pdiDloclCPV/zVlqIYNtiNdrfnmA7sVFiIrcI=; h=From:To:Cc:Subject:Date:Message-Id:MIME-Version; b=LYTFlVXDFWMGPSUdue8mKMhOBnV0mnp0VkewIrRVpN1mLwVP8Owtgj6S5N69AMvRNtV0fumoUTang5nwJaidmqo1eUDm86n4J5wjeL12v19P2aSe0DcUeckQQUTahyahc+WF8b/QNDToUFDNbmqCyJAGLxdLJgylsG8KBYtNAuc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=purestorage.com; spf=pass smtp.mailfrom=purestorage.com; dkim=pass (2048-bit key) header.d=purestorage.com header.i=@purestorage.com header.b=SE2h3KT0; arc=none smtp.client-ip=209.85.128.100 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=purestorage.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=purestorage.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=purestorage.com header.i=@purestorage.com header.b="SE2h3KT0" Received: by mail-wm1-f100.google.com with SMTP id 5b1f17b1804b1-4954f5e8020so10646845e9.2 for ; Thu, 13 Aug 2026 06:58:39 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=purestorage.com; s=google2022; t=1786629518; x=1787234318; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to:content-type; bh=1y55hehoVssabJRQyQs56fbqkSj+SkZHKxYRNj9+z/c=; b=SE2h3KT0c+mQiD2kFkBdK0wfk3IqXx83TCjYMmV+5emoz14Rg9AGj0diRBn/EVH77B GUFpNja0MJ9OO61O6JtxS1f+031i/qhtrGicMNMpaNchR7eUGK2GwMTEwM+Vnqz0sJee M3YvaJIQg1NLAxQPBYQaHdSQ6Bq5GPHOA/BUTzp1xpq+6M9eJ5UVOgzhYIHTGhFxBtnr oKMKfk5u09qhQ7YiCgBjYH+Oe+NEumq5t+6FL+DMTNex4itLvdXAAbIDzH8xR8Dr9ovD Gj58u4lfUKDRHDAmgnKZlAt3UVEEE4vQi87v8yf4Qwyn1sQWfEF3PzYhYuO74Qeqt3Ik 2VWA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786629518; x=1787234318; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=1y55hehoVssabJRQyQs56fbqkSj+SkZHKxYRNj9+z/c=; b=RJg6iuf5cj49mxVEez91lfm+J0j2QW4IvRBLOpwjMIm0fCFJOMGSDI3c02212uGmp5 ihok915f34xiXsp0rbljJ+dZsvDOA7B9W5egKHQXGQqTQlbjHIE7ilk74cj2SaXW7sqX CqAflkmfhc6wlfmfU/eYRyNfkZuUHmDJyZwfjugkRyZmnUJq8BdfzCkNK7BTB6tO+oT0 KjPSoecGFXgzHvcYMsgfsccLPVbUBDrpu+vzUjxHXdBlDkw3pBMAYw+dJFcUHkR0kw/q Ct+xEbYswm/M1o6ow2QRGSVJtwBNkS5Qhk866UCcLk+3KUrLImJ7uUhSR5NcLfF0yNnN 8ULA== X-Forwarded-Encrypted: i=1; AHgh+RpAVEupOP8usHKOxLljJqkYN03+fDhzzmiZNfsXi+WF5IxibK7WHJ3PL1XVXLWFCjGn+XL5bJG5pXQTkq4=@vger.kernel.org X-Gm-Message-State: AOJu0YyN65zExJ89xj8CDYGJ0vj6V9qfLx1dUslpcPbXp1BkCTVAkK7n 1YO22ftoT/vcjSSTKKIHbQwZn0CC2Nu1ePZp0YYoYrUrMluN6ch6p8YVnKpJfHuv2+QTdL0lasC 97WXAz5HTVU1q1uKRzvQKBPLWKF00FyS+rjYAB19ib2j3PcKGZnZk X-Gm-Gg: AR+sD11W01Fp5GQAYmFRflK2JBObn5INjgaXaBC1q9tq4V+vmAQOIbnnUx/ZTFl/CYy xROcSRB3YVD0YRrouR+FLZMrh8FxGMhu1B4JjcnO0RhfOrH+NANQtOzwcUR6wiv4K2PqOnhm5KM SNdWqUuLgqBgdHwHLLzrGpoqxcZfP4Pf80dHqohxGAZURzP+cprh6CPeMQ04yZ8Limnl7n+DVCb By+5/H844AAmiO6Brl3Vl1meJGupnTJ2BFf6Rbx2cYcEbKHvT82wu5eU5SSOnBNHvnPb0W6BLW0 Pp30YHI210ev7eENzp+ZJ38tPi89/5TOhwbJk9AB9nU+zvGHxBEEPn1yisca4gaccuZXwoSSH4K MV1/VGsWGSS952Jpzc9OlJ4BEOzEq X-Received: by 2002:a05:600c:1385:b0:495:5fdf:2075 with SMTP id 5b1f17b1804b1-49982109cc1mr75225985e9.0.1786629517600; Thu, 13 Aug 2026 06:58:37 -0700 (PDT) Received: from c14-smtp-2023.dev.purestorage.com ([2620:125:9008:8012:36:131:5:0]) by smtp-relay.gmail.com with ESMTPS id 5b1f17b1804b1-49981b16dd0sm9042585e9.7.2026.08.13.06.58.37 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 13 Aug 2026 06:58:37 -0700 (PDT) X-Relaying-Domain: purestorage.com Received: from irdv-tmenninger.dev.purestorage.com (irdv-tmenninger.dev.purestorage.com [10.32.149.15]) by c14-smtp-2023.dev.purestorage.com (Postfix) with ESMTPS id 3A43634029C; Thu, 13 Aug 2026 06:58:36 -0700 (PDT) From: tmenninger@purestorage.com To: Trond Myklebust , Anna Schumaker Cc: linux-nfs@vger.kernel.org, linux-kernel@vger.kernel.org, Tim Menninger , stable@vger.kernel.org Subject: [PATCH] pNFS: Check lseg validity before marking a layout for return Date: Thu, 13 Aug 2026 13:58:34 +0000 Message-Id: <20260813135834.22278-1-tmenninger@purestorage.com> X-Mailer: git-send-email 2.34.1 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Tim Menninger pnfs_error_mark_layout_for_return() receives the lseg associated with the failed I/O but previously used only its I/O mode, operating on the inode's current layout header regardless of whether the lseg itself was still valid. A layout stateid can be invalidated while RPCs still hold references to its lsegs. pnfs_mark_layout_stateid_invalid() clears NFS_LSEG_VALID on those lsegs through pnfs_clear_lseg_state(). A subsequent LAYOUTGET can install a replacement stateid in the same pnfs_layout_hdr. If an RPC using one of the old lsegs later reports an error, the current code can therefore mark the replacement layout for return. Once NFS_LSEG_VALID has been cleared, the lseg is no longer eligible for selection for new I/O and must not initiate another error-driven return of the inode's current layout. Fold pnfs_mark_layout_for_return() into pnfs_error_mark_layout_for_return() and check pnfs_is_valid_lseg() alongside pnfs_layout_is_valid() before setting return info, so a stale lseg cannot drive an error return of the current layout. This was reproduced by restarting a FlexFiles data server during a high-throughput read workload. Stale lsegs repeatedly caused the replacement layout to be marked for return, triggering I/O cancellation, RPC/RDMA transport reconnects, and sustained contention on inode->i_lock. The client did not recover on its own and consumed about 94 CPU cores on a 96-CPU system. With this change, I/O recovered within about 20 seconds and recovery load peaked at about 15 CPU cores. Cc: stable@vger.kernel.org Signed-off-by: Tim Menninger --- fs/nfs/pnfs.c | 28 ++++++++++------------------ 1 file changed, 10 insertions(+), 18 deletions(-) diff --git a/fs/nfs/pnfs.c b/fs/nfs/pnfs.c index 7715e2bd5871..9f32dd7c4c63 100644 --- a/fs/nfs/pnfs.c +++ b/fs/nfs/pnfs.c @@ -2708,26 +2708,30 @@ pnfs_mark_matching_lsegs_return(struct pnfs_layout_= hdr *lo, return -ENOENT; } =20 -static void -pnfs_mark_layout_for_return(struct inode *inode, - const struct pnfs_layout_range *range) +void pnfs_error_mark_layout_for_return(struct inode *inode, + struct pnfs_layout_segment *lseg) { struct pnfs_layout_hdr *lo; bool return_now =3D false; + struct pnfs_layout_range range =3D { + .iomode =3D lseg->pls_range.iomode, + .offset =3D 0, + .length =3D NFS4_MAX_UINT64, + }; =20 spin_lock(&inode->i_lock); lo =3D NFS_I(inode)->layout; - if (!pnfs_layout_is_valid(lo)) { + if (!pnfs_layout_is_valid(lo) || !pnfs_is_valid_lseg(lseg)) { spin_unlock(&inode->i_lock); return; } - pnfs_set_plh_return_info(lo, range->iomode, 0); + pnfs_set_plh_return_info(lo, range.iomode, 0); /* * mark all matching lsegs so that we are sure to have no live * segments at hand when sending layoutreturn. See pnfs_put_lseg() * for how it works. */ - if (pnfs_mark_matching_lsegs_return(lo, &lo->plh_return_segs, range, 0) != =3D -EBUSY) { + if (pnfs_mark_matching_lsegs_return(lo, &lo->plh_return_segs, &range, 0) = !=3D -EBUSY) { const struct cred *cred; nfs4_stateid stateid; enum pnfs_iomode iomode; @@ -2742,18 +2746,6 @@ pnfs_mark_layout_for_return(struct inode *inode, nfs_commit_inode(inode, 0); } } - -void pnfs_error_mark_layout_for_return(struct inode *inode, - struct pnfs_layout_segment *lseg) -{ - struct pnfs_layout_range range =3D { - .iomode =3D lseg->pls_range.iomode, - .offset =3D 0, - .length =3D NFS4_MAX_UINT64, - }; - - pnfs_mark_layout_for_return(inode, &range); -} EXPORT_SYMBOL_GPL(pnfs_error_mark_layout_for_return); =20 static bool --=20 2.34.1