From nobody Thu Sep 24 16:07:44 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DDAF13BB13A; Tue, 22 Sep 2026 11:34:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790076862; cv=none; b=c/N24lnk7/hl33W0aNOPsHLCy6lsuaDbYmkMPLVQx6mLSBS9OJJgPCmOHTwEzLBdU1kEdcoXO2RyaZKN/r5+t+mAJxqdbMa6TFqYQxawqHs0qkePG+dCwATVEZT2jVNvKVrYeSo9ZMBE8ehlK6BKYKn/ktfQb+0PsEEwsZfTiHE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790076862; c=relaxed/simple; bh=zCR/opn7RVI4+wfWUxanS7wEtFBcAfR88rNRJVt+grk=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=LVJENpdkc9qOzjUTB/7h4WeUMrSzQBbBuOZsW9DL0A83TDzsdR37htBP5nICbLeJpnponna9wkxN6T3bW+WSXwWJ93C4A3Hdnnpc99eCBBTZ8caDP7Lp4/l4cyuQLRWsd1atg8Txf0+hENNFrpwCWBMFNIlT3077w3AwX3fNOGk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=eWCJ6sBi; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="eWCJ6sBi" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 9E9E11F000FF; Tue, 22 Sep 2026 11:34:19 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1790076860; bh=jzSoL3QvVgF+2xSUs9iiTKgvCrgelMiWawC+3zEoW/0=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=eWCJ6sBiOdlMLL+N/z+7zOLvkAFrvJU5F1uWDhvM0vRYYI1Enac9B28S/cHDn6VK3 iVWDRc+DQU7arC42sxdA6AVuiIAHhMaxDKMy2g2s+x2Q8qWW+Lo/jlosThJ1oqFLJx t3MXX/bBixboGPPCMq3zUxvmlZmRT0zwWYgZ1FAKOM3TK0wp0wic2h0jS6O9CfpseP L2nYBDrSTTVjJGlk2YgR7Pn7GozvCX4Ced0KeCraTw9HuamG15L1wx3hQyriIXvpi9 CGuV9drC0i5HC/cp5g389uIF8Yq9xNSbbakkNaq7NHi96AnIjTppvx2QNWn5uB7Qpz vRtQQnTp+YSDQ== From: Jeff Layton Date: Tue, 22 Sep 2026 07:34:02 -0400 Subject: [PATCH v2 1/8] nfsd: clear NFSD_NET_UP before dropping the generic resource reference Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260922-nfsd-per-net-mutex-v2-1-4a5da8243a73@kernel.org> References: <20260922-nfsd-per-net-mutex-v2-0-4a5da8243a73@kernel.org> In-Reply-To: <20260922-nfsd-per-net-mutex-v2-0-4a5da8243a73@kernel.org> To: Chuck Lever , NeilBrown , Olga Kornievskaia , Dai Ngo , Tom Talpey , Shuah Khan Cc: linux-nfs@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, Jeff Layton X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=openpgp-sha256; l=1304; i=jlayton@kernel.org; h=from:subject:message-id; bh=zCR/opn7RVI4+wfWUxanS7wEtFBcAfR88rNRJVt+grk=; b=owEBbQKS/ZANAwAKAQAOaEEZVoIVAcsmYgBqsme4Zq8r7rpxrReBlyDJQLFo6L2pWeRg4gR/x vKfyFY3L6yJAjMEAAEKAB0WIQRLwNeyRHGyoYTq9dMADmhBGVaCFQUCarJnuAAKCRAADmhBGVaC FSLgD/0ZcAPRcHUbfk/YEfFxEUitUPai++jXThCJ739AubMcTUb9rywL+mS+ccKdIi0urni8uj5 TqWEe5CMnD3b2QOMX8aVRhPMVo4POxR5w6LnT+iyeMP3erfbOuogh1rHRA6SswnP7fyLrjoh980 6aJzIlr68FvZKFmVFH2i8vh7oEKN9bLuX2MKm3X6u7c5vucvwYP0+KvAru43J8c14LpDhonWyEp LJPlW0Ehw6G09Ks81/LMmAaBgOm4TdrKkw80yxBTDdbiEgsB0ZsJswzNTIRoV688IVGvkMoSs// KZKO6T8gvOn6yBm9Ezqu8AigWa3VJPdol2xdlG8uW7CDNiZ3VHThU35cDReuwp+5ufu/64DaxrH uA9P+GOTjeWUXRp9A1ht8yU44t4fcCvM75xn9+AX25ZIbSvDHdxL79shksX+bB1w/MTRnlOgtnx fpGg8aIkwvcoUhipN3j4k8VKNov9bKnv+zEia0Un05lM2EdXgBP+9zsmlyEgqGqoxIDrDpsXqIl FamShhaz7/cJf2hU/GpKtJW0mvO+mAn9KvI21UerhM2b3FeV4DGpY4BXQOeilJKxx+bIg0Jdi0N 1+z+q+u6DPAGsoK1ycAWKoBbm999QOUW2Upk7snH1msOGL/sMSgZsLWpyYSuCspA5iWydmReOWH lTa6KYSm4ak/Q5Q== X-Developer-Key: i=jlayton@kernel.org; a=openpgp; fpr=4BC0D7B24471B2A184EAF5D3000E684119568215 nfsd_shutdown_net() leaves NFSD_NET_UP set across nfsd_shutdown_generic(), so during that window the bit claims the host-wide file cache and NFSv4 tables are up when the last namespace may already have torn them down. Nothing observes it today only because one global mutex covers both. Fold the test and clear into test_and_clear_bit() so the invariant "NFSD_NET_UP implies the generic resources are up" holds on its own. No functional change. Assisted-by: LLM Signed-off-by: Jeff Layton --- fs/nfsd/nfssvc.c | 8 ++++++-- 1 file changed, 6 insertions(+), 2 deletions(-) diff --git a/fs/nfsd/nfssvc.c b/fs/nfsd/nfssvc.c index cbc989238710..d75fe1523436 100644 --- a/fs/nfsd/nfssvc.c +++ b/fs/nfsd/nfssvc.c @@ -433,9 +433,13 @@ static void nfsd_shutdown_net(struct net *net) =20 percpu_ref_exit(&nn->nfsd_net_ref); =20 - if (test_bit(NFSD_NET_UP, &nn->flags)) + /* + * Clear NFSD_NET_UP before dropping this namespace's reference on + * the generic (host-wide) resources, so that the bit never claims + * they are available once they may already be gone. + */ + if (test_and_clear_bit(NFSD_NET_UP, &nn->flags)) nfsd_shutdown_generic(); - clear_bit(NFSD_NET_UP, &nn->flags); } =20 static DEFINE_SPINLOCK(nfsd_notifier_lock); --=20 2.55.0 From nobody Thu Sep 24 16:07:44 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1767047CA95; Tue, 22 Sep 2026 11:34:21 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790076863; cv=none; b=KqsVqvhoqhHkbTH9Oq2DTYTofchFvuDqRQWVPpD9LLyhcWOEVJHLAZiWWGfnUgXdh/AbizOvWaahkIjZUwDRmDURQJDxSX3I09N7QRAO/IGWwo85YnXZjKOWs80Pd6KHFS+u3FjtUCQCrGe/FcBKTeh04w4JCELGlnq5yQil+Qk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790076863; c=relaxed/simple; bh=ImZu2rrTp3BgaUjvgrhT8+XgUZVUh9pI+IzGxU0cFQ8=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=FFn4MTvXAoZZJPWjLno9/JBjJ1gzd5e8YlOeokLts1++x7fC8Imb/xhhRx/HQnNJVshXcnpeBpAki61GSR+Rtav/gCC/C23EcjqAqj8Ki3XuB0Uik40gsEOdEWYTOr9rvHbhipUjNW058ZqH9QcFLBliB5SpejIpv13yi6n39GM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=AFnk/Gff; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="AFnk/Gff" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 9B6891F00898; Tue, 22 Sep 2026 11:34:20 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1790076861; bh=LsBdDva0KrdXCVvAA86ZG5+kA6vb0RfnH5D+cyesTgw=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=AFnk/GffxQ7fk6dYKEbnD1VVzyhd4LN/JNK/WaqMCPNS5DnESrdK5YKjmapjLPHJr auybBiDR4P09oGAIpeU1dSmqwMKm7MFLIgW83yUI5ss/1TxYbLwAkaurBA/Xq89kq3 D1y3Ayeij10igSB4fasuqbGxnf71biRrj4Kdj4eJX2QvY7gItevlYtw5OJy8FTeVBI +7YgeVDOctdmX48ngGXIV3Hlw8HB+pf1K9zF+d9IT3Otfk4qMd8k8rwA4zPzbFgqUf kGo4CLeH4G9pKdKCQJrkZ4K4dWtJUaR7+TQTdHSH7yVFqAXLyYpGIFRMF2eVCl6uPf Io3ro5/3JpG1g== From: Jeff Layton Date: Tue, 22 Sep 2026 07:34:03 -0400 Subject: [PATCH v2 2/8] nfsd: make max_blksize a per-namespace setting Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260922-nfsd-per-net-mutex-v2-2-4a5da8243a73@kernel.org> References: <20260922-nfsd-per-net-mutex-v2-0-4a5da8243a73@kernel.org> In-Reply-To: <20260922-nfsd-per-net-mutex-v2-0-4a5da8243a73@kernel.org> To: Chuck Lever , NeilBrown , Olga Kornievskaia , Dai Ngo , Tom Talpey , Shuah Khan Cc: linux-nfs@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, Jeff Layton X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=openpgp-sha256; l=3112; i=jlayton@kernel.org; h=from:subject:message-id; bh=ImZu2rrTp3BgaUjvgrhT8+XgUZVUh9pI+IzGxU0cFQ8=; b=owEBbQKS/ZANAwAKAQAOaEEZVoIVAcsmYgBqsme4X9HBu2mmnhbrrA8HOPjrhGAmRdUpIB1R2 2GS6zffUjOJAjMEAAEKAB0WIQRLwNeyRHGyoYTq9dMADmhBGVaCFQUCarJnuAAKCRAADmhBGVaC Fd/mEADIbLTXkkdy7b8rjk+6j854vsPUNN8dMFng7Quz8xPCs/PsCeDJD5HEXC7/LlrmF0fcYoq 0UbKmjKrUx93yTwu/qYT3Ki5fRN/7T8Z7P5e9tU4zIhWZOXz/LMS91dY7sq20aNfL6NVwEW9wYb geigzwLPplOULxAu1wLaTn/1Cc/xT3Jy5kT+7zv92ZhJu7O64Q5Gb0kbgIhrXasEra3lSZVFJJc +qQUecmB4R+XHMSvK0c3KN9Nxwuqgne78iAmQKIk7qsCga3NJnHIf1FK9rh2B+3BcywdKqdSW5+ Duz1eve01bP548EH11At7ynJ7Pe+m47s6n3MEI4+FZWIGkYtsslclLk5dzKDz3e/DoUc4VUJF0q 36bm9M3ZmAFmD4HHkGVB1Sfgj2nva+r03iSnFi19b5PwDIR1/otAAerdr+5pcrJBCX1ESztY0Vm ymVeh+CzYdkezAHNTsqEt8FcCuD9xizrG/sUoubSlzrxqirHBt67ROjelU0FS6/ODp1tOnZuxmS pTOfo0KzIqVazOYIg8368fKPv4suLfT4iOw5yUoNpH1kK7EQcFlo9RIWkGBODNTkXsN0tthxp/+ +N5zI+L2mVltr8huXlPK4CA3ePjXVD4reARFpbXeQRd2DwUVcRrH5sz2eeHvBtZIX4Jok/ITpOV p5haoA9MjvAmWZw== X-Developer-Key: i=jlayton@kernel.org; a=openpgp; fpr=4BC0D7B24471B2A184EAF5D3000E684119568215 nfsd_max_blksize is module-scope, but /proc/fs/nfsd/max_block_size is a per-netns file. A container writing it therefore changes the payload size that every other namespace's next server start will use. Move it into struct nfsd_net as ->max_blksize. The lazy nfsd_get_default_max_blksize() fill-in now happens per namespace on first nfsd_create_serv(). User-visible change: max_block_size no longer leaks across namespaces. Fixes: 11f779421a39 ("nfsd: containerize NFSd filesystem") Assisted-by: LLM Signed-off-by: Jeff Layton --- fs/nfsd/netns.h | 6 ++++++ fs/nfsd/nfsctl.c | 8 +++----- fs/nfsd/nfsd.h | 2 -- fs/nfsd/nfssvc.c | 6 +++--- 4 files changed, 12 insertions(+), 10 deletions(-) diff --git a/fs/nfsd/netns.h b/fs/nfsd/netns.h index 0ce7da20aba3..374ce83e2ba0 100644 --- a/fs/nfsd/netns.h +++ b/fs/nfsd/netns.h @@ -154,6 +154,12 @@ struct nfsd_net { */ unsigned int min_threads; =20 + /* + * Maximum size of an NFS READ or WRITE payload. Zero until the + * first server start in this namespace picks a default. + */ + unsigned int max_blksize; + u32 clientid_base; u32 clientid_counter; u32 clverifier_counter; diff --git a/fs/nfsd/nfsctl.c b/fs/nfsd/nfsctl.c index f32311f2d7cf..330d0f12e199 100644 --- a/fs/nfsd/nfsctl.c +++ b/fs/nfsd/nfsctl.c @@ -885,8 +885,6 @@ static ssize_t write_ports(struct file *file, char *buf= , size_t size) } =20 =20 -int nfsd_max_blksize; - /* * write_maxblksize - Set or report the current NFS blksize * @@ -931,12 +929,12 @@ static ssize_t write_maxblksize(struct file *file, ch= ar *buf, size_t size) mutex_unlock(&nfsd_mutex); return -EBUSY; } - nfsd_max_blksize =3D bsize; + nn->max_blksize =3D bsize; mutex_unlock(&nfsd_mutex); } =20 - return scnprintf(buf, SIMPLE_TRANSACTION_LIMIT, "%d\n", - nfsd_max_blksize); + return scnprintf(buf, SIMPLE_TRANSACTION_LIMIT, "%u\n", + nn->max_blksize); } =20 #ifdef CONFIG_NFSD_V4 diff --git a/fs/nfsd/nfsd.h b/fs/nfsd/nfsd.h index d1e413d21e76..0864d6d564e2 100644 --- a/fs/nfsd/nfsd.h +++ b/fs/nfsd/nfsd.h @@ -143,8 +143,6 @@ enum { extern u64 nfsd_io_cache_read __read_mostly; extern u64 nfsd_io_cache_write __read_mostly; =20 -extern int nfsd_max_blksize; - bool nfsd_v4client(struct svc_rqst *rqstp); =20 /* diff --git a/fs/nfsd/nfssvc.c b/fs/nfsd/nfssvc.c index d75fe1523436..77e1e6ba686d 100644 --- a/fs/nfsd/nfssvc.c +++ b/fs/nfsd/nfssvc.c @@ -630,12 +630,12 @@ int nfsd_create_serv(struct net *net, bool no_rpcbind) init_completion(&nn->nfsd_net_free_done); init_completion(&nn->nfsd_net_confirm_done); =20 - if (nfsd_max_blksize =3D=3D 0) - nfsd_max_blksize =3D nfsd_get_default_max_blksize(); + if (nn->max_blksize =3D=3D 0) + nn->max_blksize =3D nfsd_get_default_max_blksize(); nfsd_reset_versions(nn); serv =3D svc_create_pooled(nfsd_programs, ARRAY_SIZE(nfsd_programs), &nn->nfsd_svcstats, - nfsd_max_blksize, nfsd); + nn->max_blksize, nfsd); if (serv =3D=3D NULL) { percpu_ref_exit(&nn->nfsd_net_ref); return -ENOMEM; --=20 2.55.0 From nobody Thu Sep 24 16:07:44 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AA29247D936; Tue, 22 Sep 2026 11:34:22 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790076865; cv=none; b=RZYIlJSdPmHJ5zbcgd5hlfjJV0+W4n8ReSZWlwVCxkD0cOF/43r2Yh++WAKj2vt8rJ6VO401TcYShGZoIEyrA+9Wb8miLlQ3WebEV8fs7182J2pL+J3U7JAEgMhiqM3T++bCfiohFOUlo6zkU3zBS33nTHIvJLp769VE4xUgR0I= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790076865; c=relaxed/simple; bh=oP77YO0NfFcbITGL4/o1Le7e052PM++Idyl5tCNHe/k=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=roGyDgd9yNkIKT6SlNMapVXRsARAMB+ny3+SDgUwSZp3hYiWN9aoc5xQ439W1aisea23RlrT7MUdUA+r5P8irbkJIz+rMZK1bYwurVglauOD49UXGBY7szGHrkOPOS4JI5yIuqXW1U35Ki4vitQPBTm+ybJmHOBvr6ApzCgiArQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=ATpiFMGd; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="ATpiFMGd" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 98B1C1F0089D; Tue, 22 Sep 2026 11:34:21 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1790076862; bh=Rq2hjy+iWzEHQZsRdEQL+dCvx/ZmAba4xeM/zW6EjY0=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=ATpiFMGdyoBUiJkAbR2rRrUa2yyfjEDL6f0lxvhdwsTfjVPGLatpALsG85BkTGY8Z epq3JMv3gzZMOW/JB62cvz1IbBSTS+MxkndJvgpZvgC7fIX7hV2z5JbtSPRcTKcR7I Q8UlyL6bTtJUOibsIQBrPLxvkS91r/UGa5k0CuyXbVX81N8JoeHK2+z5xbeD5zix9b tdo2HUiy4BNsNdteJT9mhIPu5ODw5G9eLzKlv0Cm6sovAQcl02Bmd/5kGAIB69MMYO snQ9wkqHZceEtqmnEh5KM+KhYOJljbTA9RD9+hk6ZJDUP6OOpK/O6RlU54+tGN80q9 bwvsb6VrIXKZg== From: Jeff Layton Date: Tue, 22 Sep 2026 07:34:04 -0400 Subject: [PATCH v2 3/8] nfsd: move the control plane to a per-namespace mutex Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260922-nfsd-per-net-mutex-v2-3-4a5da8243a73@kernel.org> References: <20260922-nfsd-per-net-mutex-v2-0-4a5da8243a73@kernel.org> In-Reply-To: <20260922-nfsd-per-net-mutex-v2-0-4a5da8243a73@kernel.org> To: Chuck Lever , NeilBrown , Olga Kornievskaia , Dai Ngo , Tom Talpey , Shuah Khan Cc: linux-nfs@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, Jeff Layton X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=openpgp-sha256; l=33796; i=jlayton@kernel.org; h=from:subject:message-id; bh=oP77YO0NfFcbITGL4/o1Le7e052PM++Idyl5tCNHe/k=; b=owEBbQKS/ZANAwAKAQAOaEEZVoIVAcsmYgBqsme5fWoJ79QknL/yzfHvyVST78dMlrcpNE4Wz PplBnrQxEKJAjMEAAEKAB0WIQRLwNeyRHGyoYTq9dMADmhBGVaCFQUCarJnuQAKCRAADmhBGVaC FQfTEACSUxg8YaAzhJ6pDce5/t6uqXvbs3fXrr2jIrdIwOiPZH8ooHinhO+14q6DHh3AEY7k9Vs 0HqHGsVnOMjLtkRQ/9m2XOc/lo+vVipnxsJ5dL0ok5L1hz1+xK42KTLfS6uiKRKboHlxU0OtRH4 rfqt4EVetYlEoqF5qi/lfYOmdTsoGb+lFZugYeN0AYTeu1TpYixkOYcsuFWrGvwcIbgnVIw01/t zHc3BsJLMOk64irVdmbxLlTVwlsn0iXmlXZkrlzb/zY0i2f5JJWGClYoBcFchEBMQAnor+lMLJc AHu8dlX1lo+2zy4LCPDHQlhMxm2fHZnUTpT8iBECbzBDI70w/Q/Oz3PI6h7br9V/hxIdPYIvNMJ iovPnXQ2y5YAMgub8g8FQCKfh6I6AZjjH4y8liISgk0sgLoGyeZLsvU1ivAG0X8wQ2rTrwy+JUH iJMx0ePBCx6DqcLHU640i6btWtgBB6C3MHHyd8A/yDwnhZddihvz8VlVhHLJFNKOmaPQlaaptFS a8kVJLtl3w2EYJuVSfZ81VR80fnSeQO1IFCWz5oTmDBd0XUHs+1RiSKBBs5yFnIxk7tHwUUf/Aj NcSdLSrX0W9VLNb7THOf7oBrs7q6zGHxM+CC0cbJqf/Kh6Ys+MhWfkNqcL2lRjR0+uYQ3MLbZcR Ruh+x817rwmup+A== X-Developer-Key: i=jlayton@kernel.org; a=openpgp; fpr=4BC0D7B24471B2A184EAF5D3000E684119568215 nfsd_mutex serializes the entire NFSD control plane across every network namespace. The nfsd genl family sets .parallel_ops, so it is the only serialization there: one container starting nfsd, or one long RPC_STATUS_GET dump, stalls every other namespace's admin operations. The nfsd threads suffer too -- the dynamic-thread autoscaler trylocks the same mutex on every -ETIMEDOUT and -EBUSY, and a failed trylock skips the spawn/reap entirely. Almost none of what the mutex covers is actually shared. Add nn->nfsd_mutex for the per-namespace control plane: - nn->nfsd_serv and the svc_serv members hanging off it (->sv_permsocks, ->sv_temp_socks, per-pool thread counts) - NFSD_NET_UP / NFSD_NET_LOCKD_UP - the settables that may only change while the server is down (->nfsd_versions, ->nfsd4_lease, ->nfsd4_grace, ->max_blksize, ...) - nn->svc_export_cache / nn->svc_expkey_cache liveness - nn->conf_id_hashtbl liveness, for the state-revoke walks The global nfsd_mutex keeps only what is genuinely host-wide: the nfsd_users refcount and the resources it brings up (open file cache, NFSv4 global tables), and the address-notifier registration. nfsd_startup_generic()/nfsd_shutdown_generic() and the notifier register/unregister now take it internally, so per-net callers never see it. nfsd_file_cache_purge() likewise takes it itself, which lets expkey_flush() drop its hand-rolled lock. write_recoverydir() still takes the global mutex, but that now serializes writers only. The startup readers of user_recovery_dirname -- nfsd4_init_recdir() and check_for_legacy_methods() -- run under nn->nfsd_mutex alone, so a write from one namespace can tear the string under another namespace's startup. Worst case is a bogus path and a spurious startup error; the buffer is always NUL-terminated in bounds, so there is nothing to overrun. Left alone deliberately: that global already had lock-free readers in nfsd4_cltrack_legacy_{topdir,recdir}(), and legacy client tracking is deprecated and effectively init-netns-only -- its usermodehelper upcall always runs in the init mount namespace, so the stored path can only ever name an init-ns path. Lock ordering is nn->nfsd_mutex outside the global nfsd_mutex; nothing takes two namespaces' nfsd_mutexes. struct svc_info already indirects through a mutex pointer, so pool_stats needs only to be pointed at the new lock. The notifier refcount becomes a plain int now that it is genuinely mutex-guarded -- an atomic was never enough to serialize the register/unregister against the count. nfsd_create_serv() now takes that reference before publishing nn->nfsd_serv: namespaces are no longer serialized against each other here, so ordering it the other way would let the count dip to zero while another namespace's serv is already visible. Note: the RPC thread pool needs no global lock here. svc_pool_map has its own svc_pool_map_mutex and per-pool thread counts live in the per-net svc_serv. Assisted-by: LLM Signed-off-by: Jeff Layton --- fs/nfsd/export.c | 22 ++-- fs/nfsd/filecache.c | 5 +- fs/nfsd/netns.h | 10 ++ fs/nfsd/nfs4proc.c | 4 +- fs/nfsd/nfs4state.c | 8 +- fs/nfsd/nfsctl.c | 117 ++++++++++-------= -- fs/nfsd/nfssvc.c | 129 ++++++++++++++---= ---- .../testing/selftests/nfsd/nfsd_netlink_listener.c | 6 +- 8 files changed, 180 insertions(+), 121 deletions(-) diff --git a/fs/nfsd/export.c b/fs/nfsd/export.c index e5a0f1ababe6..265ea8fd31c6 100644 --- a/fs/nfsd/export.c +++ b/fs/nfsd/export.c @@ -248,13 +248,7 @@ static struct cache_head *expkey_alloc(void) =20 static void expkey_flush(void) { - /* - * Take the nfsd_mutex here to ensure that the file cache is not - * destroyed while we're in the middle of flushing. - */ - mutex_lock(&nfsd_mutex); nfsd_file_cache_purge(current->nsproxy->net_ns); - mutex_unlock(&nfsd_mutex); } =20 static int expkey_notify(struct cache_detail *cd, struct cache_head *h) @@ -346,7 +340,7 @@ int nfsd_nl_expkey_get_reqs_dumpit(struct sk_buff *skb, =20 nn =3D net_generic(sock_net(skb->sk), nfsd_net_id); =20 - mutex_lock(&nfsd_mutex); + mutex_lock(&nn->nfsd_mutex); =20 cd =3D nn->svc_expkey_cache; if (!cd) { @@ -425,7 +419,7 @@ int nfsd_nl_expkey_get_reqs_dumpit(struct sk_buff *skb, kfree(seqnos); kfree(items); out_unlock: - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); return ret; } =20 @@ -560,7 +554,7 @@ int nfsd_nl_expkey_set_reqs_doit(struct sk_buff *skb, =20 nn =3D net_generic(genl_info_net(info), nfsd_net_id); =20 - mutex_lock(&nfsd_mutex); + mutex_lock(&nn->nfsd_mutex); =20 cd =3D nn->svc_expkey_cache; if (!cd) { @@ -576,7 +570,7 @@ int nfsd_nl_expkey_set_reqs_doit(struct sk_buff *skb, } =20 out_unlock: - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); return ret; } =20 @@ -673,7 +667,7 @@ int nfsd_nl_svc_export_get_reqs_dumpit(struct sk_buff *= skb, =20 nn =3D net_generic(sock_net(skb->sk), nfsd_net_id); =20 - mutex_lock(&nfsd_mutex); + mutex_lock(&nn->nfsd_mutex); =20 cd =3D nn->svc_export_cache; if (!cd) { @@ -757,7 +751,7 @@ int nfsd_nl_svc_export_get_reqs_dumpit(struct sk_buff *= skb, kfree(seqnos); kfree(items); out_unlock: - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); return ret; } =20 @@ -1056,7 +1050,7 @@ int nfsd_nl_svc_export_set_reqs_doit(struct sk_buff *= skb, =20 nn =3D net_generic(genl_info_net(info), nfsd_net_id); =20 - mutex_lock(&nfsd_mutex); + mutex_lock(&nn->nfsd_mutex); =20 cd =3D nn->svc_export_cache; if (!cd) { @@ -1072,7 +1066,7 @@ int nfsd_nl_svc_export_set_reqs_doit(struct sk_buff *= skb, } =20 out_unlock: - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); return ret; } =20 diff --git a/fs/nfsd/filecache.c b/fs/nfsd/filecache.c index 17a94e6fcb15..79b9e8c92e70 100644 --- a/fs/nfsd/filecache.c +++ b/fs/nfsd/filecache.c @@ -1011,13 +1011,16 @@ nfsd_file_cache_start_net(struct net *net) * nfsd_file_cache_purge - Remove all cache items associated with @net * @net: target net namespace * + * Takes nfsd_mutex so the cache cannot be torn down underneath the + * walk. Callers must not already hold it. */ void nfsd_file_cache_purge(struct net *net) { - lockdep_assert_held(&nfsd_mutex); + mutex_lock(&nfsd_mutex); if (test_bit(NFSD_FILE_CACHE_UP, &nfsd_file_flags) =3D=3D 1) __nfsd_file_cache_purge(net); + mutex_unlock(&nfsd_mutex); } =20 void diff --git a/fs/nfsd/netns.h b/fs/nfsd/netns.h index 374ce83e2ba0..35199c17f8d1 100644 --- a/fs/nfsd/netns.h +++ b/fs/nfsd/netns.h @@ -164,6 +164,16 @@ struct nfsd_net { u32 clientid_counter; u32 clverifier_counter; =20 + /* + * Serializes this namespace's control plane: ->nfsd_serv and the + * svc_serv members that hang off it (->sv_permsocks, + * ->sv_temp_socks, thread counts), the NFSD_NET_* flags, and the + * settables above that may only change while the server is down. + * + * Nests outside the global nfsd_mutex. + */ + struct mutex nfsd_mutex; + struct svc_info nfsd_info; #define nfsd_serv nfsd_info.serv =20 diff --git a/fs/nfsd/nfs4proc.c b/fs/nfsd/nfs4proc.c index 3a82af381a8d..7df60abfbff1 100644 --- a/fs/nfsd/nfs4proc.c +++ b/fs/nfsd/nfs4proc.c @@ -1703,7 +1703,7 @@ static bool nfsd4_copy_on_sb(const struct nfsd4_copy = *copy, * @net: net namespace containing the copy operations * @sb: targeted superblock * - * Context: Caller must hold nfsd_mutex with NFSD_NET_UP set. Outside + * Context: Caller must hold nn->nfsd_mutex with NFSD_NET_UP set. Outside * that window nn->conf_id_hashtbl is unallocated or freed, * so the walk would dereference a NULL or dangling pointer. */ @@ -1715,7 +1715,7 @@ void nfsd4_cancel_copy_by_sb(struct net *net, struct = super_block *sb) unsigned int idhashval; LIST_HEAD(to_cancel); =20 - lockdep_assert_held(&nfsd_mutex); + lockdep_assert_held(&nn->nfsd_mutex); spin_lock(&nn->client_lock); for (idhashval =3D 0; idhashval < CLIENT_HASH_SIZE; idhashval++) { struct list_head *head =3D &nn->conf_id_hashtbl[idhashval]; diff --git a/fs/nfsd/nfs4state.c b/fs/nfsd/nfs4state.c index 1de6c6d757c3..0f9340eb281e 100644 --- a/fs/nfsd/nfs4state.c +++ b/fs/nfsd/nfs4state.c @@ -2102,7 +2102,7 @@ static void revoke_one_stid(struct nfsd_net *nn, stru= ct nfs4_client *clp, * The clients which own the states will subsequently be notified that the * states have been "admin-revoked". * - * Context: Caller must hold nfsd_mutex with NFSD_NET_UP set. Outside + * Context: Caller must hold nn->nfsd_mutex with NFSD_NET_UP set. Outside * that window nn->conf_id_hashtbl is unallocated or freed, * so the walk would dereference a NULL or dangling pointer. */ @@ -2111,7 +2111,7 @@ void nfsd4_revoke_states(struct nfsd_net *nn, struct = super_block *sb) unsigned int idhashval; unsigned int sc_types; =20 - lockdep_assert_held(&nfsd_mutex); + lockdep_assert_held(&nn->nfsd_mutex); =20 sc_types =3D SC_TYPE_OPEN | SC_TYPE_LOCK | SC_TYPE_DELEG | SC_TYPE_LAYOUT; =20 @@ -2190,7 +2190,7 @@ static struct nfs4_stid *find_one_export_stid(struct = nfs4_client *clp, * Userspace (exportfs -u) sends this after removing the last client * for a path, enabling the underlying filesystem to be unmounted. * - * Context: Caller must hold nfsd_mutex with NFSD_NET_UP set. Outside + * Context: Caller must hold nn->nfsd_mutex with NFSD_NET_UP set. Outside * that window nn->conf_id_hashtbl is unallocated or freed, * so the walk would dereference a NULL or dangling pointer. */ @@ -2199,7 +2199,7 @@ void nfsd4_revoke_export_states(struct nfsd_net *nn, = const struct path *path) unsigned int idhashval; unsigned int sc_types; =20 - lockdep_assert_held(&nfsd_mutex); + lockdep_assert_held(&nn->nfsd_mutex); =20 sc_types =3D SC_TYPE_OPEN | SC_TYPE_LOCK | SC_TYPE_DELEG | SC_TYPE_LAYOUT; =20 diff --git a/fs/nfsd/nfsctl.c b/fs/nfsd/nfsctl.c index 330d0f12e199..5ae33c21cf71 100644 --- a/fs/nfsd/nfsctl.c +++ b/fs/nfsd/nfsctl.c @@ -302,15 +302,15 @@ static ssize_t write_unlock_fs(struct file *file, cha= r *buf, size_t size) * 3. Is that directory the root of an exported file system? */ error =3D nlmsvc_unlock_all_by_sb(path.dentry->d_sb); - mutex_lock(&nfsd_mutex); nn =3D net_generic(netns(file), nfsd_net_id); + mutex_lock(&nn->nfsd_mutex); if (test_bit(NFSD_NET_UP, &nn->flags)) { nfsd4_cancel_copy_by_sb(netns(file), path.dentry->d_sb); nfsd4_revoke_states(nn, path.dentry->d_sb); } else { error =3D -EINVAL; } - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); =20 path_put(&path); return error; @@ -436,12 +436,12 @@ static ssize_t write_threads(struct file *file, char = *buf, size_t size) if (newthreads < 0) return -EINVAL; trace_nfsd_ctl_threads(net, newthreads); - mutex_lock(&nfsd_mutex); + mutex_lock(&nn->nfsd_mutex); if (newthreads > 0 || nn->nfsd_serv !=3D NULL) rv =3D nfsd_svc(1, &newthreads, net, file->f_cred, NULL); else rv =3D 0; - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); if (rv < 0) return rv; } else @@ -484,8 +484,9 @@ static ssize_t write_pool_threads(struct file *file, ch= ar *buf, size_t size) int npools; int *nthreads; struct net *net =3D netns(file); + struct nfsd_net *nn =3D net_generic(net, nfsd_net_id); =20 - mutex_lock(&nfsd_mutex); + mutex_lock(&nn->nfsd_mutex); npools =3D nfsd_nrpools(net); if (npools =3D=3D 0) { /* @@ -493,7 +494,7 @@ static ssize_t write_pool_threads(struct file *file, ch= ar *buf, size_t size) * writing to the threads file but NOT the pool_threads * file, sorry. Report zero threads. */ - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); strcpy(buf, "0\n"); return strlen(buf); } @@ -547,7 +548,7 @@ static ssize_t write_pool_threads(struct file *file, ch= ar *buf, size_t size) rv =3D mesg - buf; out_free: kfree(nthreads); - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); return rv; } =20 @@ -710,11 +711,12 @@ static ssize_t __write_versions(struct file *file, ch= ar *buf, size_t size) */ static ssize_t write_versions(struct file *file, char *buf, size_t size) { + struct nfsd_net *nn =3D net_generic(netns(file), nfsd_net_id); ssize_t rv; =20 - mutex_lock(&nfsd_mutex); + mutex_lock(&nn->nfsd_mutex); rv =3D __write_versions(file, buf, size); - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); return rv; } =20 @@ -876,11 +878,12 @@ static ssize_t __write_ports(struct file *file, char = *buf, size_t size, */ static ssize_t write_ports(struct file *file, char *buf, size_t size) { + struct nfsd_net *nn =3D net_generic(netns(file), nfsd_net_id); ssize_t rv; =20 - mutex_lock(&nfsd_mutex); + mutex_lock(&nn->nfsd_mutex); rv =3D __write_ports(file, buf, size, netns(file)); - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); return rv; } =20 @@ -924,13 +927,13 @@ static ssize_t write_maxblksize(struct file *file, ch= ar *buf, size_t size) bsize =3D max_t(int, bsize, 1024); bsize =3D min_t(int, bsize, NFSSVC_MAXBLKSIZE); bsize &=3D ~(1024-1); - mutex_lock(&nfsd_mutex); + mutex_lock(&nn->nfsd_mutex); if (nn->nfsd_serv) { - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); return -EBUSY; } nn->max_blksize =3D bsize; - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); } =20 return scnprintf(buf, SIMPLE_TRANSACTION_LIMIT, "%u\n", @@ -979,9 +982,9 @@ static ssize_t nfsd4_write_time(struct file *file, char= *buf, size_t size, { ssize_t rv; =20 - mutex_lock(&nfsd_mutex); + mutex_lock(&nn->nfsd_mutex); rv =3D __nfsd4_write_time(file, buf, size, time, nn); - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); return rv; } =20 @@ -1084,9 +1087,15 @@ static ssize_t write_recoverydir(struct file *file, = char *buf, size_t size) ssize_t rv; struct nfsd_net *nn =3D net_generic(netns(file), nfsd_net_id); =20 + /* + * nn->nfsd_mutex guards the nn->nfsd_serv check; the recovery + * dirname itself is still shared between namespaces. + */ + mutex_lock(&nn->nfsd_mutex); mutex_lock(&nfsd_mutex); rv =3D __write_recoverydir(file, buf, size, nn); mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); return rv; } #endif @@ -1530,9 +1539,9 @@ int nfsd_nl_rpc_status_get_dumpit(struct sk_buff *skb, int i, ret, rqstp_index =3D 0; struct nfsd_net *nn; =20 - mutex_lock(&nfsd_mutex); - nn =3D net_generic(sock_net(skb->sk), nfsd_net_id); + + mutex_lock(&nn->nfsd_mutex); if (!nn->nfsd_serv) { ret =3D -ENODEV; goto out_unlock; @@ -1649,7 +1658,7 @@ int nfsd_nl_rpc_status_get_dumpit(struct sk_buff *skb, out: rcu_read_unlock(); out_unlock: - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); =20 return ret; } @@ -1659,7 +1668,7 @@ int nfsd_nl_rpc_status_get_dumpit(struct sk_buff *skb, * @attr: nlattr NFSD_A_SERVER_FH_KEY * @nn: nfsd_net * - * Callers should hold nfsd_mutex, returns 0 on success or negative errno. + * Callers should hold nn->nfsd_mutex, returns 0 on success or negative er= rno. * Callers must ensure the server is shut down (sv_nrthreads =3D=3D 0), * userspace documentation asserts the key may only be set when the server * is not running. @@ -1714,7 +1723,7 @@ int nfsd_nl_threads_set_doit(struct sk_buff *skb, str= uct genl_info *info) GENL_HDRLEN, rem) nrpools++; =20 - mutex_lock(&nfsd_mutex); + mutex_lock(&nn->nfsd_mutex); =20 nthreads =3D kzalloc_objs(int, nrpools); if (!nthreads) { @@ -1779,7 +1788,7 @@ int nfsd_nl_threads_set_doit(struct sk_buff *skb, str= uct genl_info *info) if (ret > 0) ret =3D 0; out_unlock: - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); kfree(nthreads); return ret; } @@ -1808,7 +1817,7 @@ int nfsd_nl_threads_get_doit(struct sk_buff *skb, str= uct genl_info *info) goto err_free_msg; } =20 - mutex_lock(&nfsd_mutex); + mutex_lock(&nn->nfsd_mutex); =20 err =3D nla_put_u32(skb, NFSD_A_SERVER_GRACETIME, nn->nfsd4_grace) || @@ -1838,14 +1847,14 @@ int nfsd_nl_threads_get_doit(struct sk_buff *skb, s= truct genl_info *info) goto err_unlock; } =20 - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); =20 genlmsg_end(skb, hdr); =20 return genlmsg_reply(skb, info); =20 err_unlock: - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); err_free_msg: nlmsg_free(skb); =20 @@ -1868,11 +1877,11 @@ int nfsd_nl_version_set_doit(struct sk_buff *skb, s= truct genl_info *info) if (GENL_REQ_ATTR_CHECK(info, NFSD_A_SERVER_PROTO_VERSION)) return -EINVAL; =20 - mutex_lock(&nfsd_mutex); - nn =3D net_generic(genl_info_net(info), nfsd_net_id); + + mutex_lock(&nn->nfsd_mutex); if (nn->nfsd_serv) { - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); return -EBUSY; } =20 @@ -1915,7 +1924,7 @@ int nfsd_nl_version_set_doit(struct sk_buff *skb, str= uct genl_info *info) } } =20 - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); =20 return 0; } @@ -1943,9 +1952,9 @@ int nfsd_nl_version_get_doit(struct sk_buff *skb, str= uct genl_info *info) goto err_free_msg; } =20 - mutex_lock(&nfsd_mutex); nn =3D net_generic(genl_info_net(info), nfsd_net_id); =20 + mutex_lock(&nn->nfsd_mutex); for (i =3D 2; i <=3D 4; i++) { int j; =20 @@ -1987,13 +1996,13 @@ int nfsd_nl_version_get_doit(struct sk_buff *skb, s= truct genl_info *info) } } =20 - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); genlmsg_end(skb, hdr); =20 return genlmsg_reply(skb, info); =20 err_nfsd_unlock: - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); err_free_msg: nlmsg_free(skb); =20 @@ -2004,7 +2013,7 @@ int nfsd_nl_version_get_doit(struct sk_buff *skb, str= uct genl_info *info) * Transport classes NFSD knows how to instantiate. Vetting the name here * keeps a bogus string from reaching svc_xprt_create_from_sa(), where an * unknown name triggers a request_module("svc%s", name) upcall under - * nfsd_mutex. + * nn->nfsd_mutex. */ static bool nfsd_nl_transport_supported(const char *name) { @@ -2085,14 +2094,15 @@ static int nfsd_nl_validate_listeners(struct genl_i= nfo *info) return count; } =20 -static size_t nfsd_nl_listener_set_msgsize(struct svc_serv *serv) +static size_t nfsd_nl_listener_set_msgsize(struct nfsd_net *nn, + struct svc_serv *serv) { size_t size =3D GENL_HDRLEN + /* genlmsg_iput() */ nla_total_size(0); /* userspace-rpcbind */ struct svc_xprt *xprt; unsigned int p; =20 - lockdep_assert_held(&nfsd_mutex); + lockdep_assert_held(&nn->nfsd_mutex); =20 for (p =3D 0; p < serv->sv_nprogs; p++) size +=3D serv->sv_programs[p].pg_nvers * @@ -2118,15 +2128,16 @@ static struct sk_buff * nfsd_nl_listener_set_msg(struct genl_info *info, struct net *net, struct svc_serv *serv) { + struct nfsd_net *nn =3D net_generic(net, nfsd_net_id); struct svc_xprt *xprt; struct sk_buff *skb; unsigned int p, i; void *hdr; int err; =20 - lockdep_assert_held(&nfsd_mutex); + lockdep_assert_held(&nn->nfsd_mutex); =20 - skb =3D genlmsg_new(nfsd_nl_listener_set_msgsize(serv), GFP_KERNEL); + skb =3D genlmsg_new(nfsd_nl_listener_set_msgsize(nn, serv), GFP_KERNEL); if (!skb) return ERR_PTR(-ENOMEM); =20 @@ -2244,10 +2255,10 @@ int nfsd_nl_listener_set_doit(struct sk_buff *skb, = struct genl_info *info) =20 userspace_rpcbind =3D nla_get_flag(info->attrs[NFSD_A_SERVER_SOCK_USERSPA= CE_RPCBIND]); =20 - mutex_lock(&nfsd_mutex); - nn =3D net_generic(net, nfsd_net_id); =20 + mutex_lock(&nn->nfsd_mutex); + /* * An empty list destroys the serv, and nfsd_destroy_serv() drops * whatever svc_bind() took either way, so teardown is not an @@ -2258,13 +2269,13 @@ int nfsd_nl_listener_set_doit(struct sk_buff *skb, = struct genl_info *info) nn->nfsd_serv->sv_no_rpcbind !=3D userspace_rpcbind) { NL_SET_ERR_MSG(info->extack, "cannot change rpcbind ownership while a server exists"); - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); return -EBUSY; } =20 err =3D nfsd_create_serv(net, userspace_rpcbind); if (err) { - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); return err; } =20 @@ -2433,7 +2444,7 @@ int nfsd_nl_listener_set_doit(struct sk_buff *skb, st= ruct genl_info *info) nfsd_destroy_serv(net); =20 out_unlock_mtx: - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); =20 /* rskb is only built once err is known to be zero. */ if (rskb) @@ -2467,9 +2478,9 @@ int nfsd_nl_listener_get_doit(struct sk_buff *skb, st= ruct genl_info *info) goto err_free_msg; } =20 - mutex_lock(&nfsd_mutex); nn =3D net_generic(genl_info_net(info), nfsd_net_id); =20 + mutex_lock(&nn->nfsd_mutex); /* no nfs server? Just send empty socket list */ if (!nn->nfsd_serv) goto out_unlock_mtx; @@ -2498,14 +2509,14 @@ int nfsd_nl_listener_get_doit(struct sk_buff *skb, = struct genl_info *info) } spin_unlock_bh(&serv->sv_lock); out_unlock_mtx: - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); genlmsg_end(skb, hdr); =20 return genlmsg_reply(skb, info); =20 err_serv_unlock: spin_unlock_bh(&serv->sv_lock); - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); err_free_msg: nlmsg_free(skb); =20 @@ -2589,7 +2600,7 @@ int nfsd_nl_cache_flush_doit(struct sk_buff *skb, str= uct genl_info *info) if (info->attrs[NFSD_A_CACHE_FLUSH_MASK]) mask =3D nla_get_u32(info->attrs[NFSD_A_CACHE_FLUSH_MASK]); =20 - mutex_lock(&nfsd_mutex); + mutex_lock(&nn->nfsd_mutex); =20 if ((mask & NFSD_CACHE_TYPE_SVC_EXPORT) && nn->svc_export_cache) @@ -2599,7 +2610,7 @@ int nfsd_nl_cache_flush_doit(struct sk_buff *skb, str= uct genl_info *info) nn->svc_expkey_cache) cache_purge(nn->svc_expkey_cache); =20 - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); =20 return 0; } @@ -2947,14 +2958,14 @@ int nfsd_nl_unlock_filesystem_doit(struct sk_buff *= skb, =20 error =3D nlmsvc_unlock_all_by_sb(path.dentry->d_sb); =20 - mutex_lock(&nfsd_mutex); + mutex_lock(&nn->nfsd_mutex); if (test_bit(NFSD_NET_UP, &nn->flags)) { nfsd4_cancel_copy_by_sb(net, path.dentry->d_sb); nfsd4_revoke_states(nn, path.dentry->d_sb); } else { error =3D -EINVAL; } - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); =20 path_put(&path); return error; @@ -2994,13 +3005,13 @@ int nfsd_nl_unlock_export_doit(struct sk_buff *skb,= struct genl_info *info) if (error) return error; =20 - mutex_lock(&nfsd_mutex); + mutex_lock(&nn->nfsd_mutex); if (test_bit(NFSD_NET_UP, &nn->flags)) { nfsd_file_close_export(net, &path); nfsd4_revoke_export_states(nn, &path); } else error =3D -EINVAL; - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); =20 path_put(&path); return error; @@ -3057,7 +3068,8 @@ static __net_init int nfsd_net_init(struct net *net) nn->nfsd_versions[i] =3D nfsd_support_version(i); for (i =3D 0; i < sizeof(nn->nfsd4_minorversions); i++) nn->nfsd4_minorversions[i] =3D nfsd_support_version(4); - nn->nfsd_info.mutex =3D &nfsd_mutex; + mutex_init(&nn->nfsd_mutex); + nn->nfsd_info.mutex =3D &nn->nfsd_mutex; nn->nfsd_serv =3D NULL; nfsd4_init_leases_net(nn); get_random_bytes(&nn->siphash_key, sizeof(nn->siphash_key)); @@ -3121,6 +3133,7 @@ static __net_exit void nfsd_net_exit(struct net *net) percpu_counter_destroy_many(nn->counter, NFSD_STATS_COUNTERS_NUM); nfsd_idmap_shutdown(net); nfsd_export_shutdown(net); + mutex_destroy(&nn->nfsd_mutex); } =20 static struct pernet_operations nfsd_net_ops =3D { diff --git a/fs/nfsd/nfssvc.c b/fs/nfsd/nfssvc.c index 77e1e6ba686d..3d47e5c86bd5 100644 --- a/fs/nfsd/nfssvc.c +++ b/fs/nfsd/nfssvc.c @@ -55,16 +55,22 @@ static __be32 nfsd_init_request(struct svc_rqst *, struct svc_process_info *); =20 /* - * nfsd_mutex protects nn->nfsd_serv -- both the pointer itself and some m= embers - * of the svc_serv struct such as ->sv_temp_socks and ->sv_permsocks. + * NFSD's control plane is serialized by two mutexes. * - * Finally, the nfsd_mutex also protects some of the global variables that= are - * accessed when nfsd starts and that are settable via the write_* routine= s in - * nfsctl.c. In particular: + * Nearly everything is per-namespace and belongs to nn->nfsd_mutex: the + * nn->nfsd_serv pointer and the svc_serv members that hang off it + * (->sv_permsocks, ->sv_temp_socks, per-pool thread counts), the + * NFSD_NET_* flags, and the nfsd_net settables that may only change while + * that namespace's server is down (->nfsd_versions, ->nfsd4_lease, + * ->nfsd4_grace, ->max_blksize, ...). * - * user_recovery_dirname - * user_lease_time - * nfsd_versions + * The global nfsd_mutex covers only what is genuinely shared between + * namespaces: the nfsd_users refcount and the host-wide resources it + * brings up and tears down (the open file cache and the NFSv4 global + * tables), the address-notifier registration, and user_recovery_dirname. + * + * Lock ordering is nn->nfsd_mutex outside the global nfsd_mutex. Nothing + * takes two namespaces' nfsd_mutexes. */ DEFINE_MUTEX(nfsd_mutex); =20 @@ -251,20 +257,23 @@ int nfsd_nrthreads(struct net *net) int rv =3D 0; struct nfsd_net *nn =3D net_generic(net, nfsd_net_id); =20 - /* nfsd_mutex keeps nn->nfsd_serv valid across the read. */ - mutex_lock(&nfsd_mutex); + /* nn->nfsd_mutex keeps nn->nfsd_serv valid across the read. */ + mutex_lock(&nn->nfsd_mutex); if (nn->nfsd_serv) rv =3D svc_serv_maxthreads(nn->nfsd_serv); - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); return rv; } =20 +/* Number of namespaces holding the host-wide resources up */ static int nfsd_users =3D 0; =20 -static int nfsd_startup_generic(void) +static int __nfsd_startup_generic(void) { int ret; =20 + lockdep_assert_held(&nfsd_mutex); + if (nfsd_users++) return 0; =20 @@ -284,13 +293,24 @@ static int nfsd_startup_generic(void) return ret; } =20 -static void nfsd_shutdown_generic(void) +static int nfsd_startup_generic(void) { - if (--nfsd_users) - return; + int ret; =20 - nfs4_state_shutdown(); - nfsd_file_cache_shutdown(); + mutex_lock(&nfsd_mutex); + ret =3D __nfsd_startup_generic(); + mutex_unlock(&nfsd_mutex); + return ret; +} + +static void nfsd_shutdown_generic(void) +{ + mutex_lock(&nfsd_mutex); + if (!--nfsd_users) { + nfs4_state_shutdown(); + nfsd_file_cache_shutdown(); + } + mutex_unlock(&nfsd_mutex); } =20 static bool nfsd_needs_lockd(struct nfsd_net *nn) @@ -505,8 +525,32 @@ static struct notifier_block nfsd_inet6addr_notifier = =3D { }; #endif =20 -/* Only used under nfsd_mutex, so this atomic may be overkill: */ -static atomic_t nfsd_notifier_refcount =3D ATOMIC_INIT(0); +/* Number of namespaces with a serv, guarded by nfsd_mutex */ +static int nfsd_notifier_users; + +static void nfsd_register_notifiers(void) +{ + mutex_lock(&nfsd_mutex); + if (!nfsd_notifier_users++) { + register_inetaddr_notifier(&nfsd_inetaddr_notifier); +#if IS_ENABLED(CONFIG_IPV6) + register_inet6addr_notifier(&nfsd_inet6addr_notifier); +#endif + } + mutex_unlock(&nfsd_mutex); +} + +static void nfsd_unregister_notifiers(void) +{ + mutex_lock(&nfsd_mutex); + if (!--nfsd_notifier_users) { + unregister_inetaddr_notifier(&nfsd_inetaddr_notifier); +#if IS_ENABLED(CONFIG_IPV6) + unregister_inet6addr_notifier(&nfsd_inet6addr_notifier); +#endif + } + mutex_unlock(&nfsd_mutex); +} =20 /** * nfsd_destroy_serv - tear down NFSD's svc_serv for a namespace @@ -517,19 +561,13 @@ void nfsd_destroy_serv(struct net *net) struct nfsd_net *nn =3D net_generic(net, nfsd_net_id); struct svc_serv *serv =3D nn->nfsd_serv; =20 - lockdep_assert_held(&nfsd_mutex); + lockdep_assert_held(&nn->nfsd_mutex); =20 spin_lock(&nfsd_notifier_lock); nn->nfsd_serv =3D NULL; spin_unlock(&nfsd_notifier_lock); =20 - /* check if the notifier still has clients */ - if (atomic_dec_return(&nfsd_notifier_refcount) =3D=3D 0) { - unregister_inetaddr_notifier(&nfsd_inetaddr_notifier); -#if IS_ENABLED(CONFIG_IPV6) - unregister_inet6addr_notifier(&nfsd_inet6addr_notifier); -#endif - } + nfsd_unregister_notifiers(); =20 /* * write_ports can create the server without actually starting @@ -586,17 +624,17 @@ void nfsd_shutdown_threads(struct net *net) struct nfsd_net *nn =3D net_generic(net, nfsd_net_id); struct svc_serv *serv; =20 - mutex_lock(&nfsd_mutex); + mutex_lock(&nn->nfsd_mutex); serv =3D nn->nfsd_serv; if (serv =3D=3D NULL) { - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); return; } =20 /* Kill outstanding nfsd threads */ svc_set_num_threads(serv, 0, 0); nfsd_destroy_serv(net); - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); } =20 struct svc_rqst *nfsd_current_rqst(void) @@ -619,7 +657,7 @@ int nfsd_create_serv(struct net *net, bool no_rpcbind) struct nfsd_net *nn =3D net_generic(net, nfsd_net_id); struct svc_serv *serv; =20 - WARN_ON(!mutex_is_locked(&nfsd_mutex)); + WARN_ON(!mutex_is_locked(&nn->nfsd_mutex)); if (nn->nfsd_serv) return 0; =20 @@ -650,17 +688,18 @@ int nfsd_create_serv(struct net *net, bool no_rpcbind) percpu_ref_exit(&nn->nfsd_net_ref); return error; } + /* + * Register before publishing nn->nfsd_serv. Namespaces are only + * serialized against each other by nfsd_mutex here, so + * taking the reference first is what guarantees a visible + * nn->nfsd_serv never coincides with an unregistered notifier. + */ + nfsd_register_notifiers(); + spin_lock(&nfsd_notifier_lock); nn->nfsd_serv =3D serv; spin_unlock(&nfsd_notifier_lock); =20 - /* check if the notifier is already set */ - if (atomic_inc_return(&nfsd_notifier_refcount) =3D=3D 1) { - register_inetaddr_notifier(&nfsd_inetaddr_notifier); -#if IS_ENABLED(CONFIG_IPV6) - register_inet6addr_notifier(&nfsd_inet6addr_notifier); -#endif - } nfsd_reset_write_verifier(nn); return 0; } @@ -707,7 +746,7 @@ int nfsd_set_nrthreads(int n, int *nthreads, struct net= *net) int err =3D 0; struct nfsd_net *nn =3D net_generic(net, nfsd_net_id); =20 - lockdep_assert_held(&nfsd_mutex); + lockdep_assert_held(&nn->nfsd_mutex); =20 if (nn->nfsd_serv =3D=3D NULL || n <=3D 0) return 0; @@ -777,7 +816,7 @@ nfsd_svc(int n, int *nthreads, struct net *net, const s= truct cred *cred, const c struct nfsd_net *nn =3D net_generic(net, nfsd_net_id); struct svc_serv *serv; =20 - lockdep_assert_held(&nfsd_mutex); + lockdep_assert_held(&nn->nfsd_mutex); =20 dprintk("nfsd: creating service\n"); =20 @@ -955,13 +994,13 @@ nfsd(void *vrqstp) switch (svc_recv(rqstp, 5 * HZ)) { case -ETIMEDOUT: /* No work arrived within the timeout window */ - if (mutex_trylock(&nfsd_mutex)) { + if (mutex_trylock(&nn->nfsd_mutex)) { if (pool->sp_nrthreads > pool->sp_nrthrmin) { trace_nfsd_dynthread_kill(net, pool); set_bit(RQ_VICTIM, &rqstp->rq_flags); have_mutex =3D true; } else { - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); } } else { trace_nfsd_dynthread_trylock_fail(net, pool); @@ -970,7 +1009,7 @@ nfsd(void *vrqstp) case -EBUSY: /* No idle threads; consider spawning another */ if (pool->sp_nrthreads < pool->sp_nrthrmax) { - if (mutex_trylock(&nfsd_mutex)) { + if (mutex_trylock(&nn->nfsd_mutex)) { if (pool->sp_nrthreads < pool->sp_nrthrmax) { int ret; =20 @@ -980,7 +1019,7 @@ nfsd(void *vrqstp) pr_notice_ratelimited("%s: unable to spawn new thread: %d\n", __func__, ret); } - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); } else { trace_nfsd_dynthread_trylock_fail(net, pool); } @@ -998,7 +1037,7 @@ nfsd(void *vrqstp) /* Release the thread */ svc_exit_thread(rqstp); if (have_mutex) - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nn->nfsd_mutex); return 0; } =20 diff --git a/tools/testing/selftests/nfsd/nfsd_netlink_listener.c b/tools/t= esting/selftests/nfsd/nfsd_netlink_listener.c index bef7e8b1ee71..1294057b6f62 100644 --- a/tools/testing/selftests/nfsd/nfsd_netlink_listener.c +++ b/tools/testing/selftests/nfsd/nfsd_netlink_listener.c @@ -5,7 +5,7 @@ * * Three groups: * validation - malformed/abusive LISTENER_SET requests are rejected by - * nfsd_nl_validate_listeners(), before nfsd_mutex is take= n. + * nfsd_nl_validate_listeners(), before nn->nfsd_mutex is = taken. * functional - create/add/remove listeners and verify LISTENER_GET * reflects the set (round-trip of transport + addr:port). * semantics - once threads are running (THREADS_SET) a listener change @@ -949,8 +949,8 @@ TEST_F(nfsd_listener, val_missing_transport) } =20 /* - * A name matching no transport class must be refused before nfsd_mutex is - * taken, so it never reaches svc_xprt_create_from_sa() and its + * A name matching no transport class must be refused before nn->nfsd_mutex + * is taken, so it never reaches svc_xprt_create_from_sa() and its * request_module("svc%s", name) upcall. * * The errno cannot show that -- svc_xprt_create_from_sa() returns --=20 2.55.0 From nobody Thu Sep 24 16:07:44 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 34E1851DAF1; Tue, 22 Sep 2026 11:34:23 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790076867; cv=none; b=Jxg9tugEnl/huRtBRPdwZYhYTURwhBtOBp0jVC75Ab844buPYG/NhS6ZiMv61OJHv4s8PNjTXHiKSMaEYXhlJOTaGoAO/VS5Umi87pWYIlMQaVhyNv2/n16GFvP8jZv3FJuH0n4zO5d43Ezpbom3oVRenmAVOqWGS/WNj4WI2FI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790076867; c=relaxed/simple; bh=eccpfKgq7LHNRgtiVZPkAmXh6awU7amVgjZZIwNqGXs=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=RXyrEi7c9041S/5yWAqxMO+QIhhjS2W3XFx2/eCL58AsJu56KqjDe5HgqfDINB39mwAtCfO1Ve25PkhgGvX+TipqNKLw9HAX+gmL/820OwvqTcZbJe+NylrN5BXhqtzdBBjNj+h3fOyEikb0uimjr9ntTtSaNnL5zcJHkgf2zcY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=NFwRIb3p; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="NFwRIb3p" Received: by smtp.kernel.org (Postfix) with ESMTPSA id A7B7C1F000FF; Tue, 22 Sep 2026 11:34:22 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1790076863; bh=gYr4Al9SkWgWFj4tavPhadksvjnQShdDD9G12OjD91E=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=NFwRIb3pt6de0hgh1P8ccCfiI12RWrbnDWTuyUH/5OfbKYRzTpZKZiTDy0KRWUmcs oORX+Y6ZDVF2UvM2q8W+1KEhorvJSOttX7i3KvAACfXVIDqK4blcfesC8G4jvSn2yG OyZl2t4yKddkFkWvMAoKwef2OWSDRswj8IktkFIRMNA7RJctodpt8dDYBSSyTD81/b 04hzi+9kutkEg+TdOqfXJXvOoRfBsD/5h30tNdZIj/Tn8j1I1boRqB5qy4urb7YTK2 SlwMrnQorZXSEoh+7jDcJ6MUQAPM56bPEHwMHY5Mf4yvf4kLC7jk1KuO1dogDfxNbZ 4MfVmhbEiBY3w== From: Jeff Layton Date: Tue, 22 Sep 2026 07:34:05 -0400 Subject: [PATCH v2 4/8] nfsd: rename nfsd_mutex to nfsd_global_mutex Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260922-nfsd-per-net-mutex-v2-4-4a5da8243a73@kernel.org> References: <20260922-nfsd-per-net-mutex-v2-0-4a5da8243a73@kernel.org> In-Reply-To: <20260922-nfsd-per-net-mutex-v2-0-4a5da8243a73@kernel.org> To: Chuck Lever , NeilBrown , Olga Kornievskaia , Dai Ngo , Tom Talpey , Shuah Khan Cc: linux-nfs@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, Jeff Layton X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=openpgp-sha256; l=7432; i=jlayton@kernel.org; h=from:subject:message-id; bh=eccpfKgq7LHNRgtiVZPkAmXh6awU7amVgjZZIwNqGXs=; b=owEBbQKS/ZANAwAKAQAOaEEZVoIVAcsmYgBqsme51UfH1OjZV+Ww+CTpZ7KIKUsqcr4ohi6bI JT/DYM36niJAjMEAAEKAB0WIQRLwNeyRHGyoYTq9dMADmhBGVaCFQUCarJnuQAKCRAADmhBGVaC FfvmEACTENxEw1oFHy0u7PfzCmwP9yvuwTHi+etxs0HOVwqpTWyYRkTqUd1mgbdPHe3sdpUcNy7 kfZZIrgkiFxk1TiIfpbNechvVT+MhnoA5+r9fIkYLouR+oVapdWy7bTm7v3tW9nseihNKRhkuqS LUDDKrMz73tLTmZa0nZYMGi+zA1inJy3qRPvvGTdCrrH5FQ2qTAiLZB72gyr3ttrPuywTJetQd6 LQpX5ef9KZIZJDcSkKnBj+xxvNhKVDTHLoqUM8RjJoQYIfMtdUyZiLwF6P6+32mTenVfBshCaBQ hI9B/DJ0xWM2tMe3oZv+1AhVrCgokYWRey0wCHhwh3r+qHFaZftdb/In4oWHf0mXedBswnxJDNp YV21CSZjW6LyqNi3da2vFbloXRodzg9Yg+FwXEztFd9IbepBu+86EjBYM+qshoun3YK7cVKl4x5 14Hrof6kopv9gkbBPliKV2YXsM3KJ4j2xzEbxCwIvuKpgg/bkUoGOP56/dzashlpaiZVtAzYbXO 3fi99ho2g7kpXLfC5WrHM6WqPDJDKTNjlDHrSwErnxVNkfKrJwWff7DtyStS+tphE2EhsjcdiKx 7Mo6Fg9IYBAj3o6qVQhDYvkwtwuavKMnA5DZrTF7wUxicIvPb/6Jk+rzUCOo0QoPF9Z2XH5/LOO aQrlJj7rXrPOwrQ== X-Developer-Key: i=jlayton@kernel.org; a=openpgp; fpr=4BC0D7B24471B2A184EAF5D3000E684119568215 With nn->nfsd_mutex now carrying the per-namespace control plane, give the remaining module-wide mutex a name that says which one a call site means. Mechanical; no functional change. Assisted-by: LLM Signed-off-by: Jeff Layton --- fs/nfsd/filecache.c | 14 +++++++------- fs/nfsd/netns.h | 2 +- fs/nfsd/nfsctl.c | 4 ++-- fs/nfsd/nfsd.h | 2 +- fs/nfsd/nfssvc.c | 34 +++++++++++++++++----------------- 5 files changed, 28 insertions(+), 28 deletions(-) diff --git a/fs/nfsd/filecache.c b/fs/nfsd/filecache.c index 79b9e8c92e70..501772bc8f8d 100644 --- a/fs/nfsd/filecache.c +++ b/fs/nfsd/filecache.c @@ -869,7 +869,7 @@ nfsd_file_cache_init(void) { int ret; =20 - lockdep_assert_held(&nfsd_mutex); + lockdep_assert_held(&nfsd_global_mutex); if (test_and_set_bit(NFSD_FILE_CACHE_UP, &nfsd_file_flags) =3D=3D 1) return 0; =20 @@ -1011,16 +1011,16 @@ nfsd_file_cache_start_net(struct net *net) * nfsd_file_cache_purge - Remove all cache items associated with @net * @net: target net namespace * - * Takes nfsd_mutex so the cache cannot be torn down underneath the + * Takes nfsd_global_mutex so the cache cannot be torn down underneath the * walk. Callers must not already hold it. */ void nfsd_file_cache_purge(struct net *net) { - mutex_lock(&nfsd_mutex); + mutex_lock(&nfsd_global_mutex); if (test_bit(NFSD_FILE_CACHE_UP, &nfsd_file_flags) =3D=3D 1) __nfsd_file_cache_purge(net); - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nfsd_global_mutex); } =20 void @@ -1045,7 +1045,7 @@ nfsd_file_cache_shutdown(void) { int i; =20 - lockdep_assert_held(&nfsd_mutex); + lockdep_assert_held(&nfsd_global_mutex); if (test_and_clear_bit(NFSD_FILE_CACHE_UP, &nfsd_file_flags) =3D=3D 0) return; =20 @@ -1477,7 +1477,7 @@ int nfsd_file_cache_stats_show(struct seq_file *m, vo= id *v) unsigned long lru =3D 0, total_age =3D 0; =20 /* Serialize with server shutdown */ - mutex_lock(&nfsd_mutex); + mutex_lock(&nfsd_global_mutex); if (test_bit(NFSD_FILE_CACHE_UP, &nfsd_file_flags) =3D=3D 1) { struct bucket_table *tbl; struct rhashtable *ht; @@ -1491,7 +1491,7 @@ int nfsd_file_cache_stats_show(struct seq_file *m, vo= id *v) buckets =3D tbl->size; rcu_read_unlock(); } - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nfsd_global_mutex); =20 for_each_possible_cpu(i) { hits +=3D per_cpu(nfsd_file_cache_hits, i); diff --git a/fs/nfsd/netns.h b/fs/nfsd/netns.h index 35199c17f8d1..866d5f64641f 100644 --- a/fs/nfsd/netns.h +++ b/fs/nfsd/netns.h @@ -170,7 +170,7 @@ struct nfsd_net { * ->sv_temp_socks, thread counts), the NFSD_NET_* flags, and the * settables above that may only change while the server is down. * - * Nests outside the global nfsd_mutex. + * Nests outside nfsd_global_mutex. */ struct mutex nfsd_mutex; =20 diff --git a/fs/nfsd/nfsctl.c b/fs/nfsd/nfsctl.c index 5ae33c21cf71..e2455a549e25 100644 --- a/fs/nfsd/nfsctl.c +++ b/fs/nfsd/nfsctl.c @@ -1092,9 +1092,9 @@ static ssize_t write_recoverydir(struct file *file, c= har *buf, size_t size) * dirname itself is still shared between namespaces. */ mutex_lock(&nn->nfsd_mutex); - mutex_lock(&nfsd_mutex); + mutex_lock(&nfsd_global_mutex); rv =3D __write_recoverydir(file, buf, size, nn); - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nfsd_global_mutex); mutex_unlock(&nn->nfsd_mutex); return rv; } diff --git a/fs/nfsd/nfsd.h b/fs/nfsd/nfsd.h index 0864d6d564e2..d824d2c39029 100644 --- a/fs/nfsd/nfsd.h +++ b/fs/nfsd/nfsd.h @@ -46,7 +46,7 @@ enum { =20 extern struct svc_program nfsd_programs[]; extern const struct svc_version nfsd_version2, nfsd_version3, nfsd_version= 4; -extern struct mutex nfsd_mutex; +extern struct mutex nfsd_global_mutex; extern atomic_t nfsd_th_cnt; /* number of available threads */ =20 extern const struct seq_operations nfs_exports_op; diff --git a/fs/nfsd/nfssvc.c b/fs/nfsd/nfssvc.c index 3d47e5c86bd5..5cb92c3829f9 100644 --- a/fs/nfsd/nfssvc.c +++ b/fs/nfsd/nfssvc.c @@ -64,15 +64,15 @@ static __be32 nfsd_init_request(struct svc_rqst *, * that namespace's server is down (->nfsd_versions, ->nfsd4_lease, * ->nfsd4_grace, ->max_blksize, ...). * - * The global nfsd_mutex covers only what is genuinely shared between + * nfsd_global_mutex covers only what is genuinely shared between * namespaces: the nfsd_users refcount and the host-wide resources it * brings up and tears down (the open file cache and the NFSv4 global * tables), the address-notifier registration, and user_recovery_dirname. * - * Lock ordering is nn->nfsd_mutex outside the global nfsd_mutex. Nothing - * takes two namespaces' nfsd_mutexes. + * Lock ordering is nn->nfsd_mutex outside nfsd_global_mutex. Nothing tak= es + * two namespaces' nfsd_mutexes. */ -DEFINE_MUTEX(nfsd_mutex); +DEFINE_MUTEX(nfsd_global_mutex); =20 #if IS_ENABLED(CONFIG_NFS_LOCALIO) static const struct svc_version *localio_versions[] =3D { @@ -272,7 +272,7 @@ static int __nfsd_startup_generic(void) { int ret; =20 - lockdep_assert_held(&nfsd_mutex); + lockdep_assert_held(&nfsd_global_mutex); =20 if (nfsd_users++) return 0; @@ -297,20 +297,20 @@ static int nfsd_startup_generic(void) { int ret; =20 - mutex_lock(&nfsd_mutex); + mutex_lock(&nfsd_global_mutex); ret =3D __nfsd_startup_generic(); - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nfsd_global_mutex); return ret; } =20 static void nfsd_shutdown_generic(void) { - mutex_lock(&nfsd_mutex); + mutex_lock(&nfsd_global_mutex); if (!--nfsd_users) { nfs4_state_shutdown(); nfsd_file_cache_shutdown(); } - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nfsd_global_mutex); } =20 static bool nfsd_needs_lockd(struct nfsd_net *nn) @@ -525,31 +525,31 @@ static struct notifier_block nfsd_inet6addr_notifier = =3D { }; #endif =20 -/* Number of namespaces with a serv, guarded by nfsd_mutex */ +/* Number of namespaces with a serv, guarded by nfsd_global_mutex */ static int nfsd_notifier_users; =20 static void nfsd_register_notifiers(void) { - mutex_lock(&nfsd_mutex); + mutex_lock(&nfsd_global_mutex); if (!nfsd_notifier_users++) { register_inetaddr_notifier(&nfsd_inetaddr_notifier); #if IS_ENABLED(CONFIG_IPV6) register_inet6addr_notifier(&nfsd_inet6addr_notifier); #endif } - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nfsd_global_mutex); } =20 static void nfsd_unregister_notifiers(void) { - mutex_lock(&nfsd_mutex); + mutex_lock(&nfsd_global_mutex); if (!--nfsd_notifier_users) { unregister_inetaddr_notifier(&nfsd_inetaddr_notifier); #if IS_ENABLED(CONFIG_IPV6) unregister_inet6addr_notifier(&nfsd_inet6addr_notifier); #endif } - mutex_unlock(&nfsd_mutex); + mutex_unlock(&nfsd_global_mutex); } =20 /** @@ -690,9 +690,9 @@ int nfsd_create_serv(struct net *net, bool no_rpcbind) } /* * Register before publishing nn->nfsd_serv. Namespaces are only - * serialized against each other by nfsd_mutex here, so - * taking the reference first is what guarantees a visible - * nn->nfsd_serv never coincides with an unregistered notifier. + * serialized against each other by nfsd_global_mutex here, so taking + * the reference first is what guarantees a visible nn->nfsd_serv + * never coincides with an unregistered notifier. */ nfsd_register_notifiers(); =20 --=20 2.55.0 From nobody Thu Sep 24 16:07:44 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A2A25538D91; Tue, 22 Sep 2026 11:34:24 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790076866; cv=none; b=n9Dzu/T654eiekOKKb0gBHztQw2NLwn0kMyvNetKRVJyIii9NbhgYo/pla2OtshZivOwOZZzoXNU28G7SfwLN+rCc4lmTLQDEv66LT8HLMobxbwcOuA8R99J0V5bx/vgLBrRzP8vRA7Ufrvk0YFSGxsH+6AVxwrOryCXi6ZN2J0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790076866; c=relaxed/simple; bh=9yOyKfLbgAQCiwsB8y/tXz9WKfPm1n0GduTeDPF6LG0=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=LJZlV6loADL6TShly9oWpg+vvZXZxodXpzLS4Qe+bRHUfcamqk869OlbHoggNkDAx/dJ++4zEX08sULz9qO98M/BHvDSLCjZPrcMb9/1NePAyNL3lPr6p4M1SvJ5T8OLTiTVA8I/bQl225CL5NMhfIP2gwrrQlF0RBqT4s83GVQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=TSkWoyIo; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="TSkWoyIo" Received: by smtp.kernel.org (Postfix) with ESMTPSA id A47491F00893; Tue, 22 Sep 2026 11:34:23 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1790076864; bh=mBQU3uBQGhPYntNb5Yath9p2lTfYe1lGY1j/HkI3Jx4=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=TSkWoyIoeaOFX4inbfUq+MSUMJMNoSxanTrG2qvnlhjwVu254+C9Ub0My0gjMyodg cBHEb/MZhjumgEKSpbIz/gocwjGtEpjbfYfhXBM6m6z4bSJUZg1aqWAEp5WsBXn3u9 cAGaVomkenskgsb8Bk9LpJI9TrdCnVMDSnIt3uT7sg16rkA2KBPfjQnptMKDQe5Ib1 CkzH3k20l/RXfpOGfVW0PJvFpmsJ4SEPiBGcULJ7y6Ka2dJsHjbtd5qr1tXOOUHjCI hptjl7mswpVPaTdIUP3jb3FxRmOrbE5ZKh/qf4qLAnJzk4rRTwPWBoI6E6riAMq/QW THczv4NfLfwlw== From: Jeff Layton Date: Tue, 22 Sep 2026 07:34:06 -0400 Subject: [PATCH v2 5/8] nfsd: give the open file cache its own mutex Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260922-nfsd-per-net-mutex-v2-5-4a5da8243a73@kernel.org> References: <20260922-nfsd-per-net-mutex-v2-0-4a5da8243a73@kernel.org> In-Reply-To: <20260922-nfsd-per-net-mutex-v2-0-4a5da8243a73@kernel.org> To: Chuck Lever , NeilBrown , Olga Kornievskaia , Dai Ngo , Tom Talpey , Shuah Khan Cc: linux-nfs@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, Jeff Layton X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=openpgp-sha256; l=5329; i=jlayton@kernel.org; h=from:subject:message-id; bh=9yOyKfLbgAQCiwsB8y/tXz9WKfPm1n0GduTeDPF6LG0=; b=owEBbQKS/ZANAwAKAQAOaEEZVoIVAcsmYgBqsme50sVKamje/qJ8dxUQVn0zulVB498ElBMWG N/v0VwQd6yJAjMEAAEKAB0WIQRLwNeyRHGyoYTq9dMADmhBGVaCFQUCarJnuQAKCRAADmhBGVaC FetOEACnEZVW4f7WSCmBfIH82E1Km7q41BDWEbBNMvaGUMFxV5qrXG38gEi5NlnccncMa1TVo5f rMurSuL/r+rCCHUharONw8+KlTTTJ1oqy323NvrrGNTd/4RxfH2VD+ey5Jz/wJXdADNIUH1FUyI Jcq9Q7YCTS4qFfjwN2y/10YNSHptVkMFbvss7HkE28BmthI1uBQ4X6AWnewGj4vn7ZQ7WhVqVq2 ScbHDmvF8R2ReddIPD4YgsB4mfrWCq0QddFPEDb8akOulzzIRDt0p3aUwU+Dqqh+frYsqEElo1Y 8QbmpFGY2hdRkU1PY55kvjJPi7cbPXKT4TLC95o3fvucIINmYoCT+WTmRdAH8iAKm0uoPPksPhz kSyrEy+3/4+2pAKIb2ZyNDhCeEpQ45RkqbX1bqB0sEgcoMdVskL3mCLE603LcS0amfg8oKNUczj u89oG05TebWKTh9b22MpvrWKEuwag54XMf8qu8h4xR7Up+K6C9tGWfCqT2QbjFSW2JbX4/FEns1 5ZodEBzQCegml2IKh0W6U+By8mJaP1HDFuFQd3lAiIHvAFgW2CAz1dp8gQa0FcpEUD3uW0DgZDx kMyzJsYT4crJAxPEm1j5bXHQY4NrybKgh+LsoBCuvXa7wjQJ7hlZy6Ggb4o7PdwoSUTsklm/8kD NQUZ29Yp30Fb1Iw== X-Developer-Key: i=jlayton@kernel.org; a=openpgp; fpr=4BC0D7B24471B2A184EAF5D3000E684119568215 nfsd_global_mutex still serializes file cache flushes against server start/stop in unrelated namespaces. nfsd_file_cache_purge() is reached from expkey_flush() on every "exportfs -f", which makes it the only frequent taker of the global lock -- and it has nothing to do with nfsd_users or the NFSv4 global tables it would be waiting on. Move NFSD_FILE_CACHE_UP and the objects it covers (rhltable, LRU, slabs, shrinker, fsnotify groups) onto a nfsd_file_cache_mutex private to filecache.c. nfsd_file_cache_init() and nfsd_file_cache_shutdown() take it themselves rather than asserting the caller holds the global one, so nfssvc.c no longer needs to know how the cache locks itself. nfsd_global_mutex is left holding only nfsd_users, the notifier count and user_recovery_dirname, all of which are touched solely on server start/stop. Ordering is nn->nfsd_mutex outside nfsd_global_mutex outside nfsd_file_cache_mutex. The cache mutex is reached three ways, all consistent with that: from nfsd_startup_generic()/nfsd_shutdown_generic() under the global mutex, from nfsd_shutdown_net() under nn->nfsd_mutex, and bare from expkey_flush() and the filecache stats file. Nothing called with the cache mutex held takes either of the others. Assisted-by: LLM Signed-off-by: Jeff Layton --- fs/nfsd/filecache.c | 29 ++++++++++++++++++++--------- fs/nfsd/nfssvc.c | 9 +++++---- 2 files changed, 25 insertions(+), 13 deletions(-) diff --git a/fs/nfsd/filecache.c b/fs/nfsd/filecache.c index 501772bc8f8d..42bd8c7859cf 100644 --- a/fs/nfsd/filecache.c +++ b/fs/nfsd/filecache.c @@ -67,6 +67,17 @@ */ static DEFINE_SPINLOCK(nfsd_gc_lock); =20 +/* + * Guards NFSD_FILE_CACHE_UP and the host-wide objects it covers: the + * rhltable, the LRU, the slabs, the shrinker and the fsnotify groups. + * Held across cache bring-up and teardown, and by readers and walkers + * that need the cache to stay up for the duration (->cache_purge, the + * stats file). + * + * Nests inside nfsd_global_mutex and inside nn->nfsd_mutex. + */ +static DEFINE_MUTEX(nfsd_file_cache_mutex); + static DEFINE_PER_CPU(unsigned long, nfsd_file_cache_hits); static DEFINE_PER_CPU(unsigned long, nfsd_file_acquisitions); static DEFINE_PER_CPU(unsigned long, nfsd_file_allocations); @@ -869,7 +880,7 @@ nfsd_file_cache_init(void) { int ret; =20 - lockdep_assert_held(&nfsd_global_mutex); + guard(mutex)(&nfsd_file_cache_mutex); if (test_and_set_bit(NFSD_FILE_CACHE_UP, &nfsd_file_flags) =3D=3D 1) return 0; =20 @@ -1011,16 +1022,16 @@ nfsd_file_cache_start_net(struct net *net) * nfsd_file_cache_purge - Remove all cache items associated with @net * @net: target net namespace * - * Takes nfsd_global_mutex so the cache cannot be torn down underneath the - * walk. Callers must not already hold it. + * Takes nfsd_file_cache_mutex so the cache cannot be torn down underneath + * the walk. Callers must not already hold it. */ void nfsd_file_cache_purge(struct net *net) { - mutex_lock(&nfsd_global_mutex); + mutex_lock(&nfsd_file_cache_mutex); if (test_bit(NFSD_FILE_CACHE_UP, &nfsd_file_flags) =3D=3D 1) __nfsd_file_cache_purge(net); - mutex_unlock(&nfsd_global_mutex); + mutex_unlock(&nfsd_file_cache_mutex); } =20 void @@ -1045,7 +1056,7 @@ nfsd_file_cache_shutdown(void) { int i; =20 - lockdep_assert_held(&nfsd_global_mutex); + guard(mutex)(&nfsd_file_cache_mutex); if (test_and_clear_bit(NFSD_FILE_CACHE_UP, &nfsd_file_flags) =3D=3D 0) return; =20 @@ -1476,8 +1487,8 @@ int nfsd_file_cache_stats_show(struct seq_file *m, vo= id *v) unsigned int i, count =3D 0, buckets =3D 0; unsigned long lru =3D 0, total_age =3D 0; =20 - /* Serialize with server shutdown */ - mutex_lock(&nfsd_global_mutex); + /* Serialize with cache teardown */ + mutex_lock(&nfsd_file_cache_mutex); if (test_bit(NFSD_FILE_CACHE_UP, &nfsd_file_flags) =3D=3D 1) { struct bucket_table *tbl; struct rhashtable *ht; @@ -1491,7 +1502,7 @@ int nfsd_file_cache_stats_show(struct seq_file *m, vo= id *v) buckets =3D tbl->size; rcu_read_unlock(); } - mutex_unlock(&nfsd_global_mutex); + mutex_unlock(&nfsd_file_cache_mutex); =20 for_each_possible_cpu(i) { hits +=3D per_cpu(nfsd_file_cache_hits, i); diff --git a/fs/nfsd/nfssvc.c b/fs/nfsd/nfssvc.c index 5cb92c3829f9..b2c215775e08 100644 --- a/fs/nfsd/nfssvc.c +++ b/fs/nfsd/nfssvc.c @@ -66,11 +66,12 @@ static __be32 nfsd_init_request(struct svc_rqst *, * * nfsd_global_mutex covers only what is genuinely shared between * namespaces: the nfsd_users refcount and the host-wide resources it - * brings up and tears down (the open file cache and the NFSv4 global - * tables), the address-notifier registration, and user_recovery_dirname. + * brings up and tears down (the NFSv4 global tables, and the open file + * cache -- which guards its own internals with nfsd_file_cache_mutex), + * the address-notifier registration, and user_recovery_dirname. * - * Lock ordering is nn->nfsd_mutex outside nfsd_global_mutex. Nothing tak= es - * two namespaces' nfsd_mutexes. + * Lock ordering is nn->nfsd_mutex outside nfsd_global_mutex outside + * nfsd_file_cache_mutex. Nothing takes two namespaces' nfsd_mutexes. */ DEFINE_MUTEX(nfsd_global_mutex); =20 --=20 2.55.0 From nobody Thu Sep 24 16:07:44 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B9D3547FB0C; Tue, 22 Sep 2026 11:34:25 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790076868; cv=none; b=TEYGljP/EspiOeQe0tqM9Xyg0/d0LLEMg32tboZx0bgBnsk/gFN/2RsSxcKb0xyNT+QG+Jghcgy39JAuXecgVhNe3AdKERaAzGwUfcr+KqLqtN+EB4UThiOEQIsWPVO8bAJ2QS+FKC/Yn18IzTiBc/ppGSH4EAEc/fShpdCRUhU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790076868; c=relaxed/simple; bh=E8vGIRGIZ5Ej8iVy508sE++9uoe+ty5nWr2vlsyp4Xw=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=sl4jYqPagfZm5tPgTDBig1tqQkYW8mpbq4IMDvQ0VzDGumQFVSzgWz6H44d26rsqoMcI5rgm80FwlTYmiIp6NG8WX2SOSScBaZiQEkAoALlA0cGnyNNS4hzpBaaDMmyBLryfVdn/6Sv2tlvha/Sxn9Rk196yx3nH1/ieW2yuDXk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=AvL+l3Qj; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="AvL+l3Qj" Received: by smtp.kernel.org (Postfix) with ESMTPSA id A164E1F00898; Tue, 22 Sep 2026 11:34:24 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1790076865; bh=2aVXY9iiOPtDAAWpb3Jyxfl+AE25CNbFUHutOqPD/6c=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=AvL+l3QjUdv78sXs0ZzDvxviyFWjSHMjAn3uo1XIrC67WklTZzLYsM8RBxVm2OoLd A7fX8hSVWER4Zy5z99fOwW8Z6vKm3olj2/mH9ez3sDU0slaYG3gATiKO1fQm42T8wk Ok241egWSbp2MWA3kuXFqR500z++dEQ9Fi8jgL/iuKhUch3X9ndS99o43YzWeoahwQ a/iVqSIU/8XNbe610gmWnAYLgMHBW6MdwjfKDHDKL6DRnEY6tts6vL/75Qq0w3xhsW +ZoCjNbFiu71pUZR2p63p+/0PuLWjh2wkV7X7kppmn5tkbR9yCw3cTzfVC+nmdsndI oQJCs6V5wKGxA== From: Jeff Layton Date: Tue, 22 Sep 2026 07:34:07 -0400 Subject: [PATCH v2 6/8] selftests/nfsd: factor the netlink plumbing into a shared header Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260922-nfsd-per-net-mutex-v2-6-4a5da8243a73@kernel.org> References: <20260922-nfsd-per-net-mutex-v2-0-4a5da8243a73@kernel.org> In-Reply-To: <20260922-nfsd-per-net-mutex-v2-0-4a5da8243a73@kernel.org> To: Chuck Lever , NeilBrown , Olga Kornievskaia , Dai Ngo , Tom Talpey , Shuah Khan Cc: linux-nfs@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, Jeff Layton X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=openpgp-sha256; l=20983; i=jlayton@kernel.org; h=from:subject:message-id; bh=E8vGIRGIZ5Ej8iVy508sE++9uoe+ty5nWr2vlsyp4Xw=; b=owEBbQKS/ZANAwAKAQAOaEEZVoIVAcsmYgBqsme5PAihfAmBedzHK5Z62IPF0XRZhGv/Fx+Sy OHhWc1CYi+JAjMEAAEKAB0WIQRLwNeyRHGyoYTq9dMADmhBGVaCFQUCarJnuQAKCRAADmhBGVaC FaAkEACJwD5jhrxCvI38sNz2b8TuDlTaBafPgc4fVTgnksSXrdDRvNVKqvC7Y1EnNwf4QXNgC6+ Xb9EimX2ZM3IhkcEWZ8bbjCcKyKZecGqovEtf+sE/E10H8gqUDymHtmWenSOwcqG+dgr0VV7gGi RlX3egerXdgM9gSPUIfSo3HyS3Z9xhnyZIn3PsxEHfm+JHZYO8tXmEmMRUuWWx6c9JGAuKjJOmK +i1vXSkGYgHCwn7kaFk49otkvKS5zSSh8C7/dApWFCuasD6pO0tjOYnC2bAPEps64w31UuIhNxv 1cotobzCqEH31Amjec/lv8cmXNmR3Xk9c0j3q8z/VT8feYt/pJWQ/j5c9HEPg45jYW9qzxtHiMT Ss2n7aRCIlC46AP6XWE1P3VQCnGCJpAJeQ+pqXawtd0hK5K3Z4Pdk7/0RTTCVZXD6clXNP87Llr Q9Nv73+oYOGvuQUX1RUeA3D9YHTbnU1+te/4yqimtZBgLK+irMiDp9DGJNSJ8+pLIFY6uuRZVNw SkHiNqPAGBSXk/varH42tDpIQYOEEPEcPqUNfNM/USZomm6bZmItrM1S0MO8M7wATsL7IqlOZNK oX+l39Idv/Lka3+He6oJ1e1Z2dN5+e1bFd7LUhvNgRTOzIJyei4HOy4mgKrH19S9kP77PjBAAFZ jqXtgVYBYMZRdEA== X-Developer-Key: i=jlayton@kernel.org; a=openpgp; fpr=4BC0D7B24471B2A184EAF5D3000E684119568215 nfsd_netlink_listener.c carries a small generic-netlink client -- socket setup, attribute packing, extack parsing, family resolution -- plus the LISTENER_SET/THREADS_SET/VERSION_SET wrappers. Other nfsd tests want the same thing. Move it to nfsd_netlink.h as static inline helpers, so each test binary gets its own copy and there is nothing to link. Code is unchanged; only the three continuation lines that the static inline prefix pushed out of alignment were reflowed. Assisted-by: LLM Signed-off-by: Jeff Layton --- tools/testing/selftests/nfsd/nfsd_netlink.h | 319 +++++++++++++++++= ++++ .../testing/selftests/nfsd/nfsd_netlink_listener.c | 295 +----------------= -- 2 files changed, 322 insertions(+), 292 deletions(-) diff --git a/tools/testing/selftests/nfsd/nfsd_netlink.h b/tools/testing/se= lftests/nfsd/nfsd_netlink.h new file mode 100644 index 000000000000..c66a980bfc01 --- /dev/null +++ b/tools/testing/selftests/nfsd/nfsd_netlink.h @@ -0,0 +1,319 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Shared generic-netlink plumbing for the NFSD selftests. + * + * Header-only: every helper is static inline, so each test binary gets its + * own copy and there is nothing extra to link. nfsd_family must be set by + * calling genl_resolve_nfsd() before any of the request helpers are used. + */ +#ifndef __SELFTESTS_NFSD_NETLINK_H__ +#define __SELFTESTS_NFSD_NETLINK_H__ + +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#define NLA_ALIGN4(len) (((len) + 3) & ~3) +#define RECV_TIMEO_SEC 30 + +static int nfsd_family =3D -1; /* set per-test in FIXTURE_SETUP */ + +/* Extack message from the last genl_request(); empty if there was none. */ +static char last_extack[128]; + +static inline void die(const char *msg) +{ + perror(msg); + exit(1); +} + +/* ------------------- minimal generic-netlink plumbing ------------------= - */ + +static inline int genl_open(void) +{ + struct sockaddr_nl sa =3D { .nl_family =3D AF_NETLINK }; + struct timeval tv =3D { .tv_sec =3D RECV_TIMEO_SEC }; + int fd =3D socket(AF_NETLINK, SOCK_RAW, NETLINK_GENERIC); + int on =3D 1; + + if (fd < 0) + die("socket(NETLINK_GENERIC)"); + if (bind(fd, (void *)&sa, sizeof(sa)) < 0) + die("bind(netlink)"); + setsockopt(fd, SOL_SOCKET, SO_RCVTIMEO, &tv, sizeof(tv)); + /* + * Ask for extack, and cap the ack so the request is not echoed back: + * the TLVs then always follow the fixed part of the error message. + */ + setsockopt(fd, SOL_NETLINK, NETLINK_EXT_ACK, &on, sizeof(on)); + setsockopt(fd, SOL_NETLINK, NETLINK_CAP_ACK, &on, sizeof(on)); + return fd; +} + +/* Stash the extack message of an ack, if it carries one. */ +static inline void parse_extack(const char *rbuf) +{ + const struct nlmsghdr *nlh =3D (const void *)rbuf; + const struct nlattr *na; + int off, left; + + last_extack[0] =3D '\0'; + if (nlh->nlmsg_type !=3D NLMSG_ERROR || + !(nlh->nlmsg_flags & NLM_F_ACK_TLVS)) + return; + + off =3D NLMSG_HDRLEN + NLMSG_ALIGN(sizeof(struct nlmsgerr)); + left =3D nlh->nlmsg_len - off; + na =3D (const void *)(rbuf + off); + + while (left >=3D (int)NLA_HDRLEN) { + if ((na->nla_type & NLA_TYPE_MASK) =3D=3D NLMSGERR_ATTR_MSG) { + strncpy(last_extack, (const char *)na + NLA_HDRLEN, + sizeof(last_extack) - 1); + last_extack[sizeof(last_extack) - 1] =3D '\0'; + return; + } + left -=3D NLA_ALIGN4(na->nla_len); + na =3D (const void *)((const char *)na + NLA_ALIGN4(na->nla_len)); + } +} + +/* Append an attribute at @off; return the new (aligned) offset. */ +static inline int put_attr(char *buf, int off, uint16_t type, + const void *data, int len) +{ + struct nlattr *na =3D (void *)(buf + off); + + na->nla_type =3D type; + na->nla_len =3D NLA_HDRLEN + len; + if (len) + memcpy(buf + off + NLA_HDRLEN, data, len); + return off + NLA_ALIGN4(NLA_HDRLEN + len); +} + +/* Build a genl message header into @buf; return the offset past it. */ +static inline int genl_hdr(char *buf, uint16_t type, uint16_t flags, uint8= _t cmd) +{ + struct nlmsghdr *nlh =3D (void *)buf; + struct genlmsghdr *gnl =3D (void *)(buf + NLMSG_HDRLEN); + + memset(buf, 0, NLMSG_HDRLEN + GENL_HDRLEN); + nlh->nlmsg_type =3D type; + nlh->nlmsg_flags =3D flags; + nlh->nlmsg_seq =3D 1; + gnl->cmd =3D cmd; + gnl->version =3D 1; + return NLMSG_HDRLEN + GENL_HDRLEN; +} + +/* Send an nfsd command with an ACK; return the ACK errno (<=3D 0). */ +static inline int genl_request(uint8_t cmd, const char *attrs, int attrs_l= en) +{ + char buf[1 << 20], rbuf[4096]; + struct nlmsghdr *nlh =3D (void *)buf; + int fd =3D genl_open(); + int off, n, ret; + + off =3D genl_hdr(buf, nfsd_family, NLM_F_REQUEST | NLM_F_ACK, cmd); + if (attrs_len) { + memcpy(buf + off, attrs, attrs_len); + off +=3D attrs_len; + } + nlh->nlmsg_len =3D off; + + if (send(fd, buf, off, 0) < 0) + die("send(genl)"); + + last_extack[0] =3D '\0'; + n =3D recv(fd, rbuf, sizeof(rbuf), 0); + if (n < 0) { + ret =3D (errno =3D=3D EAGAIN || errno =3D=3D EWOULDBLOCK) ? -ETIMEDOUT := -errno; + } else if (((struct nlmsghdr *)rbuf)->nlmsg_type =3D=3D NLMSG_ERROR) { + ret =3D ((struct nlmsgerr *)NLMSG_DATA(rbuf))->error; + parse_extack(rbuf); + } else { + ret =3D 0; + } + close(fd); + return ret; +} + +/* + * Send a command with attributes and return the full reply message; -errno + * on failure. NLM_F_ACK is left off: the kernel reports an error either w= ay, + * so the first message back is the reply whenever there is one. + */ +static inline int genl_request_reply_attrs(uint8_t cmd, const char *attrs, + int attrs_len, char *rbuf, + size_t rlen) +{ + char buf[1 << 20]; + struct nlmsghdr *nlh =3D (void *)buf; + int fd =3D genl_open(); + int off, n, ret; + + off =3D genl_hdr(buf, nfsd_family, NLM_F_REQUEST, cmd); + if (attrs_len) { + memcpy(buf + off, attrs, attrs_len); + off +=3D attrs_len; + } + nlh->nlmsg_len =3D off; + + if (send(fd, buf, off, 0) < 0) + die("send(genl reply)"); + + n =3D recv(fd, rbuf, rlen, 0); + if (n < 0) + ret =3D (errno =3D=3D EAGAIN || errno =3D=3D EWOULDBLOCK) ? -ETIMEDOUT := -errno; + else if (((struct nlmsghdr *)rbuf)->nlmsg_type =3D=3D NLMSG_ERROR) + ret =3D ((struct nlmsgerr *)NLMSG_DATA(rbuf))->error; + else + ret =3D n; + close(fd); + return ret; +} + +static inline int genl_request_reply(uint8_t cmd, char *rbuf, size_t rlen) +{ + return genl_request_reply_attrs(cmd, NULL, 0, rbuf, rlen); +} + +/* Resolve the "nfsd" genl family id; -1 if not registered. */ +static inline int genl_resolve_nfsd(void) +{ + char buf[1024], rbuf[4096]; + struct nlmsghdr *nlh =3D (void *)buf; + struct nlmsghdr *rh =3D (void *)rbuf; + struct nlattr *na; + int fd, off, left, id =3D -1; + + fd =3D genl_open(); + off =3D genl_hdr(buf, GENL_ID_CTRL, NLM_F_REQUEST, CTRL_CMD_GETFAMILY); + off =3D put_attr(buf, off, CTRL_ATTR_FAMILY_NAME, + NFSD_FAMILY_NAME, sizeof(NFSD_FAMILY_NAME)); + nlh->nlmsg_len =3D off; + + if (send(fd, buf, off, 0) < 0) + die("send(GETFAMILY)"); + if (recv(fd, rbuf, sizeof(rbuf), 0) < 0) + die("recv(GETFAMILY)"); + close(fd); + + if (rh->nlmsg_type =3D=3D NLMSG_ERROR) + return -1; + + na =3D (void *)((char *)NLMSG_DATA(rh) + GENL_HDRLEN); + left =3D rh->nlmsg_len - NLMSG_HDRLEN - GENL_HDRLEN; + while (left >=3D (int)NLA_HDRLEN) { + if (na->nla_type =3D=3D CTRL_ATTR_FAMILY_ID) { + id =3D *(uint16_t *)((char *)na + NLA_HDRLEN); + break; + } + left -=3D NLA_ALIGN4(na->nla_len); + na =3D (void *)((char *)na + NLA_ALIGN4(na->nla_len)); + } + return id; +} + +/* ------------------- listener request builders ------------------- */ + +/* Fine-grained control for negative tests: any field can be omitted/malfo= rmed. */ +struct raw_listener { + const char *xprt; /* NULL -> omit NFSD_A_SOCK_TRANSPORT_NAME */ + int emit_addr; /* 0 -> omit NFSD_A_SOCK_ADDR */ + const void *addr; + int addr_len; /* bytes to emit for NFSD_A_SOCK_ADDR */ +}; + +static inline int put_raw_listener(char *buf, int off, const struct raw_li= stener *r) +{ + struct nlattr *nest =3D (void *)(buf + off); + int inner =3D off + NLA_HDRLEN; + + if (r->emit_addr) + inner =3D put_attr(buf, inner, NFSD_A_SOCK_ADDR, r->addr, r->addr_len); + if (r->xprt) + inner =3D put_attr(buf, inner, NFSD_A_SOCK_TRANSPORT_NAME, + r->xprt, strlen(r->xprt) + 1); + nest->nla_type =3D NFSD_A_SERVER_SOCK_ADDR | NLA_F_NESTED; + nest->nla_len =3D inner - off; + return off + NLA_ALIGN4(nest->nla_len); +} + +/* Well-formed loopback listener for @family (AF_INET or AF_INET6). */ +static inline int put_listener_af(char *buf, int off, const char *xprt, + int family, uint16_t port) +{ + struct sockaddr_storage ss =3D {0}; + struct raw_listener r =3D { .xprt =3D xprt, .emit_addr =3D 1, .addr =3D &= ss }; + + if (family =3D=3D AF_INET6) { + struct sockaddr_in6 *s6 =3D (void *)&ss; + + s6->sin6_family =3D AF_INET6; + s6->sin6_port =3D htons(port); + s6->sin6_addr =3D in6addr_loopback; + r.addr_len =3D sizeof(*s6); + } else { + struct sockaddr_in *s4 =3D (void *)&ss; + + s4->sin_family =3D AF_INET; + s4->sin_port =3D htons(port); + s4->sin_addr.s_addr =3D htonl(INADDR_LOOPBACK); + r.addr_len =3D sizeof(*s4); + } + return put_raw_listener(buf, off, &r); +} + +static inline int put_listener(char *buf, int off, const char *xprt, uint1= 6_t port) +{ + return put_listener_af(buf, off, xprt, AF_INET, port); +} + +static inline int listener_set(const char *attrs, int len) +{ + return genl_request(NFSD_CMD_LISTENER_SET, attrs, len); +} + +/* + * Enable exactly one NFS version in this netns. NFSD_CMD_VERSION_SET clea= rs + * every version first, so one nest is enough to leave the server v4-only. + * It refuses once a serv exists, so call it before any listener. + */ +static inline int version_set_only(uint32_t major, uint32_t minor) +{ + char attrs[64]; + struct nlattr *nest =3D (void *)attrs; + int inner =3D NLA_HDRLEN; + + inner =3D put_attr(attrs, inner, NFSD_A_VERSION_MAJOR, + &major, sizeof(major)); + inner =3D put_attr(attrs, inner, NFSD_A_VERSION_MINOR, + &minor, sizeof(minor)); + inner =3D put_attr(attrs, inner, NFSD_A_VERSION_ENABLED, NULL, 0); + nest->nla_type =3D NFSD_A_SERVER_PROTO_VERSION | NLA_F_NESTED; + nest->nla_len =3D inner; + + return genl_request(NFSD_CMD_VERSION_SET, attrs, NLA_ALIGN4(inner)); +} + +/* Start (@n > 0) or stop (@n =3D=3D 0) nfsd threads in this netns. */ +static inline int threads_set(int n) +{ + char attrs[64]; + uint32_t v =3D n; + int off =3D put_attr(attrs, 0, NFSD_A_SERVER_THREADS, &v, sizeof(v)); + + return genl_request(NFSD_CMD_THREADS_SET, attrs, off); +} + +#endif /* __SELFTESTS_NFSD_NETLINK_H__ */ diff --git a/tools/testing/selftests/nfsd/nfsd_netlink_listener.c b/tools/t= esting/selftests/nfsd/nfsd_netlink_listener.c index 1294057b6f62..f6f8aa1e1cfe 100644 --- a/tools/testing/selftests/nfsd/nfsd_netlink_listener.c +++ b/tools/testing/selftests/nfsd/nfsd_netlink_listener.c @@ -46,267 +46,13 @@ #include =20 #include "../kselftest_harness.h" +#include "nfsd_netlink.h" =20 -#define NFS_PROGRAM 100003 -#define NFS_ACL_PROGRAM 100227 - -#define NLA_ALIGN4(len) (((len) + 3) & ~3) #define TEST_PORT 20049 #define MAX_LISTENERS 8 -#define RECV_TIMEO_SEC 30 - -static int nfsd_family =3D -1; /* set per-test in FIXTURE_SETUP */ - -/* Extack message from the last genl_request(); empty if there was none. */ -static char last_extack[128]; - -static void die(const char *msg) -{ - perror(msg); - exit(1); -} - -/* ------------------- minimal generic-netlink plumbing ------------------= - */ - -static int genl_open(void) -{ - struct sockaddr_nl sa =3D { .nl_family =3D AF_NETLINK }; - struct timeval tv =3D { .tv_sec =3D RECV_TIMEO_SEC }; - int fd =3D socket(AF_NETLINK, SOCK_RAW, NETLINK_GENERIC); - int on =3D 1; - - if (fd < 0) - die("socket(NETLINK_GENERIC)"); - if (bind(fd, (void *)&sa, sizeof(sa)) < 0) - die("bind(netlink)"); - setsockopt(fd, SOL_SOCKET, SO_RCVTIMEO, &tv, sizeof(tv)); - /* - * Ask for extack, and cap the ack so the request is not echoed back: - * the TLVs then always follow the fixed part of the error message. - */ - setsockopt(fd, SOL_NETLINK, NETLINK_EXT_ACK, &on, sizeof(on)); - setsockopt(fd, SOL_NETLINK, NETLINK_CAP_ACK, &on, sizeof(on)); - return fd; -} - -/* Stash the extack message of an ack, if it carries one. */ -static void parse_extack(const char *rbuf) -{ - const struct nlmsghdr *nlh =3D (const void *)rbuf; - const struct nlattr *na; - int off, left; - - last_extack[0] =3D '\0'; - if (nlh->nlmsg_type !=3D NLMSG_ERROR || - !(nlh->nlmsg_flags & NLM_F_ACK_TLVS)) - return; - - off =3D NLMSG_HDRLEN + NLMSG_ALIGN(sizeof(struct nlmsgerr)); - left =3D nlh->nlmsg_len - off; - na =3D (const void *)(rbuf + off); - - while (left >=3D (int)NLA_HDRLEN) { - if ((na->nla_type & NLA_TYPE_MASK) =3D=3D NLMSGERR_ATTR_MSG) { - strncpy(last_extack, (const char *)na + NLA_HDRLEN, - sizeof(last_extack) - 1); - last_extack[sizeof(last_extack) - 1] =3D '\0'; - return; - } - left -=3D NLA_ALIGN4(na->nla_len); - na =3D (const void *)((const char *)na + NLA_ALIGN4(na->nla_len)); - } -} - -/* Append an attribute at @off; return the new (aligned) offset. */ -static int put_attr(char *buf, int off, uint16_t type, - const void *data, int len) -{ - struct nlattr *na =3D (void *)(buf + off); - - na->nla_type =3D type; - na->nla_len =3D NLA_HDRLEN + len; - if (len) - memcpy(buf + off + NLA_HDRLEN, data, len); - return off + NLA_ALIGN4(NLA_HDRLEN + len); -} - -/* Build a genl message header into @buf; return the offset past it. */ -static int genl_hdr(char *buf, uint16_t type, uint16_t flags, uint8_t cmd) -{ - struct nlmsghdr *nlh =3D (void *)buf; - struct genlmsghdr *gnl =3D (void *)(buf + NLMSG_HDRLEN); - - memset(buf, 0, NLMSG_HDRLEN + GENL_HDRLEN); - nlh->nlmsg_type =3D type; - nlh->nlmsg_flags =3D flags; - nlh->nlmsg_seq =3D 1; - gnl->cmd =3D cmd; - gnl->version =3D 1; - return NLMSG_HDRLEN + GENL_HDRLEN; -} - -/* Send an nfsd command with an ACK; return the ACK errno (<=3D 0). */ -static int genl_request(uint8_t cmd, const char *attrs, int attrs_len) -{ - char buf[1 << 20], rbuf[4096]; - struct nlmsghdr *nlh =3D (void *)buf; - int fd =3D genl_open(); - int off, n, ret; - - off =3D genl_hdr(buf, nfsd_family, NLM_F_REQUEST | NLM_F_ACK, cmd); - if (attrs_len) { - memcpy(buf + off, attrs, attrs_len); - off +=3D attrs_len; - } - nlh->nlmsg_len =3D off; - - if (send(fd, buf, off, 0) < 0) - die("send(genl)"); - - last_extack[0] =3D '\0'; - n =3D recv(fd, rbuf, sizeof(rbuf), 0); - if (n < 0) { - ret =3D (errno =3D=3D EAGAIN || errno =3D=3D EWOULDBLOCK) ? -ETIMEDOUT := -errno; - } else if (((struct nlmsghdr *)rbuf)->nlmsg_type =3D=3D NLMSG_ERROR) { - ret =3D ((struct nlmsgerr *)NLMSG_DATA(rbuf))->error; - parse_extack(rbuf); - } else { - ret =3D 0; - } - close(fd); - return ret; -} - -/* - * Send a command with attributes and return the full reply message; -errno - * on failure. NLM_F_ACK is left off: the kernel reports an error either w= ay, - * so the first message back is the reply whenever there is one. - */ -static int genl_request_reply_attrs(uint8_t cmd, const char *attrs, - int attrs_len, char *rbuf, size_t rlen) -{ - char buf[1 << 20]; - struct nlmsghdr *nlh =3D (void *)buf; - int fd =3D genl_open(); - int off, n, ret; - - off =3D genl_hdr(buf, nfsd_family, NLM_F_REQUEST, cmd); - if (attrs_len) { - memcpy(buf + off, attrs, attrs_len); - off +=3D attrs_len; - } - nlh->nlmsg_len =3D off; =20 - if (send(fd, buf, off, 0) < 0) - die("send(genl reply)"); - - n =3D recv(fd, rbuf, rlen, 0); - if (n < 0) - ret =3D (errno =3D=3D EAGAIN || errno =3D=3D EWOULDBLOCK) ? -ETIMEDOUT := -errno; - else if (((struct nlmsghdr *)rbuf)->nlmsg_type =3D=3D NLMSG_ERROR) - ret =3D ((struct nlmsgerr *)NLMSG_DATA(rbuf))->error; - else - ret =3D n; - close(fd); - return ret; -} - -static int genl_request_reply(uint8_t cmd, char *rbuf, size_t rlen) -{ - return genl_request_reply_attrs(cmd, NULL, 0, rbuf, rlen); -} - -/* Resolve the "nfsd" genl family id; -1 if not registered. */ -static int genl_resolve_nfsd(void) -{ - char buf[1024], rbuf[4096]; - struct nlmsghdr *nlh =3D (void *)buf; - struct nlmsghdr *rh =3D (void *)rbuf; - struct nlattr *na; - int fd, off, left, id =3D -1; - - fd =3D genl_open(); - off =3D genl_hdr(buf, GENL_ID_CTRL, NLM_F_REQUEST, CTRL_CMD_GETFAMILY); - off =3D put_attr(buf, off, CTRL_ATTR_FAMILY_NAME, - NFSD_FAMILY_NAME, sizeof(NFSD_FAMILY_NAME)); - nlh->nlmsg_len =3D off; - - if (send(fd, buf, off, 0) < 0) - die("send(GETFAMILY)"); - if (recv(fd, rbuf, sizeof(rbuf), 0) < 0) - die("recv(GETFAMILY)"); - close(fd); - - if (rh->nlmsg_type =3D=3D NLMSG_ERROR) - return -1; - - na =3D (void *)((char *)NLMSG_DATA(rh) + GENL_HDRLEN); - left =3D rh->nlmsg_len - NLMSG_HDRLEN - GENL_HDRLEN; - while (left >=3D (int)NLA_HDRLEN) { - if (na->nla_type =3D=3D CTRL_ATTR_FAMILY_ID) { - id =3D *(uint16_t *)((char *)na + NLA_HDRLEN); - break; - } - left -=3D NLA_ALIGN4(na->nla_len); - na =3D (void *)((char *)na + NLA_ALIGN4(na->nla_len)); - } - return id; -} - -/* ------------------- listener request builders ------------------- */ - -/* Fine-grained control for negative tests: any field can be omitted/malfo= rmed. */ -struct raw_listener { - const char *xprt; /* NULL -> omit NFSD_A_SOCK_TRANSPORT_NAME */ - int emit_addr; /* 0 -> omit NFSD_A_SOCK_ADDR */ - const void *addr; - int addr_len; /* bytes to emit for NFSD_A_SOCK_ADDR */ -}; - -static int put_raw_listener(char *buf, int off, const struct raw_listener = *r) -{ - struct nlattr *nest =3D (void *)(buf + off); - int inner =3D off + NLA_HDRLEN; - - if (r->emit_addr) - inner =3D put_attr(buf, inner, NFSD_A_SOCK_ADDR, r->addr, r->addr_len); - if (r->xprt) - inner =3D put_attr(buf, inner, NFSD_A_SOCK_TRANSPORT_NAME, - r->xprt, strlen(r->xprt) + 1); - nest->nla_type =3D NFSD_A_SERVER_SOCK_ADDR | NLA_F_NESTED; - nest->nla_len =3D inner - off; - return off + NLA_ALIGN4(nest->nla_len); -} - -/* Well-formed loopback listener for @family (AF_INET or AF_INET6). */ -static int put_listener_af(char *buf, int off, const char *xprt, int famil= y, - uint16_t port) -{ - struct sockaddr_storage ss =3D {0}; - struct raw_listener r =3D { .xprt =3D xprt, .emit_addr =3D 1, .addr =3D &= ss }; - - if (family =3D=3D AF_INET6) { - struct sockaddr_in6 *s6 =3D (void *)&ss; - - s6->sin6_family =3D AF_INET6; - s6->sin6_port =3D htons(port); - s6->sin6_addr =3D in6addr_loopback; - r.addr_len =3D sizeof(*s6); - } else { - struct sockaddr_in *s4 =3D (void *)&ss; - - s4->sin_family =3D AF_INET; - s4->sin_port =3D htons(port); - s4->sin_addr.s_addr =3D htonl(INADDR_LOOPBACK); - r.addr_len =3D sizeof(*s4); - } - return put_raw_listener(buf, off, &r); -} - -static int put_listener(char *buf, int off, const char *xprt, uint16_t por= t) -{ - return put_listener_af(buf, off, xprt, AF_INET, port); -} +#define NFS_PROGRAM 100003 +#define NFS_ACL_PROGRAM 100227 =20 /* ------------------- LISTENER_GET parsing ------------------- */ =20 @@ -374,32 +120,6 @@ static int parse_listener_get(const char *rbuf, int le= n, =20 /* ------------------- convenience wrappers ------------------- */ =20 -static int listener_set(const char *attrs, int len) -{ - return genl_request(NFSD_CMD_LISTENER_SET, attrs, len); -} - -/* - * Enable exactly one NFS version in this netns. NFSD_CMD_VERSION_SET clea= rs - * every version first, so one nest is enough to leave the server v4-only. - * It refuses once a serv exists, so call it before any listener. - */ -static int version_set_only(uint32_t major, uint32_t minor) -{ - char attrs[64]; - struct nlattr *nest =3D (void *)attrs; - int inner =3D NLA_HDRLEN; - - inner =3D put_attr(attrs, inner, NFSD_A_VERSION_MAJOR, - &major, sizeof(major)); - inner =3D put_attr(attrs, inner, NFSD_A_VERSION_MINOR, - &minor, sizeof(minor)); - inner =3D put_attr(attrs, inner, NFSD_A_VERSION_ENABLED, NULL, 0); - nest->nla_type =3D NFSD_A_SERVER_PROTO_VERSION | NLA_F_NESTED; - nest->nla_len =3D inner; - - return genl_request(NFSD_CMD_VERSION_SET, attrs, NLA_ALIGN4(inner)); -} =20 /* ------------------- userspace-rpcbind ------------------- */ =20 @@ -552,15 +272,6 @@ static struct listener_ent *find_listener(struct liste= ner_ent *e, int n, return NULL; } =20 -/* Start (@n > 0) or stop (@n =3D=3D 0) nfsd threads in this netns. */ -static int threads_set(int n) -{ - char attrs[64]; - uint32_t v =3D n; - int off =3D put_attr(attrs, 0, NFSD_A_SERVER_THREADS, &v, sizeof(v)); - - return genl_request(NFSD_CMD_THREADS_SET, attrs, off); -} =20 /* ------------------- per-netns local rpcbind stub ------------------- */ =20 --=20 2.55.0 From nobody Thu Sep 24 16:07:44 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 79D0C53C3CC; Tue, 22 Sep 2026 11:34:26 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790076868; cv=none; b=dwuY5rFRFxF4vWbRSlh0r3cgo/QRVMcIghNgXpNF0J7K5ehrthr7tLzePh9xeKjyxQA1snSdJBUrW+dHTKvbYrcv6qq5uLHzPdZzmbgT18Blm8W/u998rwJqjTiPSTrbGQBjN6c9CW3c1WuhhmxLN1wm2VZBY9PuUrBrXYt3VC8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790076868; c=relaxed/simple; bh=kOuMQpouklLFUeC63Q/1ytRszZTYqhpNjFTVkX/gXnY=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=WgeTQy6rkYBhpYC4GOBWFqqUB3XVtCHGvk3ZWRWvpYANi+HuPLWJF2XIJwcscTejVFdvw4dlP52Hc6oI5t3GN3cbDzyGoIe3AanX73WMKAS14GxB0tb0Qjv52j8pnUC9R9wS12eOmg6sxLahctQ0qoeL5xZkSvpMn8k2Rr8NwIU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=ZnKrPllV; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="ZnKrPllV" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 9F5C41F0089B; Tue, 22 Sep 2026 11:34:25 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1790076866; bh=DZ7DioWmuUY4s9zHzGXWjWhaeMChyVlk3Mqm1MIVMX0=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=ZnKrPllV/jfU9jLuWPc4pgPZnv4Cma5gSqzFuzaM2E/G/+Tn4aOQLbzgMEo6os3j8 YqCWz38c7U6EhuQzmVK2p/4KmC4yFnb6OGwTqbzrre3ZgAGQoSWVM5pu1nx+gfQ/Wo x5q7eZXxbV0L9RzPKYgYXynwm/b5fKN9yZXteQgMS9mqnrc6BCADtSjLv1OVuiFL6E U8RSkZ/TQLQuMGAx1lZ8rLv59gfv4j7ml3Ib2/0p85rwarM3H84enSOpCjcS1ZBZDy K0hdt06D0hjjC8iQE8ce/Z1MO52T/+05SnKFOtRRBoLyA54kc/DMXrOT/8KjsXEOdh ElRE46EA2tgNg== From: Jeff Layton Date: Tue, 22 Sep 2026 07:34:08 -0400 Subject: [PATCH v2 7/8] selftests/nfsd: add cross-namespace isolation tests Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260922-nfsd-per-net-mutex-v2-7-4a5da8243a73@kernel.org> References: <20260922-nfsd-per-net-mutex-v2-0-4a5da8243a73@kernel.org> In-Reply-To: <20260922-nfsd-per-net-mutex-v2-0-4a5da8243a73@kernel.org> To: Chuck Lever , NeilBrown , Olga Kornievskaia , Dai Ngo , Tom Talpey , Shuah Khan Cc: linux-nfs@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, Jeff Layton X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=openpgp-sha256; l=15892; i=jlayton@kernel.org; h=from:subject:message-id; bh=kOuMQpouklLFUeC63Q/1ytRszZTYqhpNjFTVkX/gXnY=; b=owEBbQKS/ZANAwAKAQAOaEEZVoIVAcsmYgBqsme6VV1ytjSUpbRIPQ0J9vQO1EO026VXdWK/7 mY9COTvkdmJAjMEAAEKAB0WIQRLwNeyRHGyoYTq9dMADmhBGVaCFQUCarJnugAKCRAADmhBGVaC FS4iEACow7yfFH9XO88jd7iO2qumFIDeQ5EWk3EzOQ0HCD1GaPug0nf4H4QZnVUVaOn5yP1Iw2R VhvQzC/0or/gdyUFaLlY3iY58L4L5jYNjOVFS5iDJE7GA6ERdix14M4BgCN7cUG/lar+v4+WwYv Ds2cuxWKB1RXPYgMUFNxquq/6GIac/gd5EyF2siBMw8VW2+/ETZEFijyYhBj15ijoQwIquLg7Uu rI60OHEiR8HsFnYaX6mA0+yPKzemobPH4i9/Qk0s4/IO94u45b38T/zAPdI9eiY0BB4ubWnrpwZ pwnbvGsRouGp9AaV5ZRddgfYLDoTnNJlsJSXllOX8vgfodPi9ZYxvYkkSICv5SSsasMX2WiJUnT OvMtHyPMOusRRvNBQVOom4Am8nG/KQ5cLhnDZ2+rrKMqHRWvlwkiWzVda/rUAwMADxmEnWs6MD/ sn0GXpPExHWKTh4a0i1y9IWsRYD/tDV1m7vsBKUNFUDM5DJ1KS3+JcmbYwt+2zGK/Xm55hTK4Gi acmMmJkiqOXE+LF8zxhbB10mJ66piWgk4SnVev/3cxsQWUKbT1oIhGJpFfO8lrk0iyhGnleWuJa LZ11kZVa21bmeOx+9HdUy3S6PDK9wRHNdnZRebtJfpANtcJuJmyU/uJvwDnrZkAhFlSGS194M1F UVVe7RwV09Wh68A== X-Developer-Key: i=jlayton@kernel.org; a=openpgp; fpr=4BC0D7B24471B2A184EAF5D3000E684119568215 Check that one namespace's NFSD settings and running server stay out of another's, now that the control plane is serialized per namespace. max_blksize_is_per_netns max_block_size is reachable only through a per-netns nfsd filesystem but used to live in a module-wide variable. Read it in a throwaway namespace, change it in a second, read it again in a third; the two reads must agree. No value is hardcoded -- the default is derived from the size of memory, so the test picks a target that differs from whatever this machine reports. *_busy_is_per_netns, pool_stats_readable_with_foreign_server, listener_get_does_not_show_foreign_listeners max_block_size, VERSION_SET and nfsv4leasetime all refuse with -EBUSY once that namespace has a serv. A peer namespace holds a server up while the test pokes its own, which must not be affected. LISTENER_GET must not report the peer's listener. expkey_flush_with_foreign_server writing /proc/net/rpc/nfsd.fh/flush is the one userspace path that reaches nfsd_file_cache_purge(), and so the one that takes the file cache lock with nothing else held. Servers are created with NFSD_A_SERVER_SOCK_USERSPACE_RPCBIND so nothing calls out to rpcbind, and restricted to NFSv4.1 so nfsd_needs_lockd() stays false. Each namespace mounts its own nfsd filesystem on a private tmpfs; the module creates /proc/fs/nfs, not /proc/fs/nfsd, so there is no mountpoint to borrow. Assisted-by: LLM Signed-off-by: Jeff Layton --- tools/testing/selftests/nfsd/.gitignore | 1 + tools/testing/selftests/nfsd/Makefile | 1 + .../testing/selftests/nfsd/nfsd_netns_isolation.c | 468 +++++++++++++++++= ++++ tools/testing/selftests/nfsd/settings | 2 +- 4 files changed, 471 insertions(+), 1 deletion(-) diff --git a/tools/testing/selftests/nfsd/.gitignore b/tools/testing/selfte= sts/nfsd/.gitignore index 19e6dec04d8e..0304b80eb844 100644 --- a/tools/testing/selftests/nfsd/.gitignore +++ b/tools/testing/selftests/nfsd/.gitignore @@ -1 +1,2 @@ nfsd_netlink_listener +nfsd_netns_isolation diff --git a/tools/testing/selftests/nfsd/Makefile b/tools/testing/selftest= s/nfsd/Makefile index 15ac65549d25..2b7c44c3fb00 100644 --- a/tools/testing/selftests/nfsd/Makefile +++ b/tools/testing/selftests/nfsd/Makefile @@ -2,5 +2,6 @@ CFLAGS +=3D $(KHDR_INCLUDES) -Wall =20 TEST_GEN_PROGS :=3D nfsd_netlink_listener +TEST_GEN_PROGS +=3D nfsd_netns_isolation =20 include ../lib.mk diff --git a/tools/testing/selftests/nfsd/nfsd_netns_isolation.c b/tools/te= sting/selftests/nfsd/nfsd_netns_isolation.c new file mode 100644 index 000000000000..b39ea6a908fd --- /dev/null +++ b/tools/testing/selftests/nfsd/nfsd_netns_isolation.c @@ -0,0 +1,468 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Namespace-isolation tests for the NFSD control plane. + * + * NFSD's per-namespace settings used to sit behind one module-wide mutex, + * and one of them -- the maximum READ/WRITE payload -- was a module-wide + * variable reachable through a per-netns file. These tests pin down the + * boundary: what one namespace does to its own server must not be visible + * to, or block, another. + * + * Every namespace here gets a private net + mount namespace, a tmpfs on + * /mnt so nothing escapes, and its own nfsd filesystem mounted on + * /mnt/nfsd. The module creates /proc/fs/nfs, not /proc/fs/nfsd, so the + * mount has to be made by hand; it is also what ties a running server to + * this test, since nfsd_umount() stops the threads. + * + * Servers are created with NFSD_A_SERVER_SOCK_USERSPACE_RPCBIND so nothing + * ever calls out to rpcbind, and restricted to NFSv4.1 so that + * nfsd_needs_lockd() stays false and no lockd instance is started. + */ +#define _GNU_SOURCE +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "../kselftest_harness.h" +#include "nfsd_netlink.h" + +#define NFSD_MNT "/mnt/nfsd" +#define TEST_PORT 20049 + +/* netns_enter() could not build a usable namespace; not a test failure. */ +#define NETNS_NO_SETUP INT_MIN + +/* + * Build a private net + mount namespace with an nfsd filesystem on + * /mnt/nfsd and loopback up. Returns 0, or NETNS_NO_SETUP when the + * environment will not allow it. + */ +static int netns_enter(void) +{ + struct ifreq ifr =3D {0}; + int s; + + if (unshare(CLONE_NEWNET | CLONE_NEWNS) < 0) + return NETNS_NO_SETUP; + if (mount("", "/", NULL, MS_REC | MS_PRIVATE, NULL) < 0) + return NETNS_NO_SETUP; + + /* + * Everything below is created inside this mount namespace only, so + * the mkdir cannot leave anything behind on the host. + */ + if (mount("tmpfs", "/mnt", "tmpfs", 0, NULL) < 0) + return NETNS_NO_SETUP; + if (mkdir(NFSD_MNT, 0755) < 0) + return NETNS_NO_SETUP; + if (mount("nfsd", NFSD_MNT, "nfsd", 0, NULL) < 0) + return NETNS_NO_SETUP; + + s =3D socket(AF_INET, SOCK_DGRAM, 0); + if (s < 0) + return NETNS_NO_SETUP; + strcpy(ifr.ifr_name, "lo"); + if (ioctl(s, SIOCGIFFLAGS, &ifr) < 0) { + close(s); + return NETNS_NO_SETUP; + } + ifr.ifr_flags |=3D IFF_UP | IFF_RUNNING; + if (ioctl(s, SIOCSIFFLAGS, &ifr) < 0) { + close(s); + return NETNS_NO_SETUP; + } + close(s); + + nfsd_family =3D genl_resolve_nfsd(); + if (nfsd_family < 0) + return NETNS_NO_SETUP; + return 0; +} + +/* ------------------- nfsdfs file access ------------------- */ + +static int nfsd_file_read(const char *name, char *buf, size_t len) +{ + char path[128]; + int fd, n; + + snprintf(path, sizeof(path), "%s/%s", NFSD_MNT, name); + fd =3D open(path, O_RDONLY); + if (fd < 0) + return -errno; + n =3D read(fd, buf, len - 1); + close(fd); + if (n < 0) + return -errno; + buf[n] =3D '\0'; + return n; +} + +static int nfsd_file_read_int(const char *name) +{ + char buf[64]; + int n =3D nfsd_file_read(name, buf, sizeof(buf)); + + if (n < 0) + return n; + return atoi(buf); +} + +static int nfsd_file_write_int(const char *name, int val) +{ + char path[128], buf[64]; + int fd, n, len; + + snprintf(path, sizeof(path), "%s/%s", NFSD_MNT, name); + len =3D snprintf(buf, sizeof(buf), "%d\n", val); + fd =3D open(path, O_WRONLY); + if (fd < 0) + return -errno; + n =3D write(fd, buf, len); + close(fd); + return n < 0 ? -errno : 0; +} + +/* ------------------- server lifecycle ------------------- */ + +/* + * Bring up a v4.1-only server on a loopback listener, owning rpcbind + * registration in userspace so the kernel never issues an rpcbind call. + */ +static int server_start(uint16_t port, int nthreads) +{ + char attrs[128]; + int off; + int ret; + + ret =3D version_set_only(4, 1); + if (ret) + return ret; + + off =3D put_listener(attrs, 0, "tcp", port); + off =3D put_attr(attrs, off, NFSD_A_SERVER_SOCK_USERSPACE_RPCBIND, + NULL, 0); + ret =3D listener_set(attrs, off); + if (ret) + return ret; + + return threads_set(nthreads); +} + +static void server_stop(void) +{ + threads_set(0); + listener_set(NULL, 0); +} + +/* ------------------- run a callback in a fresh namespace ---------------= ---- */ + +/* + * Fork a child into its own namespace, run @fn there and hand back what it + * returned. Used for the checks that only need one namespace at a time. + */ +static int netns_run(int (*fn)(long), long arg) +{ + int p[2], ret =3D -EIO; + pid_t pid; + + if (pipe(p) < 0) + return -errno; + + pid =3D fork(); + if (pid < 0) { + close(p[0]); + close(p[1]); + return -errno; + } + if (pid =3D=3D 0) { + int r =3D netns_enter(); + + if (r =3D=3D 0) + r =3D fn(arg); + if (write(p[1], &r, sizeof(r)) !=3D sizeof(r)) + _exit(1); + _exit(0); + } + + close(p[1]); + if (read(p[0], &ret, sizeof(ret)) !=3D sizeof(ret)) + ret =3D -EIO; + close(p[0]); + waitpid(pid, NULL, 0); + return ret; +} + +/* ------------------- a peer namespace held open ------------------- */ + +/* + * A second namespace running a server for as long as the test needs it. + * The peer reports readiness on a pipe and waits for a byte before tearing + * down, so the test can be sure the server is up while it pokes its own + * namespace. + */ +struct peer { + pid_t pid; + int wake; /* write here to let the peer exit */ + int ready; /* peer writes its status here */ +}; + +static int peer_start(struct peer *pr, uint16_t port) +{ + int wake[2], ready[2]; + pid_t ppid =3D getpid(); + char status; + + if (pipe(wake) < 0) + return -errno; + if (pipe(ready) < 0) { + close(wake[0]); + close(wake[1]); + return -errno; + } + + /* peer_stop() handles a dead peer itself; do not die of SIGPIPE first. */ + signal(SIGPIPE, SIG_IGN); + + pr->pid =3D fork(); + if (pr->pid < 0) { + close(wake[0]); close(wake[1]); + close(ready[0]); close(ready[1]); + return -errno; + } + if (pr->pid =3D=3D 0) { + char c; + int r; + + /* Hold no writer of our own, so read() below sees EOF. */ + close(wake[1]); + close(ready[0]); + /* + * A parent that dies without reaching peer_stop() would leave + * this namespace and its server pinned by a process blocked + * forever in read(). + */ + prctl(PR_SET_PDEATHSIG, SIGKILL); + if (getppid() !=3D ppid) + _exit(1); + + r =3D netns_enter(); + if (r =3D=3D 0) + r =3D server_start(port, 1); + status =3D r =3D=3D NETNS_NO_SETUP ? 'S' : (r ? 'E' : 'R'); + if (write(ready[1], &status, 1) !=3D 1) + _exit(1); + /* Hold the namespace open until the test is done with it. */ + if (read(wake[0], &c, 1) =3D=3D 1 && status =3D=3D 'R') + server_stop(); + _exit(0); + } + + close(wake[0]); + close(ready[1]); + pr->wake =3D wake[1]; + pr->ready =3D ready[0]; + + if (read(pr->ready, &status, 1) !=3D 1) + status =3D 'E'; + if (status =3D=3D 'S') + return NETNS_NO_SETUP; + return status =3D=3D 'R' ? 0 : -EIO; +} + +static void peer_stop(struct peer *pr) +{ + char c =3D 'x'; + + if (pr->pid <=3D 0) + return; + if (write(pr->wake, &c, 1) !=3D 1) + kill(pr->pid, SIGKILL); + close(pr->wake); + close(pr->ready); + waitpid(pr->pid, NULL, 0); + pr->pid =3D 0; +} + +/* =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D max_blo= ck_size is per-namespace =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D */ + +static int read_max_blksize(long unused) +{ + (void)unused; + return nfsd_file_read_int("max_block_size"); +} + +static int write_max_blksize(long val) +{ + int ret =3D nfsd_file_write_int("max_block_size", (int)val); + + if (ret) + return ret; + /* Report what stuck, so the caller knows the write was accepted. */ + return nfsd_file_read_int("max_block_size"); +} + +/* + * max_block_size is reachable only through a per-netns nfsd filesystem, so + * a write in one namespace must not be visible in another. Read it in a + * throwaway namespace, change it in a second, then read it again in a + * third: the two reads have to agree. + * + * No value is hardcoded. The default is derived from the size of memory, + * so the test picks a target that differs from whatever this machine uses. + */ +TEST(max_blksize_is_per_netns) +{ + int before, after, wrote, target; + + before =3D netns_run(read_max_blksize, 0); + if (before =3D=3D NETNS_NO_SETUP) + SKIP(return, "cannot set up a private nfsd namespace"); + ASSERT_GE(before, 0); + + /* Any legal value that is not the one this machine already reports. */ + target =3D (before =3D=3D 262144) ? 131072 : 262144; + + wrote =3D netns_run(write_max_blksize, target); + ASSERT_EQ(target, wrote) + TH_LOG("second namespace did not accept max_block_size=3D%d", + target); + + after =3D netns_run(read_max_blksize, 0); + ASSERT_GE(after, 0); + EXPECT_EQ(before, after) + TH_LOG("max_block_size leaked between namespaces: %d -> %d", + before, after); +} + +/* =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D a busy = namespace does not busy others =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D */ + +FIXTURE(nfsd_peer) { + struct peer pr; +}; + +FIXTURE_SETUP(nfsd_peer) +{ + int ret; + + if (geteuid() !=3D 0) + SKIP(return, "must be run as root"); + + /* The peer namespace comes first; it must not be ours. */ + memset(&self->pr, 0, sizeof(self->pr)); + ret =3D peer_start(&self->pr, TEST_PORT); + if (ret =3D=3D NETNS_NO_SETUP) + SKIP(return, "cannot set up a private nfsd namespace"); + if (ret) + SKIP(return, "peer namespace could not start a server: %d", ret); + + /* Now put this process in a namespace of its own, with no server. */ + if (netns_enter() !=3D 0) + SKIP(return, "cannot set up a private nfsd namespace"); +} + +FIXTURE_TEARDOWN(nfsd_peer) +{ + if (nfsd_family >=3D 0) + server_stop(); + peer_stop(&self->pr); +} + +/* + * write_maxblksize() refuses with -EBUSY while that namespace has a serv. + * The check is on nn->nfsd_serv, so a server belonging to someone else mu= st + * not trip it. + */ +TEST_F(nfsd_peer, maxblksize_busy_is_per_netns) +{ + int cur =3D nfsd_file_read_int("max_block_size"); + + ASSERT_GE(cur, 0); + EXPECT_EQ(0, nfsd_file_write_int("max_block_size", + cur =3D=3D 262144 ? 131072 : 262144)) + TH_LOG("max_block_size refused while another netns has a server"); +} + +/* + * NFSD_CMD_VERSION_SET returns -EBUSY once the namespace has a serv. A + * server in the peer namespace must not reach us. + */ +TEST_F(nfsd_peer, version_set_busy_is_per_netns) +{ + EXPECT_EQ(0, version_set_only(4, 1)) + TH_LOG("VERSION_SET refused while another netns has a server"); +} + +/* + * The grace and lease times are per-namespace too, and gated on the same + * nn->nfsd_serv check. + */ +TEST_F(nfsd_peer, leasetime_busy_is_per_netns) +{ + EXPECT_EQ(0, nfsd_file_write_int("nfsv4leasetime", 60)) + TH_LOG("nfsv4leasetime refused while another netns has a server"); +} + +/* + * pool_stats runs its seq_file under the mutex that also guards the serv. + * Reading it here must not be affected by the peer's server, and must + * report this namespace, which has no threads at all. + */ +TEST_F(nfsd_peer, pool_stats_readable_with_foreign_server) +{ + char buf[4096]; + int n =3D nfsd_file_read("pool_stats", buf, sizeof(buf)); + + ASSERT_GE(n, 0) + TH_LOG("pool_stats unreadable: %s", strerror(-n)); + EXPECT_NE(NULL, strstr(buf, "packets-arrived")); +} + +/* + * LISTENER_GET reports this namespace's listeners. With a server running + * next door and none here, the list must come back empty rather than + * showing the peer's. + */ +TEST_F(nfsd_peer, listener_get_does_not_show_foreign_listeners) +{ + char rbuf[8192]; + int n =3D genl_request_reply(NFSD_CMD_LISTENER_GET, rbuf, sizeof(rbuf)); + + ASSERT_GT(n, 0); + EXPECT_EQ(NULL, memmem(rbuf, n, "tcp", 4)) + TH_LOG("LISTENER_GET leaked a listener from another netns"); +} + +/* + * The file cache is one host-wide object, but the flush that reaches it is + * driven from a per-netns file. Doing it here while the peer has a server + * up must be harmless -- this is the path that takes the cache lock with + * nothing else held. + */ +TEST_F(nfsd_peer, expkey_flush_with_foreign_server) +{ + int fd =3D open("/proc/net/rpc/nfsd.fh/flush", O_WRONLY); + + if (fd < 0) + SKIP(return, "no /proc/net/rpc/nfsd.fh/flush: %s", + strerror(errno)); + EXPECT_EQ(2, write(fd, "1\n", 2)); + close(fd); +} + +TEST_HARNESS_MAIN diff --git a/tools/testing/selftests/nfsd/settings b/tools/testing/selftest= s/nfsd/settings index 6091b45d226b..694d70710ff0 100644 --- a/tools/testing/selftests/nfsd/settings +++ b/tools/testing/selftests/nfsd/settings @@ -1 +1 @@ -timeout=3D120 +timeout=3D300 --=20 2.55.0 From nobody Thu Sep 24 16:07:44 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 78FE653CA6F; Tue, 22 Sep 2026 11:34:27 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790076870; cv=none; b=JkYnlh63/50G4Afy6ebMwHWmKvmrgdgnzgPap//dKIl0mLPRX8YUNZOgBtdw41A7QGX2jzJMc1VkgfvIcCJQl0yRhgWhw+USSm9AsMJhoGB6sIGp7tlSmPao1U+Wf/Rmk1p1vhXM2myO95B/C+pa04TMET8GuNhWfcUB0gXZ+FI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790076870; c=relaxed/simple; bh=5bC76uAgFEI9MAOPPthi086OiyQQuDkGEYMVljJe+nk=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=hutBhVmq0lNZt8dY5g/8UCH4PkWiUPjZFIzeLg1CzN4iQalWgOiVFFyD6Rd5b/zrLdMCE3ghoujV51+Bd0Ij+tXY6OotMC3HfXzmAMpf6Ql1znukvORuItcTgI7Fi7mBxmF3kfnbO3UaKD780mp7EeoeGwQoKCcJE56k1QDX5Dg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=LXa+EwF8; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="LXa+EwF8" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 9CD841F00893; Tue, 22 Sep 2026 11:34:26 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1790076867; bh=psN55JlUjKry1U9zL6IGxQltn13iuretpNa5H1Ltl08=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=LXa+EwF8D1i8ylvvtkoiT/iPjsXhoIr5cGof/L0zKIznkP/X/L6Acvrzo8vfnFl7B 1fZLQHx8Di8rVpTiRlqZ8jDcQcVyXcYLISxPH8b1GGCkAH0hP2xMapa9oOEFb6GCkI o38UY09d6loJmwzinXtiflbVK9TY3M4ftEl1lA9/mDkOzHhoTx3RUFmLPGf5x+qZzE 2zmO3YIqth5WqrN20Wn46FPhxwDxNUqU+2azmae3vaVZ9PVJSLks5QmHFx7XFX49mK 4PBI4c/fOCrOFfyfruZULpvFFCbbtQjARlodRtdIHE6xxXqTJbt85f/w8I7F9YzGms O+YDVr0zx5WdQ== From: Jeff Layton Date: Tue, 22 Sep 2026 07:34:09 -0400 Subject: [PATCH v2 8/8] selftests/nfsd: add a cross-namespace control-plane soak Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260922-nfsd-per-net-mutex-v2-8-4a5da8243a73@kernel.org> References: <20260922-nfsd-per-net-mutex-v2-0-4a5da8243a73@kernel.org> In-Reply-To: <20260922-nfsd-per-net-mutex-v2-0-4a5da8243a73@kernel.org> To: Chuck Lever , NeilBrown , Olga Kornievskaia , Dai Ngo , Tom Talpey , Shuah Khan Cc: linux-nfs@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, Jeff Layton X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=openpgp-sha256; l=20610; i=jlayton@kernel.org; h=from:subject:message-id; bh=5bC76uAgFEI9MAOPPthi086OiyQQuDkGEYMVljJe+nk=; b=owEBbQKS/ZANAwAKAQAOaEEZVoIVAcsmYgBqsme6rOrlLITcd4zUW2BUyyzzmANgTrUdAAxIr hUPuoCVrN2JAjMEAAEKAB0WIQRLwNeyRHGyoYTq9dMADmhBGVaCFQUCarJnugAKCRAADmhBGVaC FVKgEADA5NDw63V8xKnFFcJIPAWwAH3KcUoUU+v9eFN1mOL766AjraogHoEbpuKIrXQdFpYS+mo Tsh8DkTtTmLlQqDupz/FFziBbfQ9W2t18TvH/TRkU7EQiZCK73/MCIVWzjkwtL9iaNeC/RCzRz2 /e+jaE0Uabyzl+v9xbhtsRCoYKi3HKe3JNIoW/8UaWp779oxD39GMZ6Zr7h2krdK2COzTOdNZys OyFZ2MOMoKXjhXjGqiQsgQ8tgXaba02A5jqv7SOUlkuVQmn2pRu1XrcZMpcQ48LaK5nVZ8tEZUs JUzSaRLOQWw8BEsoO8hi5IYZBDborMefdAhRGcVHyyPQftGxHX0rpbEXS6gD+O6GVMdcI+UDcK8 qQTHyYJNtBh2v8QXMnUdTJQ4HJ5Xh7ccT+YjjlV7klJGFLXM1qrVZ8NU/n+kuVt9sRh2DOThiUM CERUZCK/Wiubn61ASycz+BYCgT85wmpHwbijKzThWiTlqVR7aHhpnYeW1nSQc7ILSuQ0yISO+or hSLVCnqSclf1sFZe1XusHWlhPBe/cfRfX+skzdfHmelAxgCOSKJdou2xI+fxH4Z/HAKHKFZh1gJ PTWV2ZP0KRnYuUBeLT4N2SHlPMGs8fMGr++A1y1A2N/JaHiHYrquqVlB5j2YtEJ07eKm7jjNL4e av3rl+H6ygntqLA== X-Developer-Key: i=jlayton@kernel.org; a=openpgp; fpr=4BC0D7B24471B2A184EAF5D3000E684119568215 Drive the NFSD control plane from many namespaces at once, to exercise every edge of the lock nesting: nn->nfsd_mutex -> nfsd_global_mutex -> nfsd_file_cache_mutex Each worker owns a namespace and churns between "no serv" and "serv with threads", mixing in the operations that cross into the host-wide locks: start/stop a server nn -> global (nfsd_users 0->1->0, notifiers) -> cache (cache init/shutdown) empty LISTENER_SET creates and destroys a serv in one call, so the host-wide refcount goes 0->1->0 on its own nfsd.fh/flush the cache lock with nothing else held read filecache likewise read pool_stats nn, reached through svc_info.mutex Random churn alone is a weak race finder, so the run ends with synchronized rounds where every namespace attempts the host-wide 0->1 transition at the same instant. That window only opened up once the per-namespace lock stopped serializing namespaces against each other. The test asserts little by itself; the result that matters is that the kernel did not warn. The taint word is sampled before and after and a newly set TAINT_WARN fails the run, which catches lockdep splats, WARN_ON()s and refcount saturation alike. Add PROVE_LOCKING and friends to the config, and log a note if lockdep looks absent. Two things can silently gut the coverage, so both are reported rather than left to look like a pass: - without lockdep there is very little for the kernel to complain about; - a server already running outside these namespaces pins the host-wide refcount above zero for the whole run, so the 0->1 transition the synchronized rounds are built around never happens. Each worker leaves only NFSv4.1 enabled. A fresh namespace has v3 on, which makes nfsd_needs_lockd() true, and the lockd that comes up then waits out an RPC timeout against an rpcbind that is not there on every single server start. Every namespace binds the same port; they are isolated, so that must work, and a bind failure is a louder signal than a silent pass. Sized from nproc, capped at 8 namespaces and 5s of churn by default, and tunable with NFSD_STRESS_WORKERS and NFSD_STRESS_SECS. The barrier waits with a deadline so a worker that dies cannot wedge the run, and the test carries an explicit timeout because the harness otherwise caps it at TEST_TIMEOUT_DEFAULT. Assisted-by: LLM Signed-off-by: Jeff Layton --- tools/testing/selftests/nfsd/.gitignore | 1 + tools/testing/selftests/nfsd/Makefile | 1 + tools/testing/selftests/nfsd/config | 6 + tools/testing/selftests/nfsd/nfsd_netns_stress.c | 572 +++++++++++++++++++= ++++ tools/testing/selftests/nfsd/settings | 2 +- 5 files changed, 581 insertions(+), 1 deletion(-) diff --git a/tools/testing/selftests/nfsd/.gitignore b/tools/testing/selfte= sts/nfsd/.gitignore index 0304b80eb844..2347491c634d 100644 --- a/tools/testing/selftests/nfsd/.gitignore +++ b/tools/testing/selftests/nfsd/.gitignore @@ -1,2 +1,3 @@ nfsd_netlink_listener nfsd_netns_isolation +nfsd_netns_stress diff --git a/tools/testing/selftests/nfsd/Makefile b/tools/testing/selftest= s/nfsd/Makefile index 2b7c44c3fb00..b29bf642c0ad 100644 --- a/tools/testing/selftests/nfsd/Makefile +++ b/tools/testing/selftests/nfsd/Makefile @@ -3,5 +3,6 @@ CFLAGS +=3D $(KHDR_INCLUDES) -Wall =20 TEST_GEN_PROGS :=3D nfsd_netlink_listener TEST_GEN_PROGS +=3D nfsd_netns_isolation +TEST_GEN_PROGS +=3D nfsd_netns_stress =20 include ../lib.mk diff --git a/tools/testing/selftests/nfsd/config b/tools/testing/selftests/= nfsd/config index 0eef03af3503..c6407806b58b 100644 --- a/tools/testing/selftests/nfsd/config +++ b/tools/testing/selftests/nfsd/config @@ -12,3 +12,9 @@ CONFIG_INOTIFY_USER=3Dy CONFIG_SUNRPC=3Dy CONFIG_NFSD=3Dy CONFIG_NFSD_V4=3Dy + +# nfsd_netns_stress leans on lockdep to find anything; without these the +# soak still runs but proves very little. +CONFIG_PROVE_LOCKING=3Dy +CONFIG_DEBUG_MUTEXES=3Dy +CONFIG_DEBUG_ATOMIC_SLEEP=3Dy diff --git a/tools/testing/selftests/nfsd/nfsd_netns_stress.c b/tools/testi= ng/selftests/nfsd/nfsd_netns_stress.c new file mode 100644 index 000000000000..7ca278a5d4a0 --- /dev/null +++ b/tools/testing/selftests/nfsd/nfsd_netns_stress.c @@ -0,0 +1,572 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Concurrency soak for the NFSD control plane across network namespaces. + * + * NFSD's control plane is serialized by three locks, nested in this order: + * + * nn->nfsd_mutex -> nfsd_global_mutex -> nfsd_file_cache_mutex + * + * Most of what used to be one module-wide mutex is now per-namespace, so + * namespaces run their control planes in parallel and only meet on the + * host-wide bits: the refcount that brings the open file cache and the + * NFSv4 global tables up and down, and the address-notifier registration. + * + * This test exists to drive every one of those edges at once from many + * namespaces. It asserts little by itself -- the point is to give lockdep + * something to work with, so it is close to worthless without + * CONFIG_PROVE_LOCKING=3Dy. What it does check is that the kernel did not + * warn: the taint word is sampled before and after, and a newly set + * TAINT_WARN fails the run. That catches lockdep splats, WARN_ON()s and + * refcount saturation alike. + * + * Each worker drives these, which between them cover every edge: + * + * start a server nn -> global (nfsd_users 0->1, notifiers 0->1) + * -> cache (nfsd_file_cache_init) + * stop a server nn -> global (1->0) -> cache (cache_shutdown) + * empty LISTENER_SET creates and destroys a serv in one call, so the + * host-wide refcount goes 0->1->0 on its own + * nfsd.fh/flush the cache lock with nothing else held + * read filecache likewise + * read pool_stats nn, reached through svc_info.mutex + * + * Random churn alone is a weak race finder, so the run ends with + * synchronized rounds where every worker attempts the host-wide 0->1 + * transition at the same instant. That is the window that only opened up + * once the per-namespace lock stopped serializing namespaces against each + * other. + * + * Tunable with NFSD_STRESS_WORKERS and NFSD_STRESS_SECS. + */ +#define _GNU_SOURCE +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "../kselftest_harness.h" +#include "nfsd_netlink.h" + +#define NFSD_MNT "/mnt/nfsd" +#define EXPKEY_FLUSH "/proc/net/rpc/nfsd.fh/flush" + +/* + * Every worker binds the same port. They are in different namespaces, so + * that has to work; if isolation ever breaks, the second bind fails loudly + * rather than the test quietly passing. + */ +#define STRESS_PORT 20049 + +#define DEFAULT_WORKERS 8 +#define DEFAULT_SECS 5 +#define MAX_WORKERS 64 +#define SYNC_ROUNDS 20 +#define BARRIER_TIMEOUT_SEC 30 + +/* + * Generous: the harness caps a test at TEST_TIMEOUT_DEFAULT otherwise, and + * the churn duration is tunable. + */ +#define SOAK_TIMEOUT_SEC 600 + +#define TAINT_WARN_BIT (1UL << 9) + +/* Worker exit codes. */ +#define WORKER_OK 0 +#define WORKER_FAIL 1 +#define WORKER_NO_SETUP 2 + +/* ------------------- cross-process barrier ------------------- */ + +struct barrier { + unsigned int n; + unsigned int count; + unsigned int generation; + unsigned int aborted; /* latched: a rendezvous was never completed */ +}; + +/* + * Spin with a deadline rather than blocking, so a worker that dies cannot + * wedge the run. + * + * The first waiter to give up latches ->aborted, which kills the barrier + * for good: every later call returns at once instead of waiting out its + * own deadline. Without that latch a single missing worker costs + * BARRIER_TIMEOUT_SEC on every remaining rendezvous -- three per round, + * SYNC_ROUNDS rounds -- which runs into tens of minutes before the + * harness timeout fires. + * + * A giving-up waiter leaves its ->count increment behind. That is fine: + * once ->aborted is set nothing reads ->count again. + * + * Returns false if the rendezvous did not happen. + */ +static bool barrier_wait(struct barrier *b) +{ + unsigned int gen =3D __atomic_load_n(&b->generation, __ATOMIC_ACQUIRE); + time_t deadline =3D time(NULL) + BARRIER_TIMEOUT_SEC; + + if (__atomic_load_n(&b->aborted, __ATOMIC_ACQUIRE)) + return false; + + if (__atomic_add_fetch(&b->count, 1, __ATOMIC_ACQ_REL) =3D=3D b->n) { + __atomic_store_n(&b->count, 0, __ATOMIC_RELEASE); + __atomic_add_fetch(&b->generation, 1, __ATOMIC_ACQ_REL); + return true; + } + while (__atomic_load_n(&b->generation, __ATOMIC_ACQUIRE) =3D=3D gen) { + if (__atomic_load_n(&b->aborted, __ATOMIC_ACQUIRE)) + return false; + if (time(NULL) > deadline) { + __atomic_store_n(&b->aborted, 1, __ATOMIC_RELEASE); + return false; + } + sched_yield(); + } + return true; +} + +/* ------------------- namespace setup ------------------- */ + +static int netns_enter(void) +{ + struct ifreq ifr =3D {0}; + int s; + + if (unshare(CLONE_NEWNET | CLONE_NEWNS) < 0) + return -errno; + if (mount("", "/", NULL, MS_REC | MS_PRIVATE, NULL) < 0) + return -errno; + if (mount("tmpfs", "/mnt", "tmpfs", 0, NULL) < 0) + return -errno; + if (mkdir(NFSD_MNT, 0755) < 0) + return -errno; + if (mount("nfsd", NFSD_MNT, "nfsd", 0, NULL) < 0) + return -errno; + + s =3D socket(AF_INET, SOCK_DGRAM, 0); + if (s < 0) + return -errno; + strcpy(ifr.ifr_name, "lo"); + if (ioctl(s, SIOCGIFFLAGS, &ifr) < 0) { + close(s); + return -errno; + } + ifr.ifr_flags |=3D IFF_UP | IFF_RUNNING; + if (ioctl(s, SIOCSIFFLAGS, &ifr) < 0) { + close(s); + return -errno; + } + close(s); + + nfsd_family =3D genl_resolve_nfsd(); + if (nfsd_family < 0) + return -ENOENT; + return 0; +} + +/* ------------------- the operations ------------------- */ + +static int read_nfsd_file(const char *name) +{ + char path[128], buf[4096]; + int fd, n; + + snprintf(path, sizeof(path), "%s/%s", NFSD_MNT, name); + fd =3D open(path, O_RDONLY); + if (fd < 0) + return -errno; + do { + n =3D read(fd, buf, sizeof(buf)); + } while (n > 0); + close(fd); + return n < 0 ? -errno : 0; +} + +/* + * The only userspace trigger that reaches nfsd_file_cache_purge(), and so + * the only one that takes the file cache lock with no other nfsd lock + * held. NFSD_CMD_CACHE_FLUSH does not get there: cache_purge() never calls + * the cache_detail's ->flush hook. + */ +static int expkey_flush(void) +{ + int fd =3D open(EXPKEY_FLUSH, O_WRONLY); + int n; + + if (fd < 0) + return -errno; + n =3D write(fd, "1\n", 2); + close(fd); + return n < 0 ? -errno : 0; +} + +/* Create a serv and tear it straight back down inside one call. */ +static int serv_cycle(void) +{ + return listener_set(NULL, 0); +} + +static int server_up(int nthreads) +{ + char attrs[128]; + int off, ret; + + off =3D put_listener(attrs, 0, "tcp", STRESS_PORT); + off =3D put_attr(attrs, off, NFSD_A_SERVER_SOCK_USERSPACE_RPCBIND, + NULL, 0); + ret =3D listener_set(attrs, off); + if (ret) + return ret; + return threads_set(nthreads); +} + +/* threads_set(0) drops the last thread, which destroys the serv with it. = */ +static int server_down(void) +{ + return threads_set(0); +} + +/* ------------------- the worker ------------------- */ + +struct worker_err { + const char *op; + int err; +}; + +static struct worker_err worker_fault; + +static bool fail(const char *op, int err) +{ + if (err =3D=3D 0) + return false; + worker_fault.op =3D op; + worker_fault.err =3D err; + return true; +} + +/* + * Random churn between "no serv" and "serv with threads", with the + * lock-crossing reads and flushes mixed in at both ends. Every call here + * is one that must succeed in the state it is issued from, so any error is + * a real failure rather than an expected race. + */ +static bool worker_churn(time_t deadline) +{ + bool up =3D false; + + while (time(NULL) < deadline) { + int r =3D random() % 8; + + if (!up) { + switch (r) { + case 0: + case 1: + if (fail("serv_cycle", serv_cycle())) + return false; + break; + case 2: + if (fail("version_set", version_set_only(4, 1))) + return false; + break; + case 3: + if (fail("flush", expkey_flush())) + return false; + break; + case 4: + if (fail("filecache", read_nfsd_file("filecache"))) + return false; + break; + case 5: + if (fail("pool_stats", read_nfsd_file("pool_stats"))) + return false; + break; + default: + if (fail("server_up", server_up(1 + random() % 3))) + return false; + up =3D true; + break; + } + } else { + switch (r) { + case 0: + if (fail("flush", expkey_flush())) + return false; + break; + case 1: + if (fail("filecache", read_nfsd_file("filecache"))) + return false; + break; + case 2: + if (fail("pool_stats", read_nfsd_file("pool_stats"))) + return false; + break; + case 3: + if (fail("threads_set", threads_set(1 + random() % 3))) + return false; + break; + default: + if (fail("server_down", server_down())) + return false; + up =3D false; + break; + } + } + } + + if (up && fail("server_down", server_down())) + return false; + return true; +} + +/* + * Every worker arrives at the barrier with no serv, so the host-wide + * refcount is at zero, then they all try to take it to one together. The + * second barrier lines up the drop back to zero the same way. + */ +static bool worker_sync_rounds(struct barrier *b) +{ + int i; + + for (i =3D 0; i < SYNC_ROUNDS; i++) { + /* + * A failed rendezvous means a peer is gone, which its own + * exit status already reports. Stop the rounds rather than + * stalling on every remaining barrier. + */ + if (!barrier_wait(b)) + return true; + if (fail("sync server_up", server_up(1))) + return false; + + if (!barrier_wait(b)) + return true; + if (fail("sync flush", expkey_flush())) + return false; + + if (!barrier_wait(b)) + return true; + if (fail("sync server_down", server_down())) + return false; + } + return true; +} + +static int worker(int idx, struct barrier *b, unsigned int secs) +{ + int ret =3D netns_enter(); + + if (ret) { + /* + * Report setup trouble rather than a failure: a restricted + * environment is not a kernel bug. + */ + fprintf(stderr, "netns %d: setup: %s\n", idx, strerror(-ret)); + return WORKER_NO_SETUP; + } + + /* + * Leave only NFSv4.1 enabled. A fresh namespace has v3 on, which + * makes nfsd_needs_lockd() true, and the lockd that comes up then + * tries to reach an rpcbind that is not there -- so every server + * start waits out an RPC timeout and floods the log. The version + * set sticks across serv teardown, so once is enough. + */ + ret =3D version_set_only(4, 1); + if (ret) { + fprintf(stderr, "netns %d: version_set: %s\n", idx, + strerror(-ret)); + return WORKER_FAIL; + } + + srandom(getpid() ^ (unsigned int)time(NULL)); + + if (!worker_churn(time(NULL) + secs)) + goto fault; + if (!worker_sync_rounds(b)) + goto fault; + return WORKER_OK; + +fault: + fprintf(stderr, "netns %d: %s: %s\n", idx, worker_fault.op, + strerror(-worker_fault.err)); + server_down(); + return WORKER_FAIL; +} + +/* + * Threads running outside the test's namespaces. Any at all pin the + * host-wide refcount above zero for the whole run, so the 0->1 transition + * the synchronized rounds are built around never happens. + */ +static int nfsd_threads_here(void) +{ + char buf[32] =3D ""; + int fd =3D open("/proc/fs/nfsd/threads", O_RDONLY); + int n =3D 0; + + if (fd < 0) + return 0; + if (read(fd, buf, sizeof(buf) - 1) > 0) + n =3D atoi(buf); + close(fd); + return n; +} + +/* ------------------- taint ------------------- */ + +static unsigned long read_taint(void) +{ + unsigned long v =3D 0; + FILE *f =3D fopen("/proc/sys/kernel/tainted", "r"); + + if (!f) + return 0; + if (fscanf(f, "%lu", &v) !=3D 1) + v =3D 0; + fclose(f); + return v; +} + +/* ------------------- the test ------------------- */ + +static unsigned int env_uint(const char *name, unsigned int def, unsigned = int max) +{ + const char *s =3D getenv(name); + unsigned long v; + + if (!s || !*s) + return def; + v =3D strtoul(s, NULL, 0); + if (v =3D=3D 0 || v > max) + return def; + return (unsigned int)v; +} + +FIXTURE(soak) { + int unused; +}; + +FIXTURE_SETUP(soak) { } +FIXTURE_TEARDOWN(soak) { } + +TEST_F_TIMEOUT(soak, netns_control_plane_soak, SOAK_TIMEOUT_SEC) +{ + unsigned long taint_before, taint_after; + unsigned int workers, secs; + pid_t pid[MAX_WORKERS]; + int failed =3D 0, skipped =3D 0; + unsigned int aborted; + struct barrier *b; + unsigned int i; + long ncpu; + + if (geteuid() !=3D 0) + SKIP(return, "must be run as root"); + + ncpu =3D sysconf(_SC_NPROCESSORS_ONLN); + if (ncpu < 2) + ncpu =3D 2; + workers =3D env_uint("NFSD_STRESS_WORKERS", + ncpu < DEFAULT_WORKERS ? (unsigned int)ncpu + : DEFAULT_WORKERS, + MAX_WORKERS); + secs =3D env_uint("NFSD_STRESS_SECS", DEFAULT_SECS, 3600); + + b =3D mmap(NULL, sizeof(*b), PROT_READ | PROT_WRITE, + MAP_SHARED | MAP_ANONYMOUS, -1, 0); + ASSERT_NE(MAP_FAILED, b); + b->n =3D workers; + b->count =3D 0; + b->generation =3D 0; + b->aborted =3D 0; + + TH_LOG("%u namespaces, %us of churn then %d synchronized rounds", + workers, secs, SYNC_ROUNDS); + + /* + * A server already running outside these namespaces holds the + * host-wide refcount above zero for the whole run, so the 0->1 + * transition the synchronized rounds are built around never + * happens. Worth saying out loud rather than reporting coverage + * that was not there. + */ + if (nfsd_threads_here() > 0) + TH_LOG("note: nfsd runs outside these namespaces; refcount never hits 0"= ); + + taint_before =3D read_taint(); + + for (i =3D 0; i < workers; i++) { + pid[i] =3D fork(); + ASSERT_GE(pid[i], 0); + if (pid[i] =3D=3D 0) + _exit(worker(i, b, secs)); + } + + for (i =3D 0; i < workers; i++) { + int status =3D 0; + + waitpid(pid[i], &status, 0); + if (!WIFEXITED(status)) { + failed++; + continue; + } + if (WEXITSTATUS(status) =3D=3D WORKER_FAIL) + failed++; + else if (WEXITSTATUS(status) =3D=3D WORKER_NO_SETUP) + skipped++; + } + + taint_after =3D read_taint(); + aborted =3D __atomic_load_n(&b->aborted, __ATOMIC_ACQUIRE); + munmap(b, sizeof(*b)); + + if (skipped =3D=3D (int)workers) + SKIP(return, "no worker could set up a private nfsd namespace"); + + EXPECT_EQ(0, failed) + TH_LOG("%d of %u namespaces reported an error", failed, workers); + + /* + * A namespace that never reached the rounds -- one that could not + * set up, say -- leaves the rest with nobody to meet. Say so: the + * synchronized part did not run to completion. + */ + if (aborted) + TH_LOG("note: synchronized rounds cut short; a namespace did not arrive"= ); + + /* + * The real result. Without CONFIG_PROVE_LOCKING there is very little + * here for the kernel to complain about, so say so rather than + * letting a quiet pass look like coverage. + */ + EXPECT_EQ(0, (taint_after & ~taint_before) & TAINT_WARN_BIT) + TH_LOG("kernel warned during the run (taint %#lx -> %#lx); check dmesg", + taint_before, taint_after); + + /* + * lockdep_proc_init() puts these in procfs, not debugfs, and + * lockdep_chains is the one that appears only with + * CONFIG_PROVE_LOCKING -- the part that validates ordering rather + * than merely tracking. Root-only, but so is this test. + */ + if (access("/proc/lockdep_chains", R_OK) !=3D 0) + TH_LOG("note: CONFIG_PROVE_LOCKING looks absent; this proves little"); +} + +TEST_HARNESS_MAIN diff --git a/tools/testing/selftests/nfsd/settings b/tools/testing/selftest= s/nfsd/settings index 694d70710ff0..a62d2fa1275c 100644 --- a/tools/testing/selftests/nfsd/settings +++ b/tools/testing/selftests/nfsd/settings @@ -1 +1 @@ -timeout=3D300 +timeout=3D600 --=20 2.55.0