From nobody Mon Sep 28 08:02:48 2026 Received: from mail-pj1-f99.google.com (mail-pj1-f99.google.com [209.85.216.99]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 89B4D3955D2 for ; Mon, 24 Aug 2026 22:30:01 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.99 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787610603; cv=none; b=aIdpni2gvMOZSxpDVIO8ImoML7rbjN4mXdQf4mUxTQ4zKPrPc2hZwpTblObqcQJ6bFFDdYdVqP8m9lAH6hyOohYWExJQFPJfv3cIo1YzYzs87ZsqNBTEM9s3+PWjtCmg9qkOdxikZMk+iQig6e31D9f/fYULoUyWgMuE8MsOTj0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787610603; c=relaxed/simple; bh=re8SEzXHGgkELbjWHAXVAI7FdqAyApz4N7nDpXnd66A=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=iqANhtqY740yZo6tfj7T6tmBMFBgofg2M+cJx4AHOKH3hKi4R+0YXOnSKvDfvQYj6VfM166EDDf27Y0xI9/8TzRCTSJlOBEt3NxEF+s7aNOzwd/MWP+fI1El7Hn6biRTP0++5bHis0diNIiek0t4q9cH826oZexq1iTxgwk0XNY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=purestorage.com; spf=pass smtp.mailfrom=purestorage.com; dkim=pass (2048-bit key) header.d=purestorage.com header.i=@purestorage.com header.b=MsF+XxJt; arc=none smtp.client-ip=209.85.216.99 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=purestorage.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=purestorage.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=purestorage.com header.i=@purestorage.com header.b="MsF+XxJt" Received: by mail-pj1-f99.google.com with SMTP id 98e67ed59e1d1-3964e76d0f4so60189a91.3 for ; Mon, 24 Aug 2026 15:30:01 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=purestorage.com; s=google2022; t=1787610601; x=1788215401; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=6Flt4AJoH53V7PC18cjrTqknEHBQul3chjMHlAVE9x8=; b=MsF+XxJtx6uthnVY1WizlW81ya+Qth4qi+VhKc5gaDKH9KCx299A2GhBoV/bmfhep3 8nQbenECDpCBgBz+TtxNkEMXkI8CySr7+gpnzZLw9O0I0wwz+sGerGVtcIqZyM83bUWF CrWzt1TBwQHseO/ysYe78+OUJxGn/PU/4b1LWu56kWOqzaaxCQtfEVcT61jY3m50LTM6 ari7pdvKJwzaU4qBf281+LIhE2khJjV44LmlA9BnU+t6coELoX46K2ACn+BiNolURhzg IqXk/2cKJ3pN2dBI1hhPlCQvEc8djtZ0/HkCADbdQDpGKb/dgpuN0Cb4SOGHMm2zzCUf GCmA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787610601; x=1788215401; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=6Flt4AJoH53V7PC18cjrTqknEHBQul3chjMHlAVE9x8=; b=isKEehrA/i4kgR/Q5vhUKnXoujqK4Bsw0sIYwifzXMGXVguhK1MlfC5zGPBfAcgLaI U6GiOMVzv49BLFgycUEsPVjc3BwQwx7S2yaX4Dnnx6xLe+ygZbvUhOvKEfCEo0tFr2SR f8zdpXhPrhSZ++MONEAwdYAU+99bhqnfpaZ6BqurhqofCSBeyP7pAOkuxDa2IF8ILQab k4SOo6ZwkwBtoQuGMva6475XFxANj91tkzCExIYTIhUVBZBYA8ZFnasmDbbKNkpJKtab 3u1M3iyqJ5WFrxW9WrO8aSG4M8Yhpt17nDg6BQwBmUuMuef3iIPS+3dzriztect2hMXR 0UeA== X-Forwarded-Encrypted: i=1; AHgh+RoRwGzK38Q4Y6FdC8elKekT/RtHCAP6pGoNCWY64WXVDoDCAO/Pa3frwrXZmRbkgCuBNZR3Jwfm9iF4x94=@vger.kernel.org X-Gm-Message-State: AFuF++k8sq6OQa+3sMSzpE8L2u7hQohq+Lh3ju2q/R8MxJUCCrlKOwQC QaeLH6mYNa3Fh+cvsfsIxW9c+ZLS+igNwcqBuXYL1I6jJr0VmW2T69DklXI2Ahx3aSPwNYVsnMt CMPM9lE36OwkLqS9vCjAq38nqrK2xhSs7fe5d40VDqeLLCbC+f65x X-Gm-Gg: AR+sD12YL4Wb2Y/YONVOWcsBSwvjfVrwM3PyDo1QrB4xQmcsmq4uqzqRMQxZkIlaKik fvYGAFTwnD2EltrjkOJz1qfYKOl4BenSeCvsPehyO9/R6eNagbykdA4JoZCU44ZgCLXFmBfoL7c shxzrIZ1CwNj0ITWyHpY4JgqyJ0Wo8/ZuXpq1JzWYAM7WuQnu7ZpHViEIlUtWQ0cNEyxviGmFZA KUb1PHjGuwvW5g/DlCHB5+E7PC5+oQZcDIfaLpW6piwrvffhGNGBn0EXKV+imhUhNX+lMavQVLV rWMkrrhZRvxGG5fRygG0+cgDk39962JPc+ogKKzKmy3uL664zTtg7E1FjORvX815YdxgDXUD68X lDM05SzGsEeh7YwcH X-Received: by 2002:a17:90a:e70f:b0:381:a766:efcc with SMTP id 98e67ed59e1d1-395df3c6026mr42310272a91.14.1787610600667; Mon, 24 Aug 2026 15:30:00 -0700 (PDT) Received: from c7-smtp-2026.dev.purestorage.com ([2620:125:9017:12:36:3:6:0]) by smtp-relay.gmail.com with ESMTPS id 98e67ed59e1d1-39645abf830sm386992a91.6.2026.08.24.15.30.00 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 24 Aug 2026 15:30:00 -0700 (PDT) X-Relaying-Domain: purestorage.com Received: from dev-sgogte.dev.purestorage.com (bond0.slc5-n22m24-k8s.dev.purestorage.com [IPv6:2620:125:9025:20::a31:429]) by c7-smtp-2026.dev.purestorage.com (Postfix) with ESMTP id F023D402B2; Mon, 24 Aug 2026 16:29:59 -0600 (MDT) Received: by dev-sgogte.dev.purestorage.com (Postfix, from userid 1557734945) id ED70D51D15; Mon, 24 Aug 2026 16:29:59 -0600 (MDT) From: Surabhi Gogte To: Keith Busch , Jens Axboe , Christoph Hellwig , Sagi Grimberg Cc: linux-nvme@lists.infradead.org, linux-kernel@vger.kernel.org, mkhalfella@purestorage.com, randyj@purestorage.com, adailey@purestorage.com, Surabhi Gogte Subject: [PATCH] nvme-tcp: parallelize I/O queue allocation and startup Date: Mon, 24 Aug 2026 16:29:39 -0600 Message-ID: <20260824222939.301887-3-sgogte@purestorage.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260824222939.301887-1-sgogte@purestorage.com> References: <20260824222939.301887-1-sgogte@purestorage.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Similar to commit 2a8513091d2f ("nvme-rdma: parallelize I/O queue allocation and startup"), refactor nvme tcp I/O queue setup to use async API, combining allocation and startup into a single parallel operation per queue. This reduces connection and reconnection setup time when there are delays in establishing connections, which is especially important for high-core-count hosts. Key changes: - Use async API to facilitate parallel calls for io queue setup. - Add nvme_tcp_setup_ctx for propagating errors from async workers. - Remove nvme_tcp_start_io_queues() and __nvme_tcp_alloc_io_queues(); their logic is folded into nvme_tcp_setup_io_queues() and nvme_tcp_configure_io_queues(). - Allocate the io tag set before the queues so that the queue range is known, and only set up the reconnect grow case if the queue count actually increased. - Serialize the cpu scan and claim in nvme_tcp_set_queue_io_cpu() with a spinlock, as concurrent callers would otherwise select the same cpu. The per-cpu counters no longer need to be atomics. - Use init_net in nvme_tcp_alloc_queue() instead of the namespace of current, which is no longer the connecting task once the allocation runs from a worker. A controller is not guaranteed to be tied to a namespace, as the reconnect and error recovery paths already run from a workqueue in init_net. Testing on a 64-core host with 64 IO-queues shows nvme-tcp connection time reduced from 61ms to 11ms. Signed-off-by: Surabhi Gogte --- drivers/nvme/host/tcp.c | 126 +++++++++++++++++++++++++--------------- 1 file changed, 80 insertions(+), 46 deletions(-) diff --git a/drivers/nvme/host/tcp.c b/drivers/nvme/host/tcp.c index 354668ad29ac..30fe4c5abe0b 100644 --- a/drivers/nvme/host/tcp.c +++ b/drivers/nvme/host/tcp.c @@ -7,6 +7,7 @@ #include #include #include +#include #include #include #include @@ -54,7 +55,8 @@ MODULE_PARM_DESC(tls_handshake_timeout, "nvme TLS handshake timeout in seconds (default 10)"); #endif =20 -static atomic_t nvme_tcp_cpu_queues[NR_CPUS]; +static int nvme_tcp_cpu_queues[NR_CPUS]; +static DEFINE_SPINLOCK(nvme_tcp_cpu_queues_lock); =20 enum nvme_tcp_send_state { NVME_TCP_SEND_CMD_PDU =3D 0, @@ -154,6 +156,12 @@ struct nvme_tcp_queue { static DEFINE_MUTEX(nvme_tcp_ctrl_mutex); static LIST_HEAD_GUARDED(nvme_tcp_ctrl_list, nvme_tcp_ctrl_mutex); =20 +struct nvme_tcp_setup_ctx { + struct nvme_ctrl *ctrl; + int qid; + int *err; +}; + struct nvme_tcp_ctrl { /* read only in the hot path */ struct nvme_tcp_queue *queues; @@ -1718,9 +1726,10 @@ static void nvme_tcp_set_queue_io_cpu(struct nvme_tc= p_queue *queue) goto out; =20 /* Search for the least used cpu from the mq_map */ + spin_lock(&nvme_tcp_cpu_queues_lock); io_cpu =3D WORK_CPU_UNBOUND; for_each_online_cpu(cpu) { - int num_queues =3D atomic_read(&nvme_tcp_cpu_queues[cpu]); + int num_queues =3D nvme_tcp_cpu_queues[cpu]; =20 if (mq_map[cpu] !=3D qid) continue; @@ -1731,9 +1740,10 @@ static void nvme_tcp_set_queue_io_cpu(struct nvme_tc= p_queue *queue) } if (io_cpu !=3D WORK_CPU_UNBOUND) { queue->io_cpu =3D io_cpu; - atomic_inc(&nvme_tcp_cpu_queues[io_cpu]); + nvme_tcp_cpu_queues[io_cpu]++; set_bit(NVME_TCP_Q_IO_CPU_SET, &queue->flags); } + spin_unlock(&nvme_tcp_cpu_queues_lock); out: dev_dbg(ctrl->ctrl.device, "queue %d: using cpu %d\n", qid, queue->io_cpu); @@ -1846,7 +1856,7 @@ static int nvme_tcp_alloc_queue(struct nvme_ctrl *nct= rl, int qid, queue->cmnd_capsule_len =3D sizeof(struct nvme_command) + NVME_TCP_ADMIN_CCSZ; =20 - ret =3D sock_create_kern(current->nsproxy->net_ns, + ret =3D sock_create_kern(&init_net, ctrl->addr.ss_family, SOCK_STREAM, IPPROTO_TCP, &queue->sock); if (ret) { @@ -2010,8 +2020,11 @@ static void nvme_tcp_stop_queue_nowait(struct nvme_c= trl *nctrl, int qid) if (!test_bit(NVME_TCP_Q_ALLOCATED, &queue->flags)) return; =20 - if (test_and_clear_bit(NVME_TCP_Q_IO_CPU_SET, &queue->flags)) - atomic_dec(&nvme_tcp_cpu_queues[queue->io_cpu]); + if (test_and_clear_bit(NVME_TCP_Q_IO_CPU_SET, &queue->flags)) { + spin_lock(&nvme_tcp_cpu_queues_lock); + nvme_tcp_cpu_queues[queue->io_cpu]--; + spin_unlock(&nvme_tcp_cpu_queues_lock); + } =20 mutex_lock(&queue->queue_lock); if (test_and_clear_bit(NVME_TCP_Q_LIVE, &queue->flags)) @@ -2118,25 +2131,6 @@ static void nvme_tcp_stop_io_queues(struct nvme_ctrl= *ctrl) nvme_tcp_wait_queue(ctrl, i); } =20 -static int nvme_tcp_start_io_queues(struct nvme_ctrl *ctrl, - int first, int last) -{ - int i, ret; - - for (i =3D first; i < last; i++) { - ret =3D nvme_tcp_start_queue(ctrl, i); - if (ret) - goto out_stop_queues; - } - - return 0; - -out_stop_queues: - for (i--; i >=3D first; i--) - nvme_tcp_stop_queue(ctrl, i); - return ret; -} - static int nvme_tcp_alloc_admin_queue(struct nvme_ctrl *ctrl) { int ret; @@ -2197,22 +2191,64 @@ static int nvme_tcp_tls_check_psk(struct nvme_ctrl = *ctrl) return 0; } =20 -static int __nvme_tcp_alloc_io_queues(struct nvme_ctrl *ctrl) +static void nvme_tcp_setup_queue_async(void *data, async_cookie_t cookie) { - int i, ret; + struct nvme_tcp_setup_ctx *ctx =3D data; + struct nvme_ctrl *ctrl =3D ctx->ctrl; + int ret; =20 - for (i =3D 1; i < ctrl->queue_count; i++) { - ret =3D nvme_tcp_alloc_queue(ctrl, i, - ctrl->tls_pskid); - if (ret) - goto out_free_queues; + ret =3D nvme_tcp_alloc_queue(ctrl, ctx->qid, ctrl->tls_pskid); + if (ret) + goto out_err; + + ret =3D nvme_tcp_start_queue(ctrl, ctx->qid); + if (ret) + goto out_err; + + return; + +out_err: + WRITE_ONCE(*ctx->err, ret); +} + +static int nvme_tcp_setup_io_queues(struct nvme_ctrl *ctrl, unsigned int f= irst, + unsigned int last) +{ + ASYNC_DOMAIN_EXCLUSIVE(queue_domain); + struct nvme_tcp_setup_ctx *ctxs; + int nr_queues =3D last - first; + int err =3D 0, i, ret; + + ctxs =3D kmalloc_objs(*ctxs, nr_queues); + if (!ctxs) + return -ENOMEM; + + for (i =3D 0; i < nr_queues; i++) { + ctxs[i].ctrl =3D ctrl; + ctxs[i].qid =3D first + i; + ctxs[i].err =3D &err; + async_schedule_domain(nvme_tcp_setup_queue_async, &ctxs[i], + &queue_domain); } =20 + async_synchronize_full_domain(&queue_domain); + kfree(ctxs); + + ret =3D READ_ONCE(err); + if (ret) + goto out_free_queues; + return 0; =20 out_free_queues: - for (i--; i >=3D 1; i--) - nvme_tcp_free_queue(ctrl, i); + for (i =3D first; i < last; i++) { + struct nvme_tcp_queue *queue =3D &to_tcp_ctrl(ctrl)->queues[i]; + + if (test_bit(NVME_TCP_Q_LIVE, &queue->flags)) + nvme_tcp_stop_queue(ctrl, i); + if (test_bit(NVME_TCP_Q_ALLOCATED, &queue->flags)) + nvme_tcp_free_queue(ctrl, i); + } =20 return ret; } @@ -2254,10 +2290,6 @@ static int nvme_tcp_configure_io_queues(struct nvme_= ctrl *ctrl, bool new) if (ret) return ret; =20 - ret =3D __nvme_tcp_alloc_io_queues(ctrl); - if (ret) - return ret; - if (new) { ret =3D nvme_alloc_io_tag_set(ctrl, &to_tcp_ctrl(ctrl)->tag_set, &nvme_tcp_mq_ops, @@ -2268,12 +2300,12 @@ static int nvme_tcp_configure_io_queues(struct nvme= _ctrl *ctrl, bool new) } =20 /* - * Only start IO queues for which we have allocated the tagset + * Only setup IO queues for which we have allocated the tagset * and limited it to the available queues. On reconnects, the * queue number might have changed. */ nr_queues =3D min(ctrl->tagset->nr_hw_queues + 1, ctrl->queue_count); - ret =3D nvme_tcp_start_io_queues(ctrl, 1, nr_queues); + ret =3D nvme_tcp_setup_io_queues(ctrl, 1, nr_queues); if (ret) goto out_cleanup_connect_q; =20 @@ -2297,12 +2329,14 @@ static int nvme_tcp_configure_io_queues(struct nvme= _ctrl *ctrl, bool new) =20 /* * If the number of queues has increased (reconnect case) - * start all new queues now. + * setup all new queues now. */ - ret =3D nvme_tcp_start_io_queues(ctrl, nr_queues, - ctrl->tagset->nr_hw_queues + 1); - if (ret) - goto out_wait_freeze_timed_out; + if (ctrl->tagset->nr_hw_queues + 1 > nr_queues) { + ret =3D nvme_tcp_setup_io_queues(ctrl, nr_queues, + ctrl->tagset->nr_hw_queues + 1); + if (ret) + goto out_wait_freeze_timed_out; + } =20 return 0; =20 @@ -3140,7 +3174,7 @@ static int __init nvme_tcp_init_module(void) return -ENOMEM; =20 for_each_possible_cpu(cpu) - atomic_set(&nvme_tcp_cpu_queues[cpu], 0); + nvme_tcp_cpu_queues[cpu] =3D 0; =20 nvmf_register_transport(&nvme_tcp_transport); return 0; --=20 2.55.0