From nobody Fri Oct 2 02:32:12 2026 Received: from mail-pl1-f180.google.com (mail-pl1-f180.google.com [209.85.214.180]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8C1881DDA18 for ; Thu, 20 Aug 2026 08:36:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.180 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787215022; cv=none; b=omBTO6om8F8XlmULbYj3DMNzfZhFcuPpXje1hYMLtICDQT5SOMS8b20jkG/yRkfuKfzfYaYdCFAea1wKLAEK8Il6u2rjTjGXd3kQncyO0tyFnoNDBcgYv9eSYP0tPZ0HZL/PqlAO9Ow7b1ezDR9Ig6UWoluQtuly1/6MD20ewwQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787215022; c=relaxed/simple; bh=rTlrMtJqyhnV3ZS+m2IKItuBt472M1qIVRrq+KZ+mXI=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=USmoraWy+Rti16P01mbvH9Q72fWrKwpF7sN4du/GP9jmoePAU1pEPbKwbLnoUN60fgmsQAzpIICk4ysKdBPaUf7utryYj1nyR0WxhyoB73PJcOSjW69hv0hgECseF6rInUA0OQdfylLdv1AEhFlLL9VAoqFvGktgb8UssVVYGYw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=crusoe.ai; spf=pass smtp.mailfrom=crusoe.ai; dkim=pass (2048-bit key) header.d=crusoe.ai header.i=@crusoe.ai header.b=AAd6JV5y; arc=none smtp.client-ip=209.85.214.180 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=crusoe.ai Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=crusoe.ai Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=crusoe.ai header.i=@crusoe.ai header.b="AAd6JV5y" Received: by mail-pl1-f180.google.com with SMTP id d9443c01a7336-2cf50c6f235so19955555ad.0 for ; Thu, 20 Aug 2026 01:36:57 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=crusoe.ai; s=google; t=1787215014; x=1787819814; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to:content-type; bh=wuk/ok0Lrf0FihiMk3ElwPh6gTvSMLX5JskS22+aaRY=; b=AAd6JV5yZ1qgIRHj5qaYcRRD4vUUcOPqBO43IeXiZP1TgJn+5T2PPj/53Jw8gFkue7 ab6u56YW08ol59tZOb8KLSRtfNdzATwsQZv4hYwVp0S9do46z3VnfbL8Y2h49oP+CG4T 3CpjURwX+Pqquf4X91TPvopEtBau1r6Zpd6CKuGwQaie/dBSxe0eQ2/egUvCRkibujlN 6QL7YZRMP0Xr/PKKKxJ0DnsUqzgOeWVaYgRSlH+r7wv0mEBIHYOJ8x833psh2fXJdmJ9 1L51jIuer/WURLU13ZwBzOdavGhJjZycBJRvksJXedX3RunDEhiyYoynvBMLXXAvJlyl suUQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787215014; x=1787819814; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=wuk/ok0Lrf0FihiMk3ElwPh6gTvSMLX5JskS22+aaRY=; b=E//DPdmMbKZO0i5XYlDCPH7g/jlcrSAFs9P/QpxMcixPGahd82NNV6S5CBIyJj0WXJ Nb5RT9FkZCCjwNfPto7BOB81r7NYUDB1zqYSFquqWeqNNbUEYS92lnFHVEFEF4tJC0sy 3pyfP2kdpRWj2NX91Nrv7OnMX5VF307SbxJLskq3ewU310iyqbm37KebnzlRqmazx28A qm0SIfZODTAkBWv73FLJWK9nq0sC178tZZKfDTIjUJzhqcTDTTywFFgLEqEZbmAjxdVk nZF2r2h42KQ2pWGhV+qCQpnykDqkEda6ITLSh4zqkq3wNAVJ5kCZmh9+IkpFBRVTROZd mppw== X-Forwarded-Encrypted: i=1; AHgh+RqSJ/c5OyiIHbtqyJ6xn+HSBBddqs/QJI6yTasUSb3eCeYlQ9x59Akm4+BDhszy+BK/PgOJ8N/6vS6BgGs=@vger.kernel.org X-Gm-Message-State: AFuF++lyniWswYQnRArjMnhD96cklcQJYiWYIOMRRV07HXEr8wenx8vV lQxlzZyHLIsDUFLKL/nuVn3VadoGS7X4XgCAEtN26kZmzG21Q8Wy44Ea/O4gq6Zv1Gk= X-Gm-Gg: AR+sD104GEgyEXg2/xl3Q0FnNbXhnCD5/ZpLTsdqeDf7Tr3J7AaDmCSdZzRyRSQghAp xyJ8h9yDYRosjwxOcaHNuQRZNySQc37x0v/qs/edW8p1WP9IvvDv+grzymLWWxZzTdyb2e+K+wq yfNkyqqavX+O4Y7v3TzQwICG6iXkutqcg08/ka3iOm+scrxtI7Gqrckhw+Q8a1TWRjp77JI8pux PFZWcCXQ5O9jVHm/nBagIcr63wpMDv9QkVCkajBMeZ3dyHzNNL3dNhxOmQS3CcTk+6cy0/GL6Nj Fb+rSDtmBsm59iWmyjp5FeVgB/DniN9UQMdfPvKs0N0VhzVpIY1AGd8SJ5yG8RsxSZA8pvDD6On 51bM1e6MMbq4MQT18tyVdF+KnO0iKp4U9aMCgaS9R0bMD72NCYwi3lZ6xEs0eKG7Va1VyLVIfd1 Kdd2/n02whjsevLWtR3wWyASeyyMcg5CYqMTt8Xc1X4+yh6p8orvRIBi0fbEsrlKj4LpHkQ494O ApQlPv4BzaW8yQJB/8R9p3qOE204kfug9GrrgewGgDfdN0N4HeE1w== X-Received: by 2002:a17:90b:2b8e:b0:395:4de4:92be with SMTP id 98e67ed59e1d1-395810c60ffmr22022815a91.13.1787215013774; Thu, 20 Aug 2026 01:36:53 -0700 (PDT) Received: from MBP-Saravanan-D.civet-hops.ts.net ([2601:647:4380:3400:c90a:ce07:3842:3642]) by smtp.gmail.com with ESMTPSA id a92af1059eb24-1416aea2429sm13877076c88.15.2026.08.20.01.36.50 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Thu, 20 Aug 2026 01:36:52 -0700 (PDT) From: Saravanan D To: linux-nvme@lists.infradead.org Cc: kbusch@kernel.org, hch@lst.de, sagi@grimberg.me, axboe@kernel.dk, linux-kernel@vger.kernel.org, Saravanan D Subject: [PATCH v2] nvme-tcp: pin io_cpu to submitter cpu Date: Thu, 20 Aug 2026 01:36:34 -0700 Message-ID: <20260820083634.71689-1-saravanand@crusoe.ai> X-Mailer: git-send-email 2.53.0 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" nvme_tcp_set_queue_io_cpu() picks each queue's io_cpu at connect time, before any I/O exists, as the least loaded CPU in the queue's blk-mq map group, and all socket work then runs there for the connection's lifetime. On hosts that partition CPUs between pinned workloads a map group can straddle a partition boundary, so the pick can land one workload's socket processing on CPUs owned by another. On a 384 cpu multi tenant host with one VM driving ~1.1 GB/s of writes, 9% of nvme_tcp_io_work executions ran outside the submitting VM's cpuset, all on io_cpus of boundary straddling map groups, observed by the neighbor as steal time it did not cause. Adopt the submitting CPU as io_cpu for every command except the fabrics Connect. The submitter is a member of the map group by construction, and the nvme_tcp_cpu_queues accounting moves with each adoption. Connect is the only command on an I/O queue that does not represent the data path, since it is injected on an arbitrary CPU by blk_mq_alloc_request_hctx(), so it is skipped and the first real read or write decides. User passthrough is submitted from a real task on the submitting CPU and adopts like any other command. Queues outlive the workloads that submit through them, so adoption re-arms after 30 seconds of queue quiet. An idle queue is reclaimed by its next submitter, while a busy queue keeps a stable io_cpu and cannot ping pong between two live submitters. Concurrent writers on different CPUs serialize on a cmpxchg on io_cpu. The behavior is opt in per controller via the io_cpu_adopt fabrics option at connect time. wq_unbound takes precedence when set. Signed-off-by: Saravanan D --- Changes since v1 [1]: - Special case the fabrics Connect command instead of skipping all passthrough commands, so user passthrough I/O adopts too. - Make it a per-controller io_cpu_adopt fabrics option instead of a global wq_adopt module parameter, set once at connect time rather than flipped under a live connection. Both per Christoph Hellwig's review. [1] https://lore.kernel.org/linux-nvme/20260806023947.94680-2-saravanand@cr= usoe.ai/ drivers/nvme/host/fabrics.c | 4 ++ drivers/nvme/host/fabrics.h | 2 + drivers/nvme/host/tcp.c | 75 ++++++++++++++++++++++++++++++++++++- 3 files changed, 80 insertions(+), 1 deletion(-) diff --git a/drivers/nvme/host/fabrics.c b/drivers/nvme/host/fabrics.c index fd5abd04e080..26f8703744de 100644 --- a/drivers/nvme/host/fabrics.c +++ b/drivers/nvme/host/fabrics.c @@ -695,6 +695,7 @@ static const match_table_t opt_tokens =3D { { NVMF_OPT_NR_WRITE_QUEUES, "nr_write_queues=3D%d" }, { NVMF_OPT_NR_POLL_QUEUES, "nr_poll_queues=3D%d" }, { NVMF_OPT_TOS, "tos=3D%d" }, + { NVMF_OPT_IO_CPU_ADOPT, "io_cpu_adopt" }, #ifdef CONFIG_NVME_TCP_TLS { NVMF_OPT_KEYRING, "keyring=3D%d" }, { NVMF_OPT_TLS_KEY, "tls_key=3D%d" }, @@ -951,6 +952,9 @@ static int nvmf_parse_options(struct nvmf_ctrl_options = *opts, case NVMF_OPT_DATA_DIGEST: opts->data_digest =3D true; break; + case NVMF_OPT_IO_CPU_ADOPT: + opts->io_cpu_adopt =3D true; + break; case NVMF_OPT_NR_WRITE_QUEUES: if (match_int(args, &token)) { ret =3D -EINVAL; diff --git a/drivers/nvme/host/fabrics.h b/drivers/nvme/host/fabrics.h index caf5503d0833..3ecac041a628 100644 --- a/drivers/nvme/host/fabrics.h +++ b/drivers/nvme/host/fabrics.h @@ -67,6 +67,7 @@ enum { NVMF_OPT_KEYRING =3D 1 << 26, NVMF_OPT_TLS_KEY =3D 1 << 27, NVMF_OPT_CONCAT =3D 1 << 28, + NVMF_OPT_IO_CPU_ADOPT =3D 1 << 29, }; =20 /** @@ -140,6 +141,7 @@ struct nvmf_ctrl_options { unsigned int nr_poll_queues; int tos; int fast_io_fail_tmo; + bool io_cpu_adopt; }; =20 /* diff --git a/drivers/nvme/host/tcp.c b/drivers/nvme/host/tcp.c index 87d8067f3283..530e38695257 100644 --- a/drivers/nvme/host/tcp.c +++ b/drivers/nvme/host/tcp.c @@ -92,6 +92,7 @@ enum nvme_tcp_queue_flags { NVME_TCP_Q_LIVE =3D 1, NVME_TCP_Q_POLLING =3D 2, NVME_TCP_Q_IO_CPU_SET =3D 3, + NVME_TCP_Q_IO_CPU_ADOPTED =3D 4, }; =20 enum nvme_tcp_recv_state { @@ -105,6 +106,7 @@ struct nvme_tcp_queue { struct socket *sock; struct work_struct io_work; int io_cpu; + unsigned long last_data; =20 struct mutex queue_lock; struct mutex send_mutex; @@ -2783,6 +2785,74 @@ static void nvme_tcp_commit_rqs(struct blk_mq_hw_ctx= *hctx) queue_work_on(queue->io_cpu, nvme_tcp_wq, &queue->io_work); } =20 +/* Re-adopt io_cpu on the first data request after this much queue idle ti= me */ +#define NVME_TCP_READOPT_IDLE (30 * HZ) + +/* + * Adopt the CPU of the current data submission as the queue's io_cpu. + * + * The connect time choice in nvme_tcp_set_queue_io_cpu() picks the least + * loaded CPU in the queue's mq_map group before any I/O exists, so it + * cannot know which side of the group the actual submitters live on. On + * hosts that partition CPUs between pinned workloads a group that + * straddles a partition boundary can get an io_cpu on CPUs the submitting + * workload does not own, and its network processing then preempts an + * unrelated workload. The submitting CPU is in the queue's mq_map group + * by construction, so adopting it preserves the spreading property while + * landing the work on the side that generates it. + * + * Queues belong to the controller connection and outlive the workloads + * that submit through them, so adoption re-arms after NVME_TCP_READOPT_ID= LE + * of queue quiet. A successor workload reclaims an idle queue with its + * first data request, while a continuously busy queue keeps a stable + * io_cpu and cannot ping pong between two live submitters. + * + * The fabrics Connect command targets a specific queue via + * blk_mq_alloc_request_hctx() and so runs on an arbitrary CPU that does + * not represent the data path, so it is skipped and the first real read + * or write decides. All other commands, including user passthrough, + * carry a real submitting CPU and adopt. + * + * Adoption is opt in per controller via the io_cpu_adopt connect option + * and is bypassed when wq_unbound is set. + */ +static void nvme_tcp_adopt_io_cpu(struct nvme_tcp_queue *queue, + struct request *rq) +{ + struct nvme_command *cmd =3D nvme_req(rq)->cmd; + int old, new; + + if (!queue->ctrl->ctrl.opts->io_cpu_adopt || wq_unbound) + return; + if (!nvme_tcp_queue_id(queue)) + return; + if (nvme_is_fabrics(cmd) && + cmd->fabrics.fctype =3D=3D nvme_fabrics_type_connect) + return; + + if (test_bit(NVME_TCP_Q_IO_CPU_ADOPTED, &queue->flags) && + time_before(jiffies, READ_ONCE(queue->last_data) + + NVME_TCP_READOPT_IDLE)) { + WRITE_ONCE(queue->last_data, jiffies); + return; + } + + WRITE_ONCE(queue->last_data, jiffies); + set_bit(NVME_TCP_Q_IO_CPU_ADOPTED, &queue->flags); + + old =3D READ_ONCE(queue->io_cpu); + new =3D raw_smp_processor_id(); + if (old =3D=3D new || !try_cmpxchg(&queue->io_cpu, &old, new)) + return; + + if (test_bit(NVME_TCP_Q_IO_CPU_SET, &queue->flags)) { + atomic_dec(&nvme_tcp_cpu_queues[old]); + atomic_inc(&nvme_tcp_cpu_queues[new]); + } + dev_dbg(queue->ctrl->ctrl.device, "queue %d: adopted io_cpu %d\n", + nvme_tcp_queue_id(queue), new); +} + static blk_status_t nvme_tcp_queue_rq(struct blk_mq_hw_ctx *hctx, const struct blk_mq_queue_data *bd) { @@ -2802,6 +2872,8 @@ static blk_status_t nvme_tcp_queue_rq(struct blk_mq_h= w_ctx *hctx, =20 nvme_start_request(rq); =20 + nvme_tcp_adopt_io_cpu(queue, rq); + nvme_tcp_queue_request(req, bd->last); =20 return BLK_STS_OK; @@ -3047,7 +3119,8 @@ static struct nvmf_transport_ops nvme_tcp_transport = =3D { NVMF_OPT_HDR_DIGEST | NVMF_OPT_DATA_DIGEST | NVMF_OPT_NR_WRITE_QUEUES | NVMF_OPT_NR_POLL_QUEUES | NVMF_OPT_TOS | NVMF_OPT_HOST_IFACE | NVMF_OPT_TLS | - NVMF_OPT_KEYRING | NVMF_OPT_TLS_KEY | NVMF_OPT_CONCAT, + NVMF_OPT_KEYRING | NVMF_OPT_TLS_KEY | NVMF_OPT_CONCAT | + NVMF_OPT_IO_CPU_ADOPT, .create_ctrl =3D nvme_tcp_create_ctrl, }; =20 base-commit: bf881dd20062db5e951a0d0703cb476df8c9fdee --=20 2.53.0