From nobody Fri Oct 2 03:37:52 2026 Received: from mail-pj1-f41.google.com (mail-pj1-f41.google.com [209.85.216.41]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id ED5452B2D7 for ; Thu, 6 Aug 2026 02:45:32 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.41 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785984334; cv=none; b=EUWIU3UZ/eG72CplAUpcAO0wQVXvOXaIcriGkO0xyCZU0yEVB6Tewkh50RDfkTQiKfiIt9aj5YFvC87+Xviw5QcoGtHvG/PB91B4XuzHiwsBfcwaM72Hnn92Z5+ZnPG5IOLUK+/roHmYsGjQal6S0iKd68vKdc0FXHS7Rr3OHeA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785984334; c=relaxed/simple; bh=wc2ZJvk66upv6A3piooV79eTU2SbMvG3NnS1elPEUhI=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=WzvXVTeX4TVlIMncoIVA4rFWYUug7/gyU6eOwrjblrnWbJuLjR/r9Nip7e8heaI4j+pN4z8Yx9b4QeHbFvuGGDhFw58LIefAhyvAxHE7k9ROyjN3feHUUPT5/Jrm2nasorw2wrqA0N6Khhg01B+8fiKwJLRV/MLES93UF7na0aM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=crusoe.ai; spf=pass smtp.mailfrom=crusoe.ai; dkim=pass (2048-bit key) header.d=crusoe.ai header.i=@crusoe.ai header.b=K0XC00c9; arc=none smtp.client-ip=209.85.216.41 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=crusoe.ai Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=crusoe.ai Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=crusoe.ai header.i=@crusoe.ai header.b="K0XC00c9" Received: by mail-pj1-f41.google.com with SMTP id 98e67ed59e1d1-38dcbade417so1362846a91.1 for ; Wed, 05 Aug 2026 19:45:32 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=crusoe.ai; s=google; t=1785984332; x=1786589132; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to:content-type; bh=rrnbZo5m/vorntLLsKb/28XiJcabwy2i4minOlLpbkA=; b=K0XC00c9PwRcVgPWKDAWOHqkfz1mQWN3lA6XRobCuonO/X7e0Wvdy8NlrFz3Uvpor2 7McRMwXMu5LVw897OlGEtne71Z0ucz3UQUzVbjvXOaTTkIJIkqURl2W/ATqZzdy5iePs McZ7lJU4ooIqdOC6Q44wqkzIOd2QsLyhy96OrTABEfmGMsVtb2ZtKEdlGIT1AKC/OGXs DHpJ769+lqeJrDHQAZ+ctA3gRmkDOpRHCZF1OnD9M0Q9pIP+wMkvUVUEaxuFiD40DCSN jq+oYTiirETH01o1KFuCcasbzyXyftzgdfcmTVRl7S3M2Iwz4zchx3q+UoVcPEDrcGuJ AwPQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785984332; x=1786589132; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=rrnbZo5m/vorntLLsKb/28XiJcabwy2i4minOlLpbkA=; b=Ro0E3NjJsr1UIGtLDVRw5fFol9GiX8qGLW8HUVsfEP2Bb+HZWiatL8UpglVFEQWMCi dIQD6ZgfQwMqlt9pkEFdI34T3eJvlZxk2hqoh20Rtly6NmgXsv1nzmS84XPgoG23Pw0R FccVkuNWM6ERkoCxLjogrT0W3kv43EqR5LcDV71bAGHHIlDElLHnINDQFeFm8cV4b4NG 8L2cxlNT0ojBVjqRAP47hRT+2wkDSj8Gep+6H/RtCNCE6kE7ibZkewvWC7/0xzwKfV5U T58ViDSmevC9eyEUVIJ6DTXzL4V1CnSoOGLbTBD3u7vRSrOX7/rQhi/20/0sZubmF1jN VvgQ== X-Forwarded-Encrypted: i=1; AHgh+RpUzBKoKdJokhapKb6GPAkcwkSuUNqhlTF+rfLOsX/9C6i7wWu/IDewaJVm7JT5y8wfhaABNOf4+NyXliI=@vger.kernel.org X-Gm-Message-State: AOJu0YxURH47VNy/EMK7h/M2EPk/FtViu+UrN/mh/yw4tq8StNnHB2Zi LKZDymh4CNpZ7M20YlrZprk6Vscm5wScMKFdzie2eFgIMSzgMrJVJWeruNx0GEF88d4= X-Gm-Gg: AR+sD11pGZR8NeHDhjmxSxyVB/rf3ob3pO4ekenXRwxahSl78STV14jZFGjAWPNwMd2 Hu/bU1G8oVxegnIqX59dfRmD2yblHdYFda7+ew/gDeEfX9skvchE+a/UF/x5KPJoxQadq1yTv5p jmQ0mmcewk3+QMI9fljUMgOWNL2A7RSSCgFVT72uyEZsAWOk1ajoncdBpwOYzMY0fl7Nn4Ik05Y jN6vT18UkX/cdyfhlSJCqremgAG5iUdS70lO6SfHFf4KJzz9cxfPtGtv/jPbiifwVbI7AaXMQAZ WX2qgwpDLfMm76GCWwqfXwAsb32qXzkXWPJ2u98L4tOwVNAmrMliasA9TG17rCT90Fngi7GbCJk EvbJtOYfjs66KB+6+sA4Th4T8Z4zRFivs0RTSMN0syA1Whk3cBOj5MxONXxvrBm54C/AQBaTDDR U1lm4BnZ9V4m38TMJZuMzgoq3B8a+I3WtkBQv0i+86d3gPmGwGGlQ8IUO6YvfUfORrhDLQq6/Ex h5P077dQwCwVWz72EWN6zks6s2ZhEdk4Q== X-Received: by 2002:a17:90b:5112:b0:38e:542:6485 with SMTP id 98e67ed59e1d1-3903c5919c4mr11891440a91.12.1785984332151; Wed, 05 Aug 2026 19:45:32 -0700 (PDT) Received: from MBP-Saravanan-D.civet-hops.ts.net ([67.208.231.220]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-315a8179ee5sm2122066eec.18.2026.08.05.19.45.31 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Wed, 05 Aug 2026 19:45:31 -0700 (PDT) From: Saravanan D To: linux-nvme@lists.infradead.org Cc: Saravanan D , Keith Busch , Christoph Hellwig , Sagi Grimberg , Jens Axboe , linux-kernel@vger.kernel.org Subject: [PATCH] nvme-tcp: pin io_cpu to submitter cpu Date: Wed, 5 Aug 2026 19:39:44 -0700 Message-ID: <20260806023947.94680-2-saravanand@crusoe.ai> X-Mailer: git-send-email 2.53.0 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" nvme_tcp_set_queue_io_cpu() picks each queue's io_cpu at connect time, before any I/O exists, as the least loaded CPU in the queue's blk-mq map group, and all socket work then runs there for the connection's lifetime. On hosts that partition CPUs between pinned workloads a map group can straddle a partition boundary, so the pick can land one workload's socket processing on CPUs owned by another. On a 384 cpu multi tenant host with one VM driving ~1.1 GB/s of writes, 9% of nvme_tcp_io_work executions ran outside the submitting VM's cpuset, all on io_cpus of boundary straddling map groups, observed by the neighbor as steal time it did not cause. Adopt the submitting CPU of non passthrough requests as io_cpu. The submitter is a member of the map group by construction, and the nvme_tcp_cpu_queues accounting moves with each adoption. Passthrough commands do not participate since the io queue Connect arrives from an arbitrary group CPU via blk_mq_alloc_request_hctx(), so the first real read or write decides. Queues outlive the workloads that submit through them, so adoption re-arms after 30 seconds of queue quiet. An idle queue is reclaimed by its next submitter, while a busy queue keeps a stable io_cpu and cannot ping pong between two live submitters. Concurrent writers on different CPUs serialize on a cmpxchg on io_cpu. The behavior is opt in via the new wq_adopt module parameter, default off and runtime writable. wq_unbound takes precedence when both are set. Signed-off-by: Saravanan D --- drivers/nvme/host/tcp.c | 71 +++++++++++++++++++++++++++++++++++++++++ 1 file changed, 71 insertions(+) diff --git a/drivers/nvme/host/tcp.c b/drivers/nvme/host/tcp.c index 87d8067f3283..6509273f478c 100644 --- a/drivers/nvme/host/tcp.c +++ b/drivers/nvme/host/tcp.c @@ -44,6 +44,10 @@ static bool wq_unbound; module_param(wq_unbound, bool, 0644); MODULE_PARM_DESC(wq_unbound, "Use unbound workqueue for nvme-tcp IO contex= t (default false)"); =20 +static bool wq_adopt; +module_param(wq_adopt, bool, 0644); +MODULE_PARM_DESC(wq_adopt, "Adopt the submitting cpu as queue io_cpu (defa= ult false)"); + /* * TLS handshake timeout */ @@ -92,6 +96,7 @@ enum nvme_tcp_queue_flags { NVME_TCP_Q_LIVE =3D 1, NVME_TCP_Q_POLLING =3D 2, NVME_TCP_Q_IO_CPU_SET =3D 3, + NVME_TCP_Q_IO_CPU_ADOPTED =3D 4, }; =20 enum nvme_tcp_recv_state { @@ -105,6 +110,7 @@ struct nvme_tcp_queue { struct socket *sock; struct work_struct io_work; int io_cpu; + unsigned long last_data; =20 struct mutex queue_lock; struct mutex send_mutex; @@ -2783,6 +2789,69 @@ static void nvme_tcp_commit_rqs(struct blk_mq_hw_ctx= *hctx) queue_work_on(queue->io_cpu, nvme_tcp_wq, &queue->io_work); } =20 +/* Re-adopt io_cpu on the first data request after this much queue idle ti= me */ +#define NVME_TCP_READOPT_IDLE (30 * HZ) + +/* + * Adopt the CPU of the current data submission as the queue's io_cpu. + * + * The connect time choice in nvme_tcp_set_queue_io_cpu() picks the least + * loaded CPU in the queue's mq_map group before any I/O exists, so it + * cannot know which side of the group the actual submitters live on. On + * hosts that partition CPUs between pinned workloads a group that + * straddles a partition boundary can get an io_cpu on CPUs the submitting + * workload does not own, and its network processing then preempts an + * unrelated workload. The submitting CPU is in the queue's mq_map group + * by construction, so adopting it preserves the spreading property while + * landing the work on the side that generates it. + * + * Queues belong to the controller connection and outlive the workloads + * that submit through them, so adoption re-arms after NVME_TCP_READOPT_ID= LE + * of queue quiet. A successor workload reclaims an idle queue with its + * first data request, while a continuously busy queue keeps a stable + * io_cpu and cannot ping pong between two live submitters. + * + * Passthrough commands (the io queue Connect in particular) are submitted + * from an arbitrary group CPU by blk_mq_alloc_request_hctx() and do not + * represent the data path, so they are skipped and the first real read + * or write decides. + * + * Adoption is opt in via the wq_adopt module parameter and is bypassed + * when wq_unbound is set. + */ +static void nvme_tcp_adopt_io_cpu(struct nvme_tcp_queue *queue, + struct request *rq) +{ + int old, new; + + if (!wq_adopt || wq_unbound || blk_rq_is_passthrough(rq)) + return; + if (!nvme_tcp_queue_id(queue)) + return; + + if (test_bit(NVME_TCP_Q_IO_CPU_ADOPTED, &queue->flags) && + time_before(jiffies, READ_ONCE(queue->last_data) + + NVME_TCP_READOPT_IDLE)) { + WRITE_ONCE(queue->last_data, jiffies); + return; + } + + WRITE_ONCE(queue->last_data, jiffies); + set_bit(NVME_TCP_Q_IO_CPU_ADOPTED, &queue->flags); + + old =3D READ_ONCE(queue->io_cpu); + new =3D raw_smp_processor_id(); + if (old =3D=3D new || !try_cmpxchg(&queue->io_cpu, &old, new)) + return; + + if (test_bit(NVME_TCP_Q_IO_CPU_SET, &queue->flags)) { + atomic_dec(&nvme_tcp_cpu_queues[old]); + atomic_inc(&nvme_tcp_cpu_queues[new]); + } + dev_dbg(queue->ctrl->ctrl.device, "queue %d: adopted io_cpu %d\n", + nvme_tcp_queue_id(queue), new); +} + static blk_status_t nvme_tcp_queue_rq(struct blk_mq_hw_ctx *hctx, const struct blk_mq_queue_data *bd) { @@ -2802,6 +2871,8 @@ static blk_status_t nvme_tcp_queue_rq(struct blk_mq_h= w_ctx *hctx, =20 nvme_start_request(rq); =20 + nvme_tcp_adopt_io_cpu(queue, rq); + nvme_tcp_queue_request(req, bd->last); =20 return BLK_STS_OK; base-commit: bf881dd20062db5e951a0d0703cb476df8c9fdee --=20 2.53.0