From nobody Sat Jul 25 04:18:31 2026 Received: from mail-wm1-f47.google.com (mail-wm1-f47.google.com [209.85.128.47]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A8B3F3DCD8A for ; Sat, 18 Jul 2026 16:26:16 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.47 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784391978; cv=none; b=nQ5SQIGeEGw8jAy0KRHyNoXVq5bvE5U/GoRMBsUgNWXp+gr0iz7Gfa4aR/cVeqnWnzLUvtpT/HuED+j85u0qLaAQ4HXLkCv+hYbASLs7AxsNSQjCleGYbg0zHMa8poo+ptHC9N/ppQspoLPopg43LbQZF9YxXpXr3rIMLoD84sg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784391978; c=relaxed/simple; bh=vQcr8xq5kWGZvugjiEEaiZG5UL1l4rnSHXCuTtt66Sw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=p7GpYQyiRbCNbHita1XC37HluAAeRn81A8OksuuG0BWVQW4adCigOOpZ+YCBu93iBUnp2Rbvg3TZBCg9ljbET63DEftaWjJk2QUB9YdQjc7pZDthUGZ6v4PUfEXPEFMVWHJYrBi288MLmml0zocW7LqLubKLdwqGjxkM9X4nGZI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=meshstor.io; spf=pass smtp.mailfrom=meshstor.io; dkim=pass (2048-bit key) header.d=meshstor.io header.i=@meshstor.io header.b=raDKZ11L; arc=none smtp.client-ip=209.85.128.47 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=meshstor.io Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=meshstor.io Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=meshstor.io header.i=@meshstor.io header.b="raDKZ11L" Received: by mail-wm1-f47.google.com with SMTP id 5b1f17b1804b1-4921eed3fa2so68514455e9.0 for ; Sat, 18 Jul 2026 09:26:16 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=meshstor.io; s=google; t=1784391975; x=1784996775; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=v4Dvf1BFTy/Vc0nJhWIM7SPA1tR6TPAktef7abvXxvQ=; b=raDKZ11LisN1xvciqBKE3E4FmiVh6k0JJw3aZiG3DmFsTUGeGF/UIj48Q4QvywzYYw Xku4DGrRIZGL/YTY1lIogmT5AetVNkE9zmydJKhtiiAG4i1T5j9UdTolimwQDBp4pN1A bsO0tjG3cNv0sM0ltGjqTipr5kd1zAsBC5ET38abxseD+cXyokkuYQvVJ9GMsfYacuSL f0C7qZPhFAwW6ZX4qwQXs5gbQlCYaVIgd0FlXCgmigADVJjmA3m5dTIymz+X6WDjiBE1 cUtLrWI3hLUQPz7MWhF4wKCuJ+AbOfBQjOFOs/XDKAad6NhIzyicMLc7Ee/Hc42SZIII deLw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784391975; x=1784996775; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=v4Dvf1BFTy/Vc0nJhWIM7SPA1tR6TPAktef7abvXxvQ=; b=cchBeEU9wOYUm0sdUAxGDTqaeEeH1EM9r1gGE/Kw3cJYboDSCwYhjdSo0tCBveroeu RFqj2i6JQ7RVogRf1km9ndxWQVYJqjskaKPPh8f3njLnJr06k9mwjcirclflU/w/6/lN Zg8iiVnFlrpha3tSYhLRITQtxi9mzrJijn3MCQoH4+1C/Cwsb6nauFb/+UJA4azXuboN N9Wkd+QLKmimHYlPD6gJ3mqAvxMa5HkWrEWRNFkmWHITomEROgy3ymC7e2SparhDSwvn e7fDA9Vr4cSSssxTnATJMwxQp+CyZrjewJIabTRn3WjLjxKBIx/miVgvHsfWLDvz4uWI jQZg== X-Forwarded-Encrypted: i=1; AHgh+RoYPAaw/4H9ajCLw/wF6OsHJPzzA6DiK8ac7WOi2KpOwHrf5PnXjbt6Ezgyd27hrRDMK9mRBA7pmMvKFAk=@vger.kernel.org X-Gm-Message-State: AOJu0YxKxwXQVzglxOy8S6FToV/MdP7vNzaqM17ZFakYtwz7ctT3FxVM JV+RXImWRjs+qUs+lnGRRIw+y2KhJBXUp0n4kORxSGaH8ZgLtPFW17ie6qAtRiziKQ== X-Gm-Gg: AfdE7cnkh1TyAP5XUMKR+iOmO2ZxNsR+szsyF5g3QifommAVU5RdOJ3mSqKB1Adfmlz ylii/KKjRJBi6HwhibIdWHNJGDCed/DoG0o8/zleL+FRJh9Fve6udyFykj3cxfsuqiuo+NG+6+l BQzyzdl1braz9TLddOD5jjanlGb0xezPWBenvX0EEzMG8H6IXm4BbKYr9yEsOW29urCMa3JjxPh C8eJzjhEXE0c90AHPWjfm3/haybTdMMt7JyULJsUVazlAuX3HSJ7ibZ3lFR4arc5KXgqGXxx+tT /vOmABVpA12P3C9Herbdeg+RJ0h9D+LJRxOQOnmRzQZIZgnM3CDYJ3L4h7wlJzTBw0OsAxMiMzu PRxuQmm3sVDsKtUCsLeR09WVeL5leXNgsEIp3M+k9yQxe+VpKo1mg7RlqOFROCA== X-Received: by 2002:a05:600c:6b69:b0:493:e583:7053 with SMTP id 5b1f17b1804b1-4954a41431cmr58900345e9.35.1784391974915; Sat, 18 Jul 2026 09:26:14 -0700 (PDT) Received: from mf-00-01.. ([194.220.239.180]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-4954a2eddb8sm266444085e9.14.2026.07.18.09.26.13 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 18 Jul 2026 09:26:14 -0700 (PDT) From: Mykola Marzhan To: Jens Axboe , Song Liu , Yu Kuai , Keith Busch , Christoph Hellwig , Sagi Grimberg , linux-block@vger.kernel.org, linux-raid@vger.kernel.org, linux-nvme@lists.infradead.org Cc: Li Nan , Xiao Ni , Leon Romanovsky , Jason Gunthorpe , Kiran Kumar Modukuri , Chaitanya Kulkarni , Logan Gunthorpe , Bjorn Helgaas , Shivaji Kant , Pranjal Shrivastava , Henrique Carvalho , linux-kernel@vger.kernel.org, linux-rdma@vger.kernel.org, linux-pci@vger.kernel.org Subject: [PATCH 1/6] blk-mq-dma: restore BLK_STS_TARGET for unsupported P2P transfers Date: Sat, 18 Jul 2026 16:25:42 +0000 Message-ID: <20260718162547.448892-2-mykola@meshstor.io> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260718162547.448892-1-mykola@meshstor.io> References: <20260718162547.448892-1-mykola@meshstor.io> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Commit 91fb2b6052f7 ("nvme-pci: convert to using dma_map_sgtable()") deliberately mapped unsupported P2PDMA transfers to BLK_STS_TARGET, matching dma_map_sgtable()'s -EREMOTEIO: "When this happens, return BLK_STS_TARGET so the request isn't retried." The conversion to blk_rq_dma_map silently changed the status to BLK_STS_INVAL, which regresses two consumers: - md/raid1 and raid10 ignore BLK_STS_INVAL leg failures (commit f7b24c7b41f2 ("md/raid1,raid10: don't fail devices for invalid IO errors"), where it means a request-shaped error that fails identically on every member). Since commit 02666132403a ("md: propagate BLK_FEAT_PCI_P2PDMA from member devices to RAID device") P2PDMA bios reach md arrays: a peer-memory write to a member the peer cannot reach is counted as written, the master bio reports success, mirrors silently diverge, and on a topology where no member is reachable the write reports success with zero copies on stable storage. - the failure's classification flips: for direct NVMe consumers (nvme advertises BLK_FEAT_PCI_P2PDMA) the mapping failure has surfaced as EINVAL instead of the documented -EREMOTEIO since v6.17, and blk_path_error(BLK_STS_INVAL) is true, so a stacking consumer would treat it as a retryable path error (dm-mpath is the only in-tree blk_path_error() caller; P2P bios cannot currently reach it, but the classification is wrong on its face). Restore BLK_STS_TARGET. md then routes an unreachable leg through its per-device error handling (badblocks, WantReplacement, mirror retry for reads); a follow-up patch teaches raid1/raid10 to handle mapping failures without the retry storms that machinery was built around. Fixes: 858299dc6160 ("block: add scatterlist-less DMA mapping helpers") Fixes: 7ce3c1dd78fc ("nvme-pci: convert the data mapping to blk_rq_dma_map") Cc: stable@vger.kernel.org Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Mykola Marzhan Reviewed-by: Logan Gunthorpe --- block/blk-mq-dma.c | 10 +++++++++- 1 file changed, 9 insertions(+), 1 deletion(-) diff --git a/block/blk-mq-dma.c b/block/blk-mq-dma.c index bfdb9ed70741..2eed06bfe791 100644 --- a/block/blk-mq-dma.c +++ b/block/blk-mq-dma.c @@ -190,7 +190,15 @@ static bool blk_dma_map_iter_start(struct request *req= , struct device *dma_dev, case PCI_P2PDMA_MAP_NONE: break; default: - iter->status =3D BLK_STS_INVAL; + /* + * P2P transfers that the mapping layer cannot support + * report BLK_STS_TARGET, matching dma_map_sgtable()'s + * -EREMOTEIO and the pre-blk_rq_dma_map nvme behavior: + * the failure is a property of this device pairing, so + * it must not be retried on another path (blk_path_error) + * nor be mistaken for an invalid request. + */ + iter->status =3D BLK_STS_TARGET; return false; } =20 --=20 2.43.0 From nobody Sat Jul 25 04:18:31 2026 Received: from mail-wr1-f54.google.com (mail-wr1-f54.google.com [209.85.221.54]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 10B2F3CF68F for ; Sat, 18 Jul 2026 16:26:19 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.221.54 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784391983; cv=none; b=EwzTMrJfy8YGJTbsss+sFXUFX4YQvyqUxmOdWxp223rbd7PfSFkEgP3yL5ZUcKDj5tLRYcIVkLke0zXk5pKNi1gFfbrkokBNIbycA7KpEwy1eDIODXYyf0lmJ1z77kMF58rTXqNOFHI/9hLrgBy6+Tgm9wSaOXQkcCl5jzqbFWY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784391983; c=relaxed/simple; bh=vO1jpq1MeVozqKW9Fk/0t1VAaeJE5a63mWcH5DQ3WaQ=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=UCRZQjRJLDAia7A85It3bYT0ZAy+Jbh/7nENcBFSfmwwzrjtGig05wuTMVJO6++mtPwaVPSIW0tLeY4hOegjp2ACqxjZSRM8eV3Vreaa8htANJoS7G6m52X6gabBPyyAy7rz0CojjPyg5f5YNGBpM7TwYWnU2Ymdn94xXT8tgSM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=meshstor.io; spf=pass smtp.mailfrom=meshstor.io; dkim=pass (2048-bit key) header.d=meshstor.io header.i=@meshstor.io header.b=dGEd0KGF; arc=none smtp.client-ip=209.85.221.54 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=meshstor.io Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=meshstor.io Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=meshstor.io header.i=@meshstor.io header.b="dGEd0KGF" Received: by mail-wr1-f54.google.com with SMTP id ffacd0b85a97d-4799b3f7c83so6338543f8f.2 for ; Sat, 18 Jul 2026 09:26:19 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=meshstor.io; s=google; t=1784391978; x=1784996778; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=FhGxTEOOON4RkKT9KdpeyXrdqn2SrMJKFWwPDsn11x4=; b=dGEd0KGFycs70IRTPaRO3yl067VneJhz8nAfq496d1UdH5tmzBiE4NqoHRmlD1vqEe LxkfOCoJnl8dzncWTJUCWZvf3EshSpqQZiiEEwGXXu6MzrxBm8UbAvBF8xDjzq+83oif JCyq4Kmx5xPM/4PAsQjWPOZoJss+XttXYI7ZGodnqpvbYgJBbwG+GxH7jjeo6qCdQvA2 nR0RX4tJ4bopIsNJyNeBOl66rQOGEhN6O4Lfpd58TXGvYmMSsGgJw0061VsLEyLvBgCg KPhhnYTjSFu9367sE0oElHT+efoL2NEBI4HcbUz6njWBcacHaWuSjQj8eWCRCWM0Ngpt 2LaA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784391978; x=1784996778; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=FhGxTEOOON4RkKT9KdpeyXrdqn2SrMJKFWwPDsn11x4=; b=hUc3qcYjf7IGQCVpIFfj5gLEl/Md65YUIzkvc8pbfgQRZo5C9kGtNvyHSysy2pjWpj PWp4dX+ckYC8GQXDt2Xjd9B33BhyNbbk2SyFWg4pH7GiKxgua4SYTP5UzBSZ65GDB/3x tpPAafBAqXtRgpSF/Dk9FczUARTTmyh9xq251N/bjEj3n6m8n2Y8LbQuR/Jt5BfH3HiJ D8FqZTlcLOmM56TveQOGRIeR5HscIRZRTdXM+qsCllQ4Aa+45QyF7phwjPv1eHRrGO3w eB/pwNVsWVxzH9qC14RZ2vgN6sKRwN4uEb7+Ud7IZUUZSZaFmXKc1hJcvwOb6vDvrvst wx/Q== X-Forwarded-Encrypted: i=1; AHgh+Rp5jt66uuLExY6FPpH1GudtsH6jANXKceWpSuca3U4OB4aRJejeSvOveuK8Yy07eAasnzEfGn8Plq/6y5w=@vger.kernel.org X-Gm-Message-State: AOJu0YwFn56V4xV0880guwyXqPfWwJ+FrJl8d82B4uG4PbAPM/dHQMDO d7Xj2l97q4Cs1Nb7nwcBSjdYuWiqvHCTp68+LZ8hMIbPE7Dty2TAPA3NJm5FjFvkgQ== X-Gm-Gg: AfdE7cnGc9LLaQdYz5NSyWo0ZI40/2NdhnDE1NrE+4VQ0C2I/M3rlqgjvgPVESvWagT BOop0DEgEeFA/L/L6A9sLQ+gG64lglHbi76JjCuVYDGqG+QpewPoMKDMYU5CfD7CzwH9+WQG0nT x4RNfPCSpKgl9B2zY9uM/Eyhr4s7orMc9K/vuAfKdk1hl4vRcQYTtLfHEx4MmfY4V706LIzwNij j7M3ihSR2M83DmARYsrCLxnY5MhFA3O7qN3noTteB6mv7VBK3EpXyqbY+LE/sfkbjp0XWxrNs7T JjIllmRHucXbZHd7JrAlMJS0t4ws0bmavy/ZPk2Fy0YQ2SxdmdXj96Bs7oTcpUS/4xnT8RudcIS jV0551D2aQsnyt2so2guvy4wLwj8fuQuskKe1/OKlisL8bHVAWA4gWmolZfUM3Q== X-Received: by 2002:a05:600c:c84:b0:495:4689:1e98 with SMTP id 5b1f17b1804b1-4954a3ed426mr78278195e9.10.1784391978225; Sat, 18 Jul 2026 09:26:18 -0700 (PDT) Received: from mf-00-01.. ([194.220.239.180]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-4954a2eddb8sm266444085e9.14.2026.07.18.09.26.16 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 18 Jul 2026 09:26:17 -0700 (PDT) From: Mykola Marzhan To: Jens Axboe , Song Liu , Yu Kuai , Keith Busch , Christoph Hellwig , Sagi Grimberg , linux-block@vger.kernel.org, linux-raid@vger.kernel.org, linux-nvme@lists.infradead.org Cc: Li Nan , Xiao Ni , Leon Romanovsky , Jason Gunthorpe , Kiran Kumar Modukuri , Chaitanya Kulkarni , Logan Gunthorpe , Bjorn Helgaas , Shivaji Kant , Pranjal Shrivastava , Henrique Carvalho , linux-kernel@vger.kernel.org, linux-rdma@vger.kernel.org, linux-pci@vger.kernel.org Subject: [PATCH 2/6] md: ensure REQ_NOMERGE is set on P2PDMA bios Date: Sat, 18 Jul 2026 16:25:43 +0000 Message-ID: <20260718162547.448892-3-mykola@meshstor.io> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260718162547.448892-1-mykola@meshstor.io> References: <20260718162547.448892-1-mykola@meshstor.io> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" md_submit_bio() unconditionally strips REQ_NOMERGE before passing the bio to the personality, an optimization from commit 9c573de3283a ("MD: make bio mergeable"): a bio that md has split may become mergeable again below md. For PCI P2PDMA bios the flag is load-bearing, not a hint. The block layer sets REQ_NOMERGE on P2PDMA bios (__bio_add_page(), and the extraction path of bio_iov_iter_get_pages()) because the DMA mapping type of a request is resolved once, from its first segment (blk_dma_map_iter_start()), and request-level merging is prevented only by REQ_NOMERGE. Stripping it allows the member queue to merge a P2PDMA bio with a bio carrying pages of a different pgmap, or host memory, mapping the merged segments with the wrong bus address: silent data corruption on the member. Set the flag for P2PDMA bios instead of merely preserving it. No in-tree path currently submits P2PDMA pages through the bvec-iter path (bio_iov_bvec_set()), which skips the flagging -- but nothing structural prevents one, so setting rather than preserving hardens md against that gap at the cost of one branch. Everything else keeps the original optimization of clearing the flag. This covers every personality that advertises BLK_FEAT_PCI_P2PDMA (raid0, raid1, raid10), which is why the fix lives in the shared md_submit_bio() path. Fixes: 02666132403a ("md: propagate BLK_FEAT_PCI_P2PDMA from member devices= to RAID device") Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Mykola Marzhan Reviewed-by: Logan Gunthorpe --- drivers/md/md.c | 14 ++++++++++++-- drivers/md/md.h | 18 ++++++++++++++++++ 2 files changed, 30 insertions(+), 2 deletions(-) diff --git a/drivers/md/md.c b/drivers/md/md.c index d1465bcd86c8..3ae4fd4ef381 100644 --- a/drivers/md/md.c +++ b/drivers/md/md.c @@ -451,8 +451,18 @@ static void md_submit_bio(struct bio *bio) return; } =20 - /* bio could be mergeable after passing to underlayer */ - bio->bi_opf &=3D ~REQ_NOMERGE; + /* + * A bio md split could be mergeable again below md, but for P2PDMA + * bios REQ_NOMERGE is load-bearing: the DMA mapping type of a + * request is resolved once, from its first segment, so requests + * must stay single-provider (see __bio_add_page()). Set the flag + * rather than merely preserve it -- bios built through the + * bvec-iter path arrive without it. + */ + if (md_bio_is_p2pdma(bio)) + bio->bi_opf |=3D REQ_NOMERGE; + else + bio->bi_opf &=3D ~REQ_NOMERGE; =20 md_handle_request(mddev, bio); } diff --git a/drivers/md/md.h b/drivers/md/md.h index d8daf0f75cbb..140e2b3670d8 100644 --- a/drivers/md/md.h +++ b/drivers/md/md.h @@ -11,8 +11,10 @@ #include #include #include +#include #include #include +#include #include #include #include @@ -22,6 +24,22 @@ #include =20 #define MaxSector (~(sector_t)0) + +/* + * Check if the bio carries PCI P2PDMA (peer device memory) pages. Read + * bi_io_vec directly rather than using bio_first_bvec_all(), which WARNs + * on cloned bios: md routinely handles split clones, which have + * bi_vcnt =3D=3D 0 but a valid bi_io_vec shared with the parent. P2PDMA a= nd + * host pages must not be mixed within one bio, so the first bvec is + * representative. Only valid before the bio's iterator is consumed: + * bio_has_data() is false at completion time. + */ +static inline bool md_bio_is_p2pdma(struct bio *bio) +{ + return bio_has_data(bio) && bio->bi_io_vec && + is_pci_p2pdma_page(bio->bi_io_vec->bv_page); +} + /* * Number of guaranteed raid bios in case of extreme VM load: */ --=20 2.43.0 From nobody Sat Jul 25 04:18:31 2026 Received: from mail-wm1-f46.google.com (mail-wm1-f46.google.com [209.85.128.46]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BC13D3D75B5 for ; Sat, 18 Jul 2026 16:26:22 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.46 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784391984; cv=none; b=KH2rRRapuIbOlM6i0xeFsEAsA7Fk/knbgW86KpRtPES9ZXg1uJK9ovIwNwq35A9mug4+ZWmWIWgXQJbiJx1x7OCdKW/hGWdh7eK3oQJ96XrGfvEOUu6ll69vNyvCozeR0mUIbBnfbkp9n5fEwnup9VKUNsQPFlV1/+ZHzdDJC6k= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784391984; c=relaxed/simple; bh=SmuOn8Uxw8XZnxg47IbvGapBEh1SrF66f7V9e1FGwgw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Ed403ICaNepkQ682CaNoIMuJoxRpRMC5EjtDGOavrbK2kHa1HefvLglEnocGmh3+i6LkivwDuPRKuA+9QAoo52CEQ7vUpempdhPecvedObpUK9b5SS4XQHkNozVTqEtu8FhOCGlaSaCFJN+tw9asdF8HXZMeDdHSjz3G/pbwi74= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=meshstor.io; spf=pass smtp.mailfrom=meshstor.io; dkim=pass (2048-bit key) header.d=meshstor.io header.i=@meshstor.io header.b=QAcSNHXR; arc=none smtp.client-ip=209.85.128.46 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=meshstor.io Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=meshstor.io Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=meshstor.io header.i=@meshstor.io header.b="QAcSNHXR" Received: by mail-wm1-f46.google.com with SMTP id 5b1f17b1804b1-49556f97a9dso606105e9.1 for ; Sat, 18 Jul 2026 09:26:22 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=meshstor.io; s=google; t=1784391981; x=1784996781; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=95R+xIP/87RyRb0dz+s7R+IaY6bqD6sOy0TrsJCA7Fk=; b=QAcSNHXR7WJ8lkEc1DCyHf8bkHY2j6tNo2m/A1JddgbVd55sspZVEr/ZvfDHZZHaTJ dvhFfMWHO8cLD/Qr++X44q4PoP5dAE0nuovSogtsp1HNJp1fgaC4xMUWgS2uu3+hcIZR 9G1GqBdTCT1Qh/iit0Q2yoTLqBJuzinVb+Qa27A4HH9Pw5qJh4Wk8uGppINWvReA0xa3 M6B+nQEsu7u1uHpgw2223OCsHGKHQ90lQto6gvxxEAZ1I1vFelss9cy2KEYsg10SPEgB uy8PbEwJ1e5SdKdqzQgS6uxgP7A1pyPnxbeQwSCMorxdVrNgB5HygORPPahjkrRw4DAQ PB0w== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784391981; x=1784996781; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=95R+xIP/87RyRb0dz+s7R+IaY6bqD6sOy0TrsJCA7Fk=; b=REyc/hsKLufzK+dSUxb7l0mHQqJhjywrZBRlzfziRn2cGmI/A+3qW+2GhYGB4r/lkh RHx9ZHWGxoirCLzNy7XdsAKn+GMHOQWWks1Kquu52dBsZL3PzSYFV925+kBLxnPJzqLl YSl7IE1rL6cVDjFZBYpTRZoZGJXMatakGoE+r+8LkfgOWrFH48IOfSvS25QAITSfa4xZ zbSu9v28P/V1bdSu4LLtduBU1ZDYH/614B5EFCnZ3SloTgbui6w+KiVyQkxQVq+E2tDo ZxpFNLlL/zFKHBZER1QwuhbfqBfPZSjS8tgHSagNrXmlzwsJ6rkeolTu5PInsFoB56nm QNeg== X-Forwarded-Encrypted: i=1; AHgh+RpZVokKsm0tmbCt0z8/8MGChNNgHEQs/hjwK/kZMp8bvD7QbJ7bRnNAidit5Z8scEI5k/4uWflntCclpF8=@vger.kernel.org X-Gm-Message-State: AOJu0YzrgW3Wf+QrFC+tftIi9siABqlxa4fK59cL4quDINhIUSwx/2B/ CnkrWfd0AQzvgHgsoMqBMtEtvZ4e0/cYGDbKXU8BimQXY6bF4H/FKw0sELb2PqPkoQ== X-Gm-Gg: AfdE7clSKcCo6B8DzhfkGGr/0htdF5Kl0Hk7/YP+mAq6y9e/6UzREETMXSK8NiGxnss K86uvul43wsE5Qq63n7yvn90tzQ36sNIWd7ZPm05sQgccx2EzH7H4zklMHmcS2GBUqJFGGN8Ffp E8zjzSDlSnyKKJa3x/88V9GIl3P1XIh8+jZDIv4tCYMIZ3dp0hwhJfSFJw2Tyw1EiB4/oZkVUyj +FGUEQROJ2dTVlsRzVFeRhc68tUg+ay51dSOS3R3oBQYEffdn54WRZkDqhVIygxrG2elwM2cEnr FTOa7KuZL6ZZjMNbYWSyC5f4LSRCzNQQm8ucryt4UkH8GngTylN2zpIGrvleobkOj5nI6hw/tX/ Zhg+mNusOegBfQod8T2G+WWlT0okNjNDQ2KOtlu2c01rhR5LoEVzypRgetGOHuA== X-Received: by 2002:a7b:c054:0:b0:495:401d:9f4f with SMTP id 5b1f17b1804b1-4954a40513fmr53645805e9.25.1784391980997; Sat, 18 Jul 2026 09:26:20 -0700 (PDT) Received: from mf-00-01.. ([194.220.239.180]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-4954a2eddb8sm266444085e9.14.2026.07.18.09.26.19 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 18 Jul 2026 09:26:20 -0700 (PDT) From: Mykola Marzhan To: Jens Axboe , Song Liu , Yu Kuai , Keith Busch , Christoph Hellwig , Sagi Grimberg , linux-block@vger.kernel.org, linux-raid@vger.kernel.org, linux-nvme@lists.infradead.org Cc: Li Nan , Xiao Ni , Leon Romanovsky , Jason Gunthorpe , Kiran Kumar Modukuri , Chaitanya Kulkarni , Logan Gunthorpe , Bjorn Helgaas , Shivaji Kant , Pranjal Shrivastava , Henrique Carvalho , linux-kernel@vger.kernel.org, linux-rdma@vger.kernel.org, linux-pci@vger.kernel.org Subject: [PATCH 3/6] md/raid1: don't use write-behind for P2PDMA bios Date: Sat, 18 Jul 2026 16:25:44 +0000 Message-ID: <20260718162547.448892-4-mykola@meshstor.io> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260718162547.448892-1-mykola@meshstor.io> References: <20260718162547.448892-1-mykola@meshstor.io> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" alloc_behind_master_bio() copies the bio's data with bio_copy_data(), a CPU copy. P2PDMA pages are peer device (BAR) memory; generic code must not assume CPU load/store access to them is safe or fast on every architecture, and bouncing peer memory through the CPU defeats the point of a peer-to-peer transfer. Skip write-behind for P2PDMA bios: they are written directly to all members, including write-mostly ones. Fixes: 02666132403a ("md: propagate BLK_FEAT_PCI_P2PDMA from member devices= to RAID device") Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Mykola Marzhan --- drivers/md/raid1.c | 7 +++++-- 1 file changed, 5 insertions(+), 2 deletions(-) diff --git a/drivers/md/raid1.c b/drivers/md/raid1.c index afe2ca96ad8c..57f64e890102 100644 --- a/drivers/md/raid1.c +++ b/drivers/md/raid1.c @@ -1575,9 +1575,12 @@ static bool raid1_write_request(struct mddev *mddev,= struct bio *bio, /* * The write-behind io is only attempted on drives marked as * write-mostly, which means we could allocate write behind - * bio later. + * bio later. P2PDMA bios are excluded: write-behind copies + * the data with bio_copy_data(), a CPU copy that cannot be + * assumed safe or fast on P2PDMA (device BAR) pages. */ - if (!is_discard && rdev && test_bit(WriteMostly, &rdev->flags)) + if (!is_discard && rdev && test_bit(WriteMostly, &rdev->flags) && + !md_bio_is_p2pdma(bio)) write_behind =3D true; =20 r1_bio->bios[i] =3D NULL; --=20 2.43.0 From nobody Sat Jul 25 04:18:31 2026 Received: from mail-wm1-f41.google.com (mail-wm1-f41.google.com [209.85.128.41]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 802133DCD97 for ; Sat, 18 Jul 2026 16:26:25 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.41 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784391990; cv=none; b=l9NlCfCBSSnmztjWhzFpfr58kb6CioOozra7R9SHsaRV8PyQDzi9pGp+wzBRjYu/1vKRfSrKiw8AImvEZowizihH/NV+fSPTr9KBRvg0SRqE58/w2wrQJH7pzbuWmphaATb33Doww3waW+LSiGdEjp5eL9dYqdbHMCBXqwU3ljk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784391990; c=relaxed/simple; bh=bc8MxzvbVIXgCq+hE70J7nSNv3Wiq2lrvNKvz8Yb5i4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=maAu6wcYR+gAKb4Ez77Llp3XAt+flbnRk6Mho6CYE9EbiQb6ViwcDB2fLU3uMWoCyi/vttChEhU1IWeDJ9rbWnlVVajzqp4Vq3uONDUeoD6DI40DKSOCPYdRG/M9SAHslm5iDpzjygbYAM01RJSwiMPDQf1oTSdIExx1rvGsdJg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=meshstor.io; spf=pass smtp.mailfrom=meshstor.io; dkim=pass (2048-bit key) header.d=meshstor.io header.i=@meshstor.io header.b=murysiT0; arc=none smtp.client-ip=209.85.128.41 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=meshstor.io Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=meshstor.io Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=meshstor.io header.i=@meshstor.io header.b="murysiT0" Received: by mail-wm1-f41.google.com with SMTP id 5b1f17b1804b1-4953de5be0aso21979135e9.0 for ; Sat, 18 Jul 2026 09:26:25 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=meshstor.io; s=google; t=1784391984; x=1784996784; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=CUxR31Xh85+TAm3j4nZpX4fvNBvJG7NmZg0i4br8yQ0=; b=murysiT0EDR3QJc8dX8UEXj7Qx7YqS5luH0QXYnY2LoVxmqXB/hDbwz0lDT6/bOMYY 93zuVvzwfXfgn9G06TCyC/PW//Wya0zWw7DKa5gxIPyd+OOT7ul8Xosjv22ImWvth730 70xJA+mdQ209vab47/Msu78SF9sXPHI1/YZ9Mt0wZMLAtc2d5p6sE7wg9FXRFWT7bb6X LzJVCw+yJZ9/ruXNJh6/96vHdWB3zWb+YKVwPqwrSSE89v7FVwnsywXJJGydbYlDTh/R +ZSMX7W0dYfO54nN7TOV3g6Jsf0udkt4L5Q0jRzxjR+ONr1vcKTJouI7LI6HQJHYQQi5 92lQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784391984; x=1784996784; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=CUxR31Xh85+TAm3j4nZpX4fvNBvJG7NmZg0i4br8yQ0=; b=bjXwWub5SUn9tZElUByGg2G+CNpH3GRsf49Z2odYznZQdc4o6y+S4ellezdoGsd1Wg 5OL9TLyqeba7uhV4pb98MaCZKdu5p0nPW1w9TSe+bYBrQqHli+JCijOsJtGjW8E9ntEf Yxu7mH6E9dCfoNV9DhBEU0XoDdMbqD/xm+Ixw/O2Rar9QO3aTmLzefYXicgGJ62WsPbl akcxedIDGJIXdeaj8V6Xg877XOQSxgL8RRNuc6OzFS+VDaRuLidcoegQO9C/p5u4HVX7 KVeFErVRF0EeStITSi5VbF2CBPn1EIXWfR1jsFZoYumrg0BBuR2dNhdQnyEg7+PWQ1Yc HDDg== X-Forwarded-Encrypted: i=1; AHgh+RoNqge/tb8ACRUIC6fwxlO+JowNNK1bHqBTNJzsx2NVB4i/zOobNCvgLganih0dHWdbF3py0r068WB8ewU=@vger.kernel.org X-Gm-Message-State: AOJu0YzO8wtwbbw3FV/IuMwQkhCp9AyQnib+eGwxaMFHWaNEpOJNmD7t zMSXu0sKp43Cie+/qCHKo37gu6Yn/QLGA3NRR5O+a2LW/jO8q3HUulTwjcC7/rV0+g== X-Gm-Gg: AfdE7cngUUktLg88rOjWqq1P25hmvoav8vECzJ7Av30jvGMiT2+XzbZ5q/uu6i3Zah5 6McfRix9ZQ7uqi6T3IKyBUDzySAC3FsEaGj23ueUrWaNh9Xr6CzQ5dGWhZdrxZru1SgXWfP0BNs jekHmZ+7ky6/ppXDDup0m2A+WgeDScAs/hfSCKE27tQxNeme6ju9yg1zb0hFopnvFRHe3KX9/7G H7JEVySAxPl348vdmBnWug7yV0SY1zvR2pEc+pgOZEq52jn7Kb5HK06NdFsn68MxOLKpa9rn2dC bW+MnqIE1CKrMBA7Kxr+zEsFWo6BBIRqzf97Idq+fZmCtcSu41praRTdbj6zdRep30wRaGAOiQl xajfTTaFPTrwq9UfXxODz5/h76YANCwU3iC4YyTdxO5I6CeJ6j7wHH4C++W8oRw== X-Received: by 2002:a05:600c:1914:b0:492:7101:3d88 with SMTP id 5b1f17b1804b1-4954a402e03mr75887675e9.24.1784391983856; Sat, 18 Jul 2026 09:26:23 -0700 (PDT) Received: from mf-00-01.. ([194.220.239.180]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-4954a2eddb8sm266444085e9.14.2026.07.18.09.26.22 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 18 Jul 2026 09:26:23 -0700 (PDT) From: Mykola Marzhan To: Jens Axboe , Song Liu , Yu Kuai , Keith Busch , Christoph Hellwig , Sagi Grimberg , linux-block@vger.kernel.org, linux-raid@vger.kernel.org, linux-nvme@lists.infradead.org Cc: Li Nan , Xiao Ni , Leon Romanovsky , Jason Gunthorpe , Kiran Kumar Modukuri , Chaitanya Kulkarni , Logan Gunthorpe , Bjorn Helgaas , Shivaji Kant , Pranjal Shrivastava , Henrique Carvalho , linux-kernel@vger.kernel.org, linux-rdma@vger.kernel.org, linux-pci@vger.kernel.org Subject: [PATCH 4/6] md/raid1,raid10: keep REQ_NOMERGE on narrow_write_error() retry clones Date: Sat, 18 Jul 2026 16:25:45 +0000 Message-ID: <20260718162547.448892-5-mykola@meshstor.io> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260718162547.448892-1-mykola@meshstor.io> References: <20260718162547.448892-1-mykola@meshstor.io> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" narrow_write_error() re-issues a failed write in badblock-granularity chunks, cloning from the master bio and resetting bi_opf to a bare REQ_OP_WRITE. For a P2PDMA bio that reset drops REQ_NOMERGE, which is the only request-level protection against the member queue merging P2PDMA segments across pgmaps or with host memory (see the preceding md_submit_bio() fix): the retry path would quietly reopen the hole the submission path closes. Restore the flag on P2PDMA retry clones. Fixes: 02666132403a ("md: propagate BLK_FEAT_PCI_P2PDMA from member devices= to RAID device") Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Mykola Marzhan --- drivers/md/raid1.c | 3 +++ drivers/md/raid10.c | 3 +++ 2 files changed, 6 insertions(+) diff --git a/drivers/md/raid1.c b/drivers/md/raid1.c index 57f64e890102..a30032321191 100644 --- a/drivers/md/raid1.c +++ b/drivers/md/raid1.c @@ -2565,6 +2565,9 @@ static void narrow_write_error(struct r1bio *r1_bio, = int i) } =20 wbio->bi_opf =3D REQ_OP_WRITE; + /* Keep P2PDMA retry bios unmergeable, like the original */ + if (md_bio_is_p2pdma(wbio)) + wbio->bi_opf |=3D REQ_NOMERGE; wbio->bi_iter.bi_sector =3D r1_bio->sector; wbio->bi_iter.bi_size =3D r1_bio->sectors << 9; =20 diff --git a/drivers/md/raid10.c b/drivers/md/raid10.c index 0a3cfdd3f5df..f7ef903a3d4e 100644 --- a/drivers/md/raid10.c +++ b/drivers/md/raid10.c @@ -2831,6 +2831,9 @@ static void narrow_write_error(struct r10bio *r10_bio= , int i) wbio->bi_iter.bi_sector =3D wsector + choose_data_offset(r10_bio, rdev); wbio->bi_opf =3D REQ_OP_WRITE; + /* Keep P2PDMA retry bios unmergeable, like the original */ + if (md_bio_is_p2pdma(wbio)) + wbio->bi_opf |=3D REQ_NOMERGE; =20 if (submit_bio_wait(wbio) && !rdev_set_badblocks(rdev, wsector, sectors, 0)) { --=20 2.43.0 From nobody Sat Jul 25 04:18:31 2026 Received: from mail-wm1-f52.google.com (mail-wm1-f52.google.com [209.85.128.52]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AD1BA3E022C for ; Sat, 18 Jul 2026 16:26:30 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.52 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784391993; cv=none; b=hPBjTeYg0aB6O6bBkMxUgvxtBhmSPYwsNWb4/slo1kmumHYbIu6X6/ucZ6kq4cbxglCxaVmREuPxHhDm8xla+fwMxyGWumU3cGWtlidIpDRBNJlqlHkKq4xcB3fjDs/EPJXj0UsqwZuk32cwdqTJdTlRXxHggCosar1dKhhlqeg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784391993; c=relaxed/simple; bh=pPYwsyVqfl2OK9H2o/7sp9g85/sHCZY91iO6JM3eDjc=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=fOXNVYKJ46dcgDYYPbjy0ilY41looDRGVGw8SYPoVyFI6M/m30pEVYpXo8OMwSl8d9ybP2W5VihhKOvTuQ93zR99HKyAuf4zc+SRHBABbBkNeTeypCLL5/2Yp65HGiP679dPvVBYmw6r7OiuwrfXx1axOIcTY3zzeN7eL5EJCe0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=meshstor.io; spf=pass smtp.mailfrom=meshstor.io; dkim=pass (2048-bit key) header.d=meshstor.io header.i=@meshstor.io header.b=btTbUjrP; arc=none smtp.client-ip=209.85.128.52 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=meshstor.io Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=meshstor.io Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=meshstor.io header.i=@meshstor.io header.b="btTbUjrP" Received: by mail-wm1-f52.google.com with SMTP id 5b1f17b1804b1-4955484387cso1717585e9.1 for ; Sat, 18 Jul 2026 09:26:30 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=meshstor.io; s=google; t=1784391987; x=1784996787; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=0MziNpD1hbsOVuiqO219eZKFpG+B3dliOI3daA6CJJA=; b=btTbUjrPdF1AVJ8VeHoDv+5va8KY9Ahik2HkJSbcJQxsflWSoTPRZ1v6FYPnL2L01s InFpM8F4PQuSTWX2MJgIyreKQKw1da2pJZHX1Wt590co7m5EZyiqf1MXp2Gj+IFJ4Y5u 9EiEHPx2DFfaD44k5Mi4d5v/E3iVDGKPhMhkAQrwSvasNejP8hTmHLgQC6plK2wkAola mA4RkeXHNEaEa3UYuZ2K1KVEyV0JfHfDshFJ/EY63siNQTkfJTwa7MySg0GDwngWsmLZ MqECF1uDlLQevFYSC+Y7EvhPcNNyDK6496bsqCbrkV2B0Xq1H6/BpH9A0gIov/qRi8YQ z8eg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784391987; x=1784996787; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=0MziNpD1hbsOVuiqO219eZKFpG+B3dliOI3daA6CJJA=; b=MkD0CPLhs1M1ASItWeFEUELqeLC86HVJ2CeDahrM9Ge2cJqQGscrRUeYTmL63EWBBX zhDAjUHgB1lUNyc2UPlrW1Xzm/X3HnZZXOILo2ruLEDw5kCLza3TOwqoHqx9v5UsW7+r fKYCG+W3wfRBuu1vj8gd2ElwIxnTg60Zd4WL/yqknw8/t6FWDanX2kNC9bTV7+VEpz7V EF1HxGPv3fHwNdcDFVwzIxmmwZx/ukmOG0HePgECIU8IcM7jnBR9TxqkzjQ8OQMG8qKU wKSRSLD0SV3EeKhI6MrkDxGVHuhwGNfPFg6wWI3B3RYeWj3IfnY4jHeH+nsnQilwh7sY bMjA== X-Forwarded-Encrypted: i=1; AHgh+Rq1bG53N6+PA7x+6CXRuU9JQ3X+yotg89094fv1qNBoQiGJvchhmEIi4UVX+60w3W6gLIkz+tEUriOKfKI=@vger.kernel.org X-Gm-Message-State: AOJu0Yw/XLb+mPHh4XV+QfzBQLEOSB1PDCMfSXwwH/ZehmWpeZxD0x33 CQaAbU27tecyA/yHd/41wSlsCcTw8XQd4ZmrKSZbhC9YW00z/Cijdlxuqk1ubNYqMw== X-Gm-Gg: AfdE7cn7f0nHN3Z70AAyCAmGjdMGEDtfOsQQM/DkDOd9iD7TNnP/MbGHkYatEwk2Ci3 uySF0xTE/nVCqYVmpeow+sduj64ccb5WGFdFD0w8GiA9Q0f2zfsVG+gHG7ZQMmbtyejREcAp+WY Tpv0bgkqUjLKrZab2ophmJzGuRzGZCpNpgmWTpM4TaN1caTPyrM4g4JQJ3TGpmpg+2pxXQDvqhb Q2sGVTe4u4VgQWaKAANhKxTYx3WpqV4e8I0iZKN2kMt46JB64/toLAGujIpkn0Kbh9nkMaKi1kh z5v9lToEJuPTahO7+gM8OIBfV8dcUVZw8epiPLDH2ks62fc0cK40fykZkJJ31XuLg3289s+SvjI CyrpKcCm0Pae5UwGvv6axR21zGItrX0/AaebaZTf+S901SKYn++pIPQG85gPKMg== X-Received: by 2002:a05:600c:19c8:b0:495:43ab:4f78 with SMTP id 5b1f17b1804b1-4954a3d0ad9mr79922245e9.7.1784391986433; Sat, 18 Jul 2026 09:26:26 -0700 (PDT) Received: from mf-00-01.. ([194.220.239.180]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-4954a2eddb8sm266444085e9.14.2026.07.18.09.26.24 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 18 Jul 2026 09:26:26 -0700 (PDT) From: Mykola Marzhan To: Jens Axboe , Song Liu , Yu Kuai , Keith Busch , Christoph Hellwig , Sagi Grimberg , linux-block@vger.kernel.org, linux-raid@vger.kernel.org, linux-nvme@lists.infradead.org Cc: Li Nan , Xiao Ni , Leon Romanovsky , Jason Gunthorpe , Kiran Kumar Modukuri , Chaitanya Kulkarni , Logan Gunthorpe , Bjorn Helgaas , Shivaji Kant , Pranjal Shrivastava , Henrique Carvalho , linux-kernel@vger.kernel.org, linux-rdma@vger.kernel.org, linux-pci@vger.kernel.org Subject: [PATCH 5/6] md/raid1,raid10: skip futile retries on P2PDMA mapping failures Date: Sat, 18 Jul 2026 16:25:46 +0000 Message-ID: <20260718162547.448892-6-mykola@meshstor.io> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260718162547.448892-1-mykola@meshstor.io> References: <20260718162547.448892-1-mykola@meshstor.io> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Since commit 02666132403a ("md: propagate BLK_FEAT_PCI_P2PDMA from member devices to RAID device") raid1 and raid10 arrays accept P2PDMA (peer device memory) bios, and a member that cannot DMA-map the peer's pages fails its leg bio with BLK_STS_TARGET (from blk_dma_map_iter_start()). That mapping failure is a property of the peer-device/member pairing, not of the medium: re-submitting the same peer pages to the same member cannot succeed, and there is nothing on the disk to repair. Routing it through the stock error machinery misfires on all three paths: - narrow_write_error() re-issues the failed write in badblock-sized chunks, serializing hundreds of guaranteed-to-fail synchronous bios through raid1d/raid10d per failed write (256 for a 1 MiB write on 4K logical blocks, 2048 on 512e). - fix_read_error() probes members with kernel pages, starting with the failing member itself. That probe succeeds (the member is healthy for host memory), so the routine rewrites nothing, records nothing, and logs nothing -- but each invocation costs a full freeze_array() quiesce and a tick of the read-error budget. The budget exists to evict members whose medium keeps producing corrected errors; its hourly decay is sized for sporadic medium errors, while mapping failures are deterministic and arrive at I/O rate. On an asymmetric PCIe topology (peer reaches one member but not the other) a mixed host/P2P read workload charges 20 errors to the unreachable leg within a second and kicks a perfectly healthy member -- destroying redundancy for all I/O because some I/O cannot route, while every P2P read is meanwhile served correctly by the reachable mirror. - On FailFast members both paths short-circuit into md_error(), so a single unroutable peer-memory I/O evicts a healthy mirror outright. Handle the mapping failure -- BLK_STS_TARGET on a bio carrying PCI P2PDMA pages -- explicitly: - Writes: probe the whole range with one retry and record one bad range if it also fails. BLK_STS_TARGET is also produced for device conditions that narrow_write_error()'s retry loop recovers from (e.g. NVME_SC_NS_NOT_READY and NVME_SC_CMD_INTERRUPTED via nvme_error_status()), and md cannot tell the two apart from bi_status alone, so it must not skip the retry outright. A single whole-range retry observes the outcome instead of predicting it: a mapping failure fails deterministically and is recorded in one call, collapsing the chunked storm to a single bio, while a transient error that has since cleared succeeds and records nothing, exactly as the chunked loop behaves today. - Reads: mark the leg IO_BLOCKED so read_balance() picks another mirror, and skip the freeze/fix cycle and the budget charge. If no mirror can serve the read the master bio fails with EIO as before, now without kicking healthy members on the way. The P2P page check is what keeps this narrow: a member that keeps failing with a device-produced BLK_STS_TARGET (e.g. stuck not-ready) must still be evictable, so the status alone cannot gate the skip. - FailFast: a mapping failure is rejected at map time and never reaches the wire, so it is no evidence of device unreliability; don't let it trigger the FailFast md_error() -- the same reasoning commit f7b24c7b41f2 ("md/raid1,raid10: don't fail devices for invalid IO errors") applied to BLK_STS_INVAL. The first bvec is representative of the whole bio: the block layer never mixes P2P pages from different pgmaps with each other or with host pages in one bio (zone_device_pages_compatible()), and reads bi_io_vec[0] the same way in bio_iov_iter_get_pages(). md's P2PDMA support is otherwise declared at array setup time through queue limits, but whether one leg's completion was a mapping refusal is a property of that bio alone, so the completion path has to look at the bio. A distinct block status naming the mapping failure exactly would remove the remaining ambiguity against device-produced BLK_STS_TARGET and is left as a follow-up. This belongs in the same release as the BLK_STS_TARGET restoration: with that fix alone, an unroutable leg costs the retry storms and the healthy-member evictions above. Measured on QEMU q35 rigs (one CMB provider, two NVMe members), 8KiB peer-memory I/O, reads x90 under concurrent host reads: scenario before (patches 1-4) after (this patch) asym reads 90/90 ok, healthy far 90/90 ok, no eviction, leg kicked after ~20 no budget charge unreach reads all EIO, one healthy all EIO, no members member kicked kicked p2p TARGET 16 chunk retries from 1 whole-range probe, write raid1d, then badblocks same badblocks, master ok A transient TARGET recovers through the single probe with no badblocks recorded, and a mapping failure no longer trips FailFast eviction while a genuine I/O error still does -- both verified on the same rig. Fixes: 02666132403a ("md: propagate BLK_FEAT_PCI_P2PDMA from member devices= to RAID device") Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Mykola Marzhan --- drivers/md/md.h | 13 +++++++++++ drivers/md/raid1.c | 46 +++++++++++++++++++++++++++++++----- drivers/md/raid10.c | 57 ++++++++++++++++++++++++++++++++++++++------- 3 files changed, 101 insertions(+), 15 deletions(-) diff --git a/drivers/md/md.h b/drivers/md/md.h index 140e2b3670d8..b1abfc89cfc0 100644 --- a/drivers/md/md.h +++ b/drivers/md/md.h @@ -40,6 +40,19 @@ static inline bool md_bio_is_p2pdma(struct bio *bio) is_pci_p2pdma_page(bio->bi_io_vec->bv_page); } =20 +/* + * True when a leg bio failed because its P2PDMA pages cannot be DMA-mapped + * to this member (BLK_STS_TARGET from blk_dma_map_iter_start()). Usable at + * completion time, unlike md_bio_is_p2pdma(): leg bios are clones sharing + * the master bio's bvec table, which outlives the leg's end_io, but their + * iterator is consumed by then, so bio_has_data() cannot gate the access. + */ +static inline bool md_bio_p2pdma_mapping_error(struct bio *bio) +{ + return bio->bi_status =3D=3D BLK_STS_TARGET && bio->bi_io_vec && + is_pci_p2pdma_page(bio->bi_io_vec->bv_page); +} + /* * Number of guaranteed raid bios in case of extreme VM load: */ diff --git a/drivers/md/raid1.c b/drivers/md/raid1.c index a30032321191..241c2994b6a3 100644 --- a/drivers/md/raid1.c +++ b/drivers/md/raid1.c @@ -491,7 +491,9 @@ static void raid1_end_write_request(struct bio *bio) if (test_bit(FailFast, &rdev->flags) && (bio->bi_opf & MD_FAILFAST) && /* We never try FailFast to WriteMostly devices */ - !test_bit(WriteMostly, &rdev->flags)) { + !test_bit(WriteMostly, &rdev->flags) && + /* A mapping failure says nothing about device health */ + !md_bio_p2pdma_mapping_error(bio)) { md_error(r1_bio->mddev, rdev); } =20 @@ -2517,7 +2519,7 @@ static void fix_read_error(struct r1conf *conf, struc= t r1bio *r1_bio) } } =20 -static void narrow_write_error(struct r1bio *r1_bio, int i) +static void narrow_write_error(struct r1bio *r1_bio, int i, bool coarse) { struct mddev *mddev =3D r1_bio->mddev; struct r1conf *conf =3D mddev->private; @@ -2531,6 +2533,12 @@ static void narrow_write_error(struct r1bio *r1_bio,= int i) * It is conceivable that the bio doesn't exactly align with * blocks. We must handle this somehow. * + * With 'coarse', retry the whole range as one bio and record + * one bad range if it fails: used when per-block narrowing + * cannot find a good block (P2PDMA mapping failures fail the + * whole range identically), while a single retry still tells + * a since-cleared transient error apart. + * * We currently own a reference on the rdev. */ =20 @@ -2545,9 +2553,12 @@ static void narrow_write_error(struct r1bio *r1_bio,= int i) block_sectors =3D roundup(1 << rdev->badblocks.shift, lbs); =20 sector =3D r1_bio->sector; - sectors =3D ((sector + block_sectors) - & ~(sector_t)(block_sectors - 1)) - - sector; + if (coarse) + sectors =3D sect_to_write; + else + sectors =3D ((sector + block_sectors) + & ~(sector_t)(block_sectors - 1)) + - sector; =20 while (sect_to_write) { struct bio *wbio; @@ -2629,7 +2640,20 @@ static void handle_write_finished(struct r1conf *con= f, struct r1bio *r1_bio) * errors. */ fail =3D true; - narrow_write_error(r1_bio, m); + if (md_bio_p2pdma_mapping_error(r1_bio->bios[m])) + /* + * A P2PDMA mapping failure fails the whole + * range identically, so narrowing block by + * block cannot find a good block -- but a + * transient device error also surfaces as + * BLK_STS_TARGET, so don't assume. Retry + * the range once: if it fails, record it in + * one go; if it succeeds, there was nothing + * wrong with the medium. + */ + narrow_write_error(r1_bio, m, true); + else + narrow_write_error(r1_bio, m, false); rdev_dec_pending(conf->mirrors[m].rdev, conf->mddev); } @@ -2656,6 +2680,7 @@ static void handle_read_error(struct r1conf *conf, st= ruct r1bio *r1_bio) { struct md_rdev *rdev =3D conf->mirrors[r1_bio->read_disk].rdev; struct bio *bio =3D r1_bio->bios[r1_bio->read_disk]; + bool p2pdma_error =3D md_bio_p2pdma_mapping_error(bio); struct mddev *mddev =3D conf->mddev; sector_t sector; =20 @@ -2675,6 +2700,15 @@ static void handle_read_error(struct r1conf *conf, s= truct r1bio *r1_bio) */ if (mddev->ro) { r1_bio->bios[r1_bio->read_disk] =3D IO_BLOCKED; + } else if (p2pdma_error) { + /* + * The peer pages cannot be DMA-mapped to this member; + * there is nothing to fix on the medium and the member + * is healthy for host I/O: don't charge the read-error + * budget or fail a FailFast member, just keep this leg + * out of the retry. + */ + r1_bio->bios[r1_bio->read_disk] =3D IO_BLOCKED; } else if (test_bit(FailFast, &rdev->flags)) { md_error(mddev, rdev); } else { diff --git a/drivers/md/raid10.c b/drivers/md/raid10.c index f7ef903a3d4e..b24429a19254 100644 --- a/drivers/md/raid10.c +++ b/drivers/md/raid10.c @@ -489,7 +489,9 @@ static void raid10_end_write_request(struct bio *bio) =20 dec_rdev =3D 0; if (test_bit(FailFast, &rdev->flags) && - (bio->bi_opf & MD_FAILFAST)) { + (bio->bi_opf & MD_FAILFAST) && + /* A mapping failure says nothing about device health */ + !md_bio_p2pdma_mapping_error(bio)) { md_error(rdev->mddev, rdev); } =20 @@ -2786,7 +2788,7 @@ static void fix_read_error(struct r10conf *conf, stru= ct mddev *mddev, struct r10 } } =20 -static void narrow_write_error(struct r10bio *r10_bio, int i) +static void narrow_write_error(struct r10bio *r10_bio, int i, bool coarse) { struct bio *bio =3D r10_bio->master_bio; struct mddev *mddev =3D r10_bio->mddev; @@ -2800,6 +2802,12 @@ static void narrow_write_error(struct r10bio *r10_bi= o, int i) * It is conceivable that the bio doesn't exactly align with * blocks. We must handle this. * + * With 'coarse', retry the whole range as one bio and record + * one bad range if it fails: used when per-block narrowing + * cannot find a good block (P2PDMA mapping failures fail the + * whole range identically), while a single retry still tells + * a since-cleared transient error apart. + * * We currently own a reference to the rdev. */ =20 @@ -2814,9 +2822,12 @@ static void narrow_write_error(struct r10bio *r10_bi= o, int i) block_sectors =3D roundup(1 << rdev->badblocks.shift, lbs); =20 sector =3D r10_bio->sector; - sectors =3D ((r10_bio->sector + block_sectors) - & ~(sector_t)(block_sectors - 1)) - - sector; + if (coarse) + sectors =3D sect_to_write; + else + sectors =3D ((r10_bio->sector + block_sectors) + & ~(sector_t)(block_sectors - 1)) + - sector; =20 while (sect_to_write) { struct bio *wbio; @@ -2856,6 +2867,7 @@ static void handle_read_error(struct mddev *mddev, st= ruct r10bio *r10_bio) { int slot =3D r10_bio->read_slot; struct bio *bio; + bool p2pdma_error; struct r10conf *conf =3D mddev->private; struct md_rdev *rdev =3D r10_bio->devs[slot].rdev; =20 @@ -2868,17 +2880,28 @@ static void handle_read_error(struct mddev *mddev, = struct r10bio *r10_bio) * frozen. */ bio =3D r10_bio->devs[slot].bio; + p2pdma_error =3D md_bio_p2pdma_mapping_error(bio); bio_put(bio); r10_bio->devs[slot].bio =3D NULL; =20 if (mddev->ro) r10_bio->devs[slot].bio =3D IO_BLOCKED; - else if (!test_bit(FailFast, &rdev->flags)) { + else if (p2pdma_error) { + /* + * The peer pages cannot be DMA-mapped to this member; + * there is nothing to fix on the medium and the member + * is healthy for host I/O: don't charge the read-error + * budget or fail a FailFast member, just keep this leg + * out of the retry. + */ + r10_bio->devs[slot].bio =3D IO_BLOCKED; + } else if (test_bit(FailFast, &rdev->flags)) { + md_error(mddev, rdev); + } else { freeze_array(conf, 1); fix_read_error(conf, mddev, r10_bio); unfreeze_array(conf); - } else - md_error(mddev, rdev); + } =20 rdev_dec_pending(rdev, mddev); r10_bio->state =3D 0; @@ -2948,7 +2971,23 @@ static void handle_write_completed(struct r10conf *c= onf, struct r10bio *r10_bio) rdev_dec_pending(rdev, conf->mddev); } else if (bio !=3D NULL && bio->bi_status) { fail =3D true; - narrow_write_error(r10_bio, m); + if (md_bio_p2pdma_mapping_error(bio)) + /* + * A P2PDMA mapping failure fails + * the whole range identically, so + * narrowing block by block cannot + * find a good block -- but a + * transient device error also + * surfaces as BLK_STS_TARGET, so + * don't assume. Retry the range + * once: if it fails, record it in + * one go; if it succeeds, there + * was nothing wrong with the + * medium. + */ + narrow_write_error(r10_bio, m, true); + else + narrow_write_error(r10_bio, m, false); rdev_dec_pending(rdev, conf->mddev); } bio =3D r10_bio->devs[m].repl_bio; --=20 2.43.0 From nobody Sat Jul 25 04:18:31 2026 Received: from mail-wr1-f53.google.com (mail-wr1-f53.google.com [209.85.221.53]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B391F3E0233 for ; Sat, 18 Jul 2026 16:26:30 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.221.53 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784391994; cv=none; b=YCEhPocLfHbumWkOLROuGeffEfg1FIo+ORCnUAyrXs0si0fDEiItgItFMnnozI7dGUqfSlUbZBIVDgmtEHbXTL6k9dLJeW+yeSbp6jD0zNWfIh+pf07b/iRLmZ36HbfR/jeiz8EX3ZgUKDK0grvy77T+l0k3F0qkSqmF4eWKufA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784391994; c=relaxed/simple; bh=JHE3cgZBQuIzT0YLZJnapzbT3KS4Yu6cNSzhz+WZwJs=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=UI4ifmfjkkkMHgzATWh8fgaRu+tKu4BpSORDJs8TDU8z3T1/CX6PXVxWr4WO83TVQrjhfDTjz3etMVS3uiBoijcwmFud1SkanwUoZcmBy4yzyFGVvhbZKa45RvyEkLHm/HtCAF+Bv6u5n7hIHdD5RC42TqBXkDYHKuM6QpmNNZo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=meshstor.io; spf=pass smtp.mailfrom=meshstor.io; dkim=pass (2048-bit key) header.d=meshstor.io header.i=@meshstor.io header.b=Ik4jzfuI; arc=none smtp.client-ip=209.85.221.53 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=meshstor.io Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=meshstor.io Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=meshstor.io header.i=@meshstor.io header.b="Ik4jzfuI" Received: by mail-wr1-f53.google.com with SMTP id ffacd0b85a97d-4798bea72f9so4564311f8f.1 for ; Sat, 18 Jul 2026 09:26:30 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=meshstor.io; s=google; t=1784391989; x=1784996789; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=dagmucy3ghWM9/yf2A3Q3PwS7L9VrlH2jBuQsiQB1xw=; b=Ik4jzfuI2OC3tJDIWdBrOtKEGBnKFGa+39d+vEvYcxV04LRRmwPN20ctAHgRVj1niX EiV9aIEQ2mLaPV372bAo0xSOF2QXE9Yszg2/FXVRQ4KAETSuVNx8L58Des6aBUOg05Vo EgyPY3X28/yAIy+cxvUNZEV9aGfsZkxJfExXiiKglWawDqmy+9x2UZrnRBjXq0YYNNeF BegFi4GAdcuB1Th51AELhSGM8lrtlT1f80vyMGCDMFn9hPnibOO+ErScPkZAQEibNwL1 1cNgVx/jDGOTUZ3VleHvTVRtPLEfOkBUIJ3VVJxm2BqZSGEMLMq5jRYreAlT3xQthuuE 3NmQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784391989; x=1784996789; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=dagmucy3ghWM9/yf2A3Q3PwS7L9VrlH2jBuQsiQB1xw=; b=kC8h1LfVP8wCDhRTZFLFqiLpZMg30mSuDpXVN2+w5MNGmSUB2ZC3hCpF3Ht7d7lF9j oht4FrITZTlIW850XqmjnyABvDmg1XizMAKvO2R2uFytjH9gfFA6Ah3lM6IHFVzKO/OQ qXZJiyIUjJznGFuRbkp1si9gzwH8StX4jGKDnUwjVC77b4whoblRvid4PPjQkbwtQnLK vDFKMNPVkrCXJND/ICVl7mU6y2ZoRQaSiygC8X5bIGtJ/mpiVkUPa40Mge+E/iDlmJ5V W1yWCMQ/2JWmlsjpg4Du2swMYAK3oH7nxKO+lebllPA6P79ah66sN/qkfgCVyuUmYZIs 3AdA== X-Forwarded-Encrypted: i=1; AHgh+RraBKESO9T0gfUUcBSNY7/JM4HvHV2wUQDE7arMQ3FiGFWLK6ZzKRv29Tv6bX7zQ08vI+zd/3n3FenE3Xw=@vger.kernel.org X-Gm-Message-State: AOJu0YxYtmkpxXNiWRWqSDqKp3R8636XSwst2hu3YzsS02BZCYCu20hP tbd2PYdAJDlJ6+CBcKJtFXXVbD5Gn/jB8gNLLusVtAPtyNCCHZ1kLkdwioxQ11UW1w== X-Gm-Gg: AfdE7cmeKOEems5q6gEZdWnDoKaC3vKXaaZjSqjAs/TF1lfwLgoo5SvryXfeRA6Q+9k GZHVhwPITJQbiXhsXagf/P+iBtEGOtqnm6KdTBmtfITvpeQP7hZ2iDgnX0Bii4eaTBV65KGOHQZ +6r6c+4/lJChrDIMtIvrckHFVnW+qpMS2marwuCES+cnAMxBvyaHaVtTTeFpDCK+6kLhrMvHmGW fkH/b5Tt7Nk5OrKm8810orr8Wqg4RKNuz671eW5S6oc9iLWIUCa/bFkJwZ7sri37duSLbNPz7h3 6SnQV/cpd7Vs38GhGlhm+iCIxznD3VbQRBHuchRggR3bM1QN+Sh+ngu0baO6OL7b+aPEXJmJucd W10k4HjdHL+j1BxUKczsIAgHZ0fo9FcFyjUEryEdtQgk5aJ/iFeCNDrpbulnJvQ== X-Received: by 2002:a05:600c:1549:b0:493:c601:3e23 with SMTP id 5b1f17b1804b1-4954a3d08c0mr76801375e9.5.1784391988987; Sat, 18 Jul 2026 09:26:28 -0700 (PDT) Received: from mf-00-01.. ([194.220.239.180]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-4954a2eddb8sm266444085e9.14.2026.07.18.09.26.27 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 18 Jul 2026 09:26:28 -0700 (PDT) From: Mykola Marzhan To: Jens Axboe , Song Liu , Yu Kuai , Keith Busch , Christoph Hellwig , Sagi Grimberg , linux-block@vger.kernel.org, linux-raid@vger.kernel.org, linux-nvme@lists.infradead.org Cc: Li Nan , Xiao Ni , Leon Romanovsky , Jason Gunthorpe , Kiran Kumar Modukuri , Chaitanya Kulkarni , Logan Gunthorpe , Bjorn Helgaas , Shivaji Kant , Pranjal Shrivastava , Henrique Carvalho , linux-kernel@vger.kernel.org, linux-rdma@vger.kernel.org, linux-pci@vger.kernel.org Subject: [PATCH 6/6] nvme-rdma: return BLK_STS_TARGET for unsupported P2P transfers Date: Sat, 18 Jul 2026 16:25:47 +0000 Message-ID: <20260718162547.448892-7-mykola@meshstor.io> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260718162547.448892-1-mykola@meshstor.io> References: <20260718162547.448892-1-mykola@meshstor.io> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Since commit 23528aa3320a ("nvme: enable PCI P2PDMA support for RDMA transport") nvme-rdma accepts P2PDMA bios, but a mapping failure for peer memory the HCA cannot reach is misreported as a path error: ib_dma_map_sg() returns 0, discarding the -EREMOTEIO that dma_map_sgtable() documents for exactly this case, and the driver converts it to -EIO -> nvme_host_path_error(). Under the default multipath configuration (the multipath head node advertises BLK_FEAT_PCI_P2PDMA since commit fb0eeeed91f3 ("nvme-multipath: enable PCI P2PDMA for multipath devices")) the path-error status makes nvme_failover_req() requeue the bios with a fresh retry budget each cycle, and since the mapping failure is a deterministic property of the peer/device pairing the I/O simply never completes: a hot requeue livelock, and a stacked md mirror hangs instead of failing over to its other leg. With nvme_core.multipath=3DN the request burns nvme_max_retries requeues -- nothing ever reaching the wire -- and completes as BLK_STS_TRANSPORT, which blk_path_error() classifies as retryable and md/raid1,raid10 treat as a genuine device error: retry storms on writes, read-error-budget eviction of a healthy member on reads (see the preceding md patches, whose mapping-failure handling keys on BLK_STS_TARGET). Map the data scatterlist with ib_dma_map_sgtable_attrs() so the DMA layer's error code is preserved, and translate -EREMOTEIO to BLK_STS_TARGET: the classification nvme-pci established in commit 91fb2b6052f7 ("nvme-pci: convert to using dma_map_sgtable()") and the preceding blk-mq-dma patch restores (the blk_rq_dma_map conversion had changed it to BLK_STS_INVAL). The failure is a property of the peer/device pairing, so it must not be retried on this or another path. Start the request only after mapping succeeds, as nvme-pci does: a mapping failure now errors out of queue_rq on a not-yet-started request, so nvme_mpath_start_request() accounting is never taken and cannot leak on the direct blk-mq completion (this also closes the same latent leak for the existing BLK_STS_IOERR returns). A mapping -ENOMEM now takes the existing BLK_STS_RESOURCE branch instead of masquerading as a path error, and a DMA-layer -EINVAL completes as BLK_STS_IOERR instead of a path error; generic -EIO keeps today's host-path-error behavior, and virt-DMA devices are unaffected. The metadata scatterlist keeps ib_dma_map_sg(): integrity buffers are host memory and cannot produce -EREMOTEIO. Ratelimit the map-failure message -- with an md mirror steering peer-memory I/O around an unreachable leg it fires per redirected I/O, not per rare event. Fixes: 23528aa3320a ("nvme: enable PCI P2PDMA support for RDMA transport") Cc: stable@vger.kernel.org # v7.1 Assisted-by: Claude:claude-fable-5 Signed-off-by: Mykola Marzhan --- drivers/nvme/host/rdma.c | 26 +++++++++++++++++--------- 1 file changed, 17 insertions(+), 9 deletions(-) diff --git a/drivers/nvme/host/rdma.c b/drivers/nvme/host/rdma.c index 6909e3542794..b8642cd2fb79 100644 --- a/drivers/nvme/host/rdma.c +++ b/drivers/nvme/host/rdma.c @@ -1469,6 +1469,7 @@ static int nvme_rdma_dma_map_req(struct ib_device *ib= dev, struct request *rq, int *count, int *pi_count) { struct nvme_rdma_request *req =3D blk_mq_rq_to_pdu(rq); + struct sg_table sgt; int ret; =20 req->data_sgl.sg_table.sgl =3D (struct scatterlist *)(req + 1); @@ -1480,12 +1481,12 @@ static int nvme_rdma_dma_map_req(struct ib_device *= ibdev, struct request *rq, =20 req->data_sgl.nents =3D blk_rq_map_sg(rq, req->data_sgl.sg_table.sgl); =20 - *count =3D ib_dma_map_sg(ibdev, req->data_sgl.sg_table.sgl, - req->data_sgl.nents, rq_dma_dir(rq)); - if (unlikely(*count <=3D 0)) { - ret =3D -EIO; + sgt.sgl =3D req->data_sgl.sg_table.sgl; + sgt.orig_nents =3D req->data_sgl.nents; + ret =3D ib_dma_map_sgtable_attrs(ibdev, &sgt, rq_dma_dir(rq), 0); + if (unlikely(ret)) goto out_free_table; - } + *count =3D sgt.nents; =20 if (blk_integrity_rq(rq)) { req->metadata_sgl->sg_table.sgl =3D @@ -2026,8 +2027,6 @@ static blk_status_t nvme_rdma_queue_rq(struct blk_mq_= hw_ctx *hctx, if (ret) goto unmap_qe; =20 - nvme_start_request(rq); - if (IS_ENABLED(CONFIG_BLK_DEV_INTEGRITY) && queue->pi_support && (c->common.opcode =3D=3D nvme_cmd_write || @@ -2039,11 +2038,13 @@ static blk_status_t nvme_rdma_queue_rq(struct blk_m= q_hw_ctx *hctx, =20 err =3D nvme_rdma_map_data(queue, rq, c); if (unlikely(err < 0)) { - dev_err(queue->ctrl->ctrl.device, - "Failed to map data (%d)\n", err); + dev_err_ratelimited(queue->ctrl->ctrl.device, + "Failed to map data (%d)\n", err); goto err; } =20 + nvme_start_request(rq); + sqe->cqe.done =3D nvme_rdma_send_done; =20 ib_dma_sync_single_for_device(dev, sqe->dma, @@ -2063,6 +2064,13 @@ static blk_status_t nvme_rdma_queue_rq(struct blk_mq= _hw_ctx *hctx, ret =3D nvme_host_path_error(rq); else if (err =3D=3D -ENOMEM || err =3D=3D -EAGAIN) ret =3D BLK_STS_RESOURCE; + /* + * The DMA layer refused to map peer memory to this device: a + * property of the pairing, not a path failure. Match nvme-pci + * and do not retry (see blk_path_error()). + */ + else if (err =3D=3D -EREMOTEIO) + ret =3D BLK_STS_TARGET; else ret =3D BLK_STS_IOERR; nvme_cleanup_cmd(rq); --=20 2.43.0