From nobody Sat Jul 25 03:47:43 2026 Received: from mail-wm1-f47.google.com (mail-wm1-f47.google.com [209.85.128.47]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7E6E439150D for ; Sun, 19 Jul 2026 10:53:38 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.47 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784458421; cv=none; b=eMxub8lxPpjpoU61d139dA3fzMPJWhAD+LUxekyzGnR7eD+tr7NTc0/tukzLtPWQnzO5WXbobdlFxRpl7xbIEx8y11dnJTQlkMEbaozceVNtVfcRjU9EdxxhV2dAguJQCyoPyRZ3Ixtzv3ZnIV+CwIw1ssFkmWnWyVJqkNZ/qEw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784458421; c=relaxed/simple; bh=I0++bTZ6dP7SwW+DMBEsuDxoAUGip3+1I9WU1guq83I=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=sEdJbMjlKMTC0Cl3pdWdfW6BAmRuCiAyJPkuaBliXRYqxXSxPvjP5ReoPIrPyKkctKxuPUhsc632Vzv6VUs+KJOXGGn9PMW//t3+wBsn2gSh1OhtSOqRSYjPSZiVdmOCt/i1HgrspAiVF1Xog5Ul2+lGHdF0ZXZ0hCmmxB38u3A= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=meshstor.io; spf=pass smtp.mailfrom=meshstor.io; dkim=pass (2048-bit key) header.d=meshstor.io header.i=@meshstor.io header.b=M0TUPaja; arc=none smtp.client-ip=209.85.128.47 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=meshstor.io Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=meshstor.io Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=meshstor.io header.i=@meshstor.io header.b="M0TUPaja" Received: by mail-wm1-f47.google.com with SMTP id 5b1f17b1804b1-4955aa106b1so2252265e9.0 for ; Sun, 19 Jul 2026 03:53:38 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=meshstor.io; s=google; t=1784458417; x=1785063217; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=hybfSCWaTjres2ZZlI9Lu6CV5Kqe33y6xILjvKI8G00=; b=M0TUPajadEG7LHHvuhNik2cKvB9QyqXj/5nvWQcAIOmHx8E0g1cZAfgZreCmzneDyg 9G0AwF7BSiHMJ0P1gY+yGp+2omGaaVJmBDb7alJtSiX01mqMtsfHKohUBxRzuE9bqjSb eAgdCAZeHcTlIXJC40LXL5ewCCPIUePXGTVfjPh7Je+v2K1NrhQnC6Z8NVcAGLKcc0kK clSeAw1o6Vfo+lqbE5YRphPOpubboNdK0PhtyqGKJbL4dtXd2IQaVCrZRyt0gqB3CVEl p4ouF6AFBdyf05SEZNQT4H0s2Y2BwsfN0yIbIdZgX3AkuBE9jOCeHcqyZb1BZ9DDtqOT x/7w== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784458417; x=1785063217; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=hybfSCWaTjres2ZZlI9Lu6CV5Kqe33y6xILjvKI8G00=; b=hyNwxuHalEe0ksE+FekwqqQpc4IdUiONSIQU5M5pNP/9ox4BUlMaTZujJWg1lDTBW+ MhgpGLWBm5M7c/wkuLwkmxVI8crYL0phXY6cHXAAQdZ+lx9k08TZNJF6VNtEGyILfhTC CzwkkeapEtn2U8JRW3YWnST8jCRJljs3c4upevbvms5C5vT9g0MlxzvFXtekETWExdxJ jF41vJeeJWCvX/CNJLOmaHbfgxD4XlOwLA/ytAFjqxymVCornUk/r3CCUAkX77+VDSHs McUXxQLcSebfnklQ5rnTaSLQuRtVIEqgd2V7U5//hbexoRE2O34ONaBcQt+lEeDYkfpx 6pOQ== X-Forwarded-Encrypted: i=1; AHgh+RoYhpu1zi3C2wjQS1t8n1YkjmsUSwASkHM977ULxzMqZCW/jf4aMQKmsVdGIWSsQBS3CE7eYP4MwQ7UMp0=@vger.kernel.org X-Gm-Message-State: AOJu0YwtvvSau6EGUiwEGuhRWlhVSVyEQKp2mXd/Pw145gmu/ejtHPCF PA+9Y7ILnwOLy9mz4SE05sYMOqrUfj0sKk/g9u5p+HFmBo2n+NDDGnlnR1uumCjDlw== X-Gm-Gg: AfdE7cnp6YAbkdFXkBsLb+X1O56ACm4BltppdS/Opfc8bSomKrg3hkYm3sO1QpNMwMH eEC3p+c7MBkStyyUskLdTVaBi33VtZLrET5p/NCyBbHZjbxO2kfNRxIosTiAHLZxQdMqJxDQM69 1gotUSuAfT4w6Rt98LU2SilhH1+Q6uYnpGpP7JXbPDM+ZZcI3DkNEBP7fR8W+e5vT7AMK9koSXH bqu2cLMuplsy5bNfne03mgShtm3XJrQ0WWRuorMIFwL8K4VEi29ix3mYg/lQrCc2XkLAq24cpWQ ZBch1SSsiHGIFJkAqzE0/rgQkiDARdogLV5BgZtVLQ6/B2cwqcj2/0vOkZUCxA2Z5YaBty8l9GK dNICaojnAoabe8fXEWYHL6zcZNFBacEvXHNzlHMc6EeGe3dql/xlPQiEiT1WDPA== X-Received: by 2002:a05:600c:e54a:10b0:490:5057:f5f7 with SMTP id 5b1f17b1804b1-4954a3ed516mr73355155e9.11.1784458416521; Sun, 19 Jul 2026 03:53:36 -0700 (PDT) Received: from mf-00-01.. ([194.220.239.180]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-4954a2a24b6sm197679575e9.3.2026.07.19.03.53.34 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sun, 19 Jul 2026 03:53:35 -0700 (PDT) From: Mykola Marzhan To: Jens Axboe , Song Liu , Yu Kuai , Keith Busch , Christoph Hellwig , Sagi Grimberg , linux-block@vger.kernel.org, linux-raid@vger.kernel.org, linux-nvme@lists.infradead.org Cc: Li Nan , Xiao Ni , Guoqing Jiang , Leon Romanovsky , Jason Gunthorpe , Kiran Kumar Modukuri , Chaitanya Kulkarni , Logan Gunthorpe , Bjorn Helgaas , Shivaji Kant , Pranjal Shrivastava , Henrique Carvalho , linux-kernel@vger.kernel.org, linux-rdma@vger.kernel.org, linux-pci@vger.kernel.org Subject: [PATCH v2 1/7] blk-mq-dma: restore BLK_STS_TARGET for unsupported P2P transfers Date: Sun, 19 Jul 2026 10:53:21 +0000 Message-ID: <20260719105327.864949-2-mykola@meshstor.io> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260719105327.864949-1-mykola@meshstor.io> References: <20260719105327.864949-1-mykola@meshstor.io> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Commit 91fb2b6052f7 ("nvme-pci: convert to using dma_map_sgtable()") deliberately mapped unsupported P2PDMA transfers to BLK_STS_TARGET, matching dma_map_sgtable()'s -EREMOTEIO: "When this happens, return BLK_STS_TARGET so the request isn't retried." The conversion to blk_rq_dma_map silently changed the status to BLK_STS_INVAL, which regresses two consumers: - md/raid1 and raid10 ignore BLK_STS_INVAL leg failures (commit f7b24c7b41f2 ("md/raid1,raid10: don't fail devices for invalid IO errors"), where it means a request-shaped error that fails identically on every member). Since commit 02666132403a ("md: propagate BLK_FEAT_PCI_P2PDMA from member devices to RAID device") P2PDMA bios reach md arrays: a peer-memory write to a member the peer cannot reach is counted as written, the master bio reports success, mirrors silently diverge, and on a topology where no member is reachable the write reports success with zero copies on stable storage. - the failure's classification flips: for direct NVMe consumers (nvme advertises BLK_FEAT_PCI_P2PDMA) the mapping failure has surfaced as EINVAL instead of the documented -EREMOTEIO since v6.17, and blk_path_error(BLK_STS_INVAL) is true, so a stacking consumer would treat it as a retryable path error (dm-mpath is the only in-tree blk_path_error() caller; P2P bios cannot currently reach it, but the classification is wrong on its face). Restore BLK_STS_TARGET. md then routes an unreachable leg through its per-device error handling (badblocks, mirror retry for reads); a later patch in this series teaches raid1/raid10 to handle mapping failures without the retry storms that machinery was built around. Fixes: 858299dc6160 ("block: add scatterlist-less DMA mapping helpers") Fixes: 7ce3c1dd78fc ("nvme-pci: convert the data mapping to blk_rq_dma_map") Cc: stable@vger.kernel.org # v6.17 Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Mykola Marzhan --- block/blk-mq-dma.c | 10 +++++++++- 1 file changed, 9 insertions(+), 1 deletion(-) diff --git a/block/blk-mq-dma.c b/block/blk-mq-dma.c index bfdb9ed70741..2eed06bfe791 100644 --- a/block/blk-mq-dma.c +++ b/block/blk-mq-dma.c @@ -190,7 +190,15 @@ static bool blk_dma_map_iter_start(struct request *req= , struct device *dma_dev, case PCI_P2PDMA_MAP_NONE: break; default: - iter->status =3D BLK_STS_INVAL; + /* + * P2P transfers that the mapping layer cannot support + * report BLK_STS_TARGET, matching dma_map_sgtable()'s + * -EREMOTEIO and the pre-blk_rq_dma_map nvme behavior: + * the failure is a property of this device pairing, so + * it must not be retried on another path (blk_path_error) + * nor be mistaken for an invalid request. + */ + iter->status =3D BLK_STS_TARGET; return false; } =20 --=20 2.43.0 From nobody Sat Jul 25 03:47:43 2026 Received: from mail-wm1-f53.google.com (mail-wm1-f53.google.com [209.85.128.53]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7C6CF3921DB for ; Sun, 19 Jul 2026 10:53:40 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.53 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784458422; cv=none; b=hZz4BLx5xeAltFyTw/y255R8PqDMJnRXPqncBlkXkPTDlCioJGn57IDXcMnshMqJMirdoLQf0KQbx4fCQtJN4irxHglrPXJP/0PiHIXVK1COslzaAAlAhw3h5GTrfNj2HEMRmzPop2QlyH3Ls32kWbzamaYz9UXOW1GW71a7q0k= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784458422; c=relaxed/simple; bh=vO1jpq1MeVozqKW9Fk/0t1VAaeJE5a63mWcH5DQ3WaQ=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=sdPryJZOUYsMDzykP2eAGeKnpZ+/9it5caxcovyhPmP4TLUKWDrwmUcocCT4H22fyJiSwwFALV3k2QIFs9LIml8ATFgjLhuqsflx54AWNCTNFs0jBAJuvFEYcI0MEfnQy1zecWs90dVW00X6ne1PTYBGrFu2jK3ptQDmH/0dkQ0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=meshstor.io; spf=pass smtp.mailfrom=meshstor.io; dkim=pass (2048-bit key) header.d=meshstor.io header.i=@meshstor.io header.b=TMywT7bt; arc=none smtp.client-ip=209.85.128.53 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=meshstor.io Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=meshstor.io Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=meshstor.io header.i=@meshstor.io header.b="TMywT7bt" Received: by mail-wm1-f53.google.com with SMTP id 5b1f17b1804b1-49550ec592cso3992095e9.0 for ; Sun, 19 Jul 2026 03:53:40 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=meshstor.io; s=google; t=1784458418; x=1785063218; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=FhGxTEOOON4RkKT9KdpeyXrdqn2SrMJKFWwPDsn11x4=; b=TMywT7btQU4u3spSyoWLDJ/468nq6xhx+jMZZrf3gwhl9IYxrBjulIblMspX9W3cL+ c4qpDCm76ZfB8AfEzft9P0ancqXdpJ8tdP4WnwlgoHW4R+hRErCYJD4q7tAuc27A87bO SBvwhFhntnGnFMFENvJ+vCXn86i2W7mQoDeUJ1esqLo974aYVdfIEyQaqx8dU89GQCt3 zH5lVyZ2yoOHv7X3HC8Oe9unYFTTurcnR4FirWfH/zRj+Ou0qpYioCXEaulEBDDUlM9X X+dBEEUhzK3xWTr056T3l561Ni0CAyWqQ0SZgmNwnDnZC5wZg/isXfeHhlLM0GfB6v/f ux9A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784458418; x=1785063218; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=FhGxTEOOON4RkKT9KdpeyXrdqn2SrMJKFWwPDsn11x4=; b=N7hvWj40jWxEDsTK2MpVLAtoMHFYk/Db49reCzkKsBDRXr9aApb4p7F4eiPqfvM26s Q8FrWnvc+63VHA3KzpAWFPJ0udlEXDG04l5geSPFm/383XUwcSMRmOoMzGKSf/k9XuHG DgjesE/czKgHByNo8Hkum7CsCe10pV31LG1LVMGw+k0hf3Jy3gVGxkxg2b4LS3bUiTQT 99MYISTi7/D60dUa5wuzhsN7wM8wwivOWsCFhtHfue6hcTCZR9CIc7hSj8Qzv+G4NEnU O+gEdONoJRZq/+sdH4O+VTyzfH9GvPhZuAfVw482sSgIfrJTChANXLxN6hUZFEXXMlj1 t1oA== X-Forwarded-Encrypted: i=1; AHgh+RrwGc28HV9Pha6VqXh3I2jvx3dNxFzgXxbkMMk9/T4t4HUyATPao3PL+WsTytY7Jqoyw2xH2h4C+COHu0g=@vger.kernel.org X-Gm-Message-State: AOJu0Yz8CqgjapMKDr9c3MKsg2z9ngCrZuMTBex3bHoofEHUBK6ZkQ8F 9x/45U7hplQdkWx7igVa48YrqfSnoCMmraPkOQJHY4lT6WD71KyNYuDdpNR+tygXdg== X-Gm-Gg: AfdE7cl06dnljeOikvcTzSBuehEivrPjT0hk16Ytoci+poK0jFsHY/5kWZt0gE4X/Hs ITHjAZ07fgV0Bg7TLfdUKDJxuR/+FAANB+6VF/mwwlu/PoP5/uUKkbTsyjKcSQHpryiU8xx1d4+ fc5h0D/wIc+BSPVin99fFphMLRHaB/hr9dPSgBu4K42yOR8+E4z6mOyacYDtIBaJrUHTyCcfyut wxuB+xudP2O81loJia9q/xrjycHI6fiz4bWJaVa3DiFPViV7UtXRgLyvZ7Wmucl4QEX3teejsbA o2f03wdjXl/LYY6jqA58JcTJN+9/54+pKUg4rmCt0sTsN+Qj0VXQ6ZPkDgo7ce6RQF88HxhUZFg NRf/+ws65DBIY1o2LMGuP9R4+JVK9L8IrdItSgOlq+AiMup/SD7vXYS02RZ93aQ== X-Received: by 2002:a05:600c:35d1:b0:495:3a52:71b1 with SMTP id 5b1f17b1804b1-4954aa1a34fmr92326745e9.5.1784458418163; Sun, 19 Jul 2026 03:53:38 -0700 (PDT) Received: from mf-00-01.. ([194.220.239.180]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-4954a2a24b6sm197679575e9.3.2026.07.19.03.53.36 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sun, 19 Jul 2026 03:53:37 -0700 (PDT) From: Mykola Marzhan To: Jens Axboe , Song Liu , Yu Kuai , Keith Busch , Christoph Hellwig , Sagi Grimberg , linux-block@vger.kernel.org, linux-raid@vger.kernel.org, linux-nvme@lists.infradead.org Cc: Li Nan , Xiao Ni , Guoqing Jiang , Leon Romanovsky , Jason Gunthorpe , Kiran Kumar Modukuri , Chaitanya Kulkarni , Logan Gunthorpe , Bjorn Helgaas , Shivaji Kant , Pranjal Shrivastava , Henrique Carvalho , linux-kernel@vger.kernel.org, linux-rdma@vger.kernel.org, linux-pci@vger.kernel.org Subject: [PATCH v2 2/7] md: ensure REQ_NOMERGE is set on P2PDMA bios Date: Sun, 19 Jul 2026 10:53:22 +0000 Message-ID: <20260719105327.864949-3-mykola@meshstor.io> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260719105327.864949-1-mykola@meshstor.io> References: <20260719105327.864949-1-mykola@meshstor.io> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" md_submit_bio() unconditionally strips REQ_NOMERGE before passing the bio to the personality, an optimization from commit 9c573de3283a ("MD: make bio mergeable"): a bio that md has split may become mergeable again below md. For PCI P2PDMA bios the flag is load-bearing, not a hint. The block layer sets REQ_NOMERGE on P2PDMA bios (__bio_add_page(), and the extraction path of bio_iov_iter_get_pages()) because the DMA mapping type of a request is resolved once, from its first segment (blk_dma_map_iter_start()), and request-level merging is prevented only by REQ_NOMERGE. Stripping it allows the member queue to merge a P2PDMA bio with a bio carrying pages of a different pgmap, or host memory, mapping the merged segments with the wrong bus address: silent data corruption on the member. Set the flag for P2PDMA bios instead of merely preserving it. No in-tree path currently submits P2PDMA pages through the bvec-iter path (bio_iov_bvec_set()), which skips the flagging -- but nothing structural prevents one, so setting rather than preserving hardens md against that gap at the cost of one branch. Everything else keeps the original optimization of clearing the flag. This covers every personality that advertises BLK_FEAT_PCI_P2PDMA (raid0, raid1, raid10), which is why the fix lives in the shared md_submit_bio() path. Fixes: 02666132403a ("md: propagate BLK_FEAT_PCI_P2PDMA from member devices= to RAID device") Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Mykola Marzhan --- drivers/md/md.c | 14 ++++++++++++-- drivers/md/md.h | 18 ++++++++++++++++++ 2 files changed, 30 insertions(+), 2 deletions(-) diff --git a/drivers/md/md.c b/drivers/md/md.c index d1465bcd86c8..3ae4fd4ef381 100644 --- a/drivers/md/md.c +++ b/drivers/md/md.c @@ -451,8 +451,18 @@ static void md_submit_bio(struct bio *bio) return; } =20 - /* bio could be mergeable after passing to underlayer */ - bio->bi_opf &=3D ~REQ_NOMERGE; + /* + * A bio md split could be mergeable again below md, but for P2PDMA + * bios REQ_NOMERGE is load-bearing: the DMA mapping type of a + * request is resolved once, from its first segment, so requests + * must stay single-provider (see __bio_add_page()). Set the flag + * rather than merely preserve it -- bios built through the + * bvec-iter path arrive without it. + */ + if (md_bio_is_p2pdma(bio)) + bio->bi_opf |=3D REQ_NOMERGE; + else + bio->bi_opf &=3D ~REQ_NOMERGE; =20 md_handle_request(mddev, bio); } diff --git a/drivers/md/md.h b/drivers/md/md.h index d8daf0f75cbb..140e2b3670d8 100644 --- a/drivers/md/md.h +++ b/drivers/md/md.h @@ -11,8 +11,10 @@ #include #include #include +#include #include #include +#include #include #include #include @@ -22,6 +24,22 @@ #include =20 #define MaxSector (~(sector_t)0) + +/* + * Check if the bio carries PCI P2PDMA (peer device memory) pages. Read + * bi_io_vec directly rather than using bio_first_bvec_all(), which WARNs + * on cloned bios: md routinely handles split clones, which have + * bi_vcnt =3D=3D 0 but a valid bi_io_vec shared with the parent. P2PDMA a= nd + * host pages must not be mixed within one bio, so the first bvec is + * representative. Only valid before the bio's iterator is consumed: + * bio_has_data() is false at completion time. + */ +static inline bool md_bio_is_p2pdma(struct bio *bio) +{ + return bio_has_data(bio) && bio->bi_io_vec && + is_pci_p2pdma_page(bio->bi_io_vec->bv_page); +} + /* * Number of guaranteed raid bios in case of extreme VM load: */ --=20 2.43.0 From nobody Sat Jul 25 03:47:43 2026 Received: from mail-wr1-f47.google.com (mail-wr1-f47.google.com [209.85.221.47]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A9ED3393DF5 for ; Sun, 19 Jul 2026 10:53:41 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.221.47 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784458423; cv=none; b=csePE6L7Mr6WzpKyQ8UUuVAfGqt8Q5e+jx3/gAgLTNURTaJgn+Lj7aPI1vLSbiHtmdS5IhbwxwaxtlodvRlHNi3v3zEJFTqaRtB2vJz5V6p13Wk5PPh40G8qRvWK7DXLmtkJkVJyYPsdEK+AIgD/dv6Ir06rqJ+AsHIh5PV0j70= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784458423; c=relaxed/simple; bh=RcnkLcv6uBny2kPlRJHtExCJJhAqkz60vxDTP6tEfQw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=afy6f/fv6ksTFuH6ou1JAftQvsEfSInQp96OmRWK1DtlNhbyz9M5K7j82PznNDDAwlgrfyZHaOE6DRvui4SXNQKnNLhsmda3SWwt6pk1Vm+EWxY8Nsuklb+qCGBVxbWgQt1LRc510G82VSTiRHx3yEvwnSiwdvzZlG6vagYMXQY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=meshstor.io; spf=pass smtp.mailfrom=meshstor.io; dkim=pass (2048-bit key) header.d=meshstor.io header.i=@meshstor.io header.b=iNZFLepK; arc=none smtp.client-ip=209.85.221.47 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=meshstor.io Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=meshstor.io Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=meshstor.io header.i=@meshstor.io header.b="iNZFLepK" Received: by mail-wr1-f47.google.com with SMTP id ffacd0b85a97d-47c6e9a694bso5013088f8f.1 for ; Sun, 19 Jul 2026 03:53:41 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=meshstor.io; s=google; t=1784458420; x=1785063220; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=wPPSB8e7R/CQvh84euJ6RH38MKEu1i6y4HBVSmJRVhc=; b=iNZFLepKnvAsq0juCcoWumRJg8mGfwZKiSYkEY7iYKb6HCFI+nwlM5dMrkSlpulB/W lG4DA583q9wxQGR780qk7oGdD+KvYUdRboKpW6Bh+x9dyKTVNR9lov1mY7apNAo2opTa Bz66dKi82M69rfpSZiml14/h0kMsQ2Y+983O3HmgEdDpy7admSBKp8aIM/uwNvsKGD5B 8FyRXqFVrl504DikToQC6qNLvkCQfnG5Ll3o5v+aUmTo4GCx3Z/1VMRFAEP3YijzyKOR aQEXsGmZ8RK2jmubHlPxzUUlCz5jNqih6jLB81iA+gmgGN00+dQrhfXP9y8gKxG90jQz b8NA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784458420; x=1785063220; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=wPPSB8e7R/CQvh84euJ6RH38MKEu1i6y4HBVSmJRVhc=; b=SEp9v2Y0/DujhDJI8OAJ6DCnnozr++5g88902sH7Vq4qLO4VyrgFt7GB9yA03zaWbQ CtttpLk2pIb9CBCECaDfuDhlW3zcvijfC+Hx8zuYdxBJ+Rz1vu7SdVMjy4X82WPgBdbo zOStB7xIPCDh/w9+rA2Ch2O4ExcIk7VkD+/LdbQe9UqpKt7QT8tshHNGHb1Al7dXfX9H yz17M+PQIN6UY/XD6MUglS2mdmV9sXky6Vrf9jsWx079ttTC5vo+qlPQDTGiFfdigbks dOo2IZxj5USazrX0unzq1U/MuiZ9JRLx0GhXghCN8zlbIgWA4kQaCZYYEpr7q9ALtSL7 gWxA== X-Forwarded-Encrypted: i=1; AHgh+Rr+QY0UUspsMk2a+7xTraPE8Tjq2+TvR5VSlye3wvq+ctDvz89vpIYt/UYp/Gnl8YdQFK9hByawXMx3oxg=@vger.kernel.org X-Gm-Message-State: AOJu0YwLXA66GqohSa/g9Ai2QECzm8zPMq1OTAaZf2MDtXWArrODevQv 7imimNHaHLzowhY6KJ2zoJuN3620bv7Pndbg/Sz7+zP0yInXohPxoKFhRB/Hc6r6alpdTGQJqd/ ipLKd8Xr9 X-Gm-Gg: AfdE7clCi7/9aB/4KrcF3wOpA3bNdtI4G1vfPA3hAihy9JGTqmaJ3S1/QIfEXRh86lY gyZyw8HzNI93z4NgAiAUVOUamTFZQ4JRtcH3RxC9hnK4M8flffz1KNKvfFOd2/8sREsXUUgczqY mMdfpkzf88vBOcz1wwfEecJ5mS/ylLDdW+UMZ4AMvsmNsF+YREaukfqneTIsNSx5GklW5nNh9dQ ToZ2AkPcsX8u4HdZzBlX0MTD+d9B4UdaX8iBcgx8qy1+sJMswQKhpmpngjdog/JWPZ4KwRI948l kmLjBOETPzslypZ2AfSzGuVIjquD3U5FmaLi16c523G1VLq946V/crilzSsmE0EI329N79IAU9t hN7YHqhYlJT/drv7NtKfPxRIJyO3BfQz6tpLXg54TYyy5KkeOY3B7SZn6BPdovg== X-Received: by 2002:a05:600c:b90:b0:493:c42c:7e87 with SMTP id 5b1f17b1804b1-4954a50885emr97913395e9.33.1784458419927; Sun, 19 Jul 2026 03:53:39 -0700 (PDT) Received: from mf-00-01.. ([194.220.239.180]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-4954a2a24b6sm197679575e9.3.2026.07.19.03.53.38 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sun, 19 Jul 2026 03:53:39 -0700 (PDT) From: Mykola Marzhan To: Jens Axboe , Song Liu , Yu Kuai , Keith Busch , Christoph Hellwig , Sagi Grimberg , linux-block@vger.kernel.org, linux-raid@vger.kernel.org, linux-nvme@lists.infradead.org Cc: Li Nan , Xiao Ni , Guoqing Jiang , Leon Romanovsky , Jason Gunthorpe , Kiran Kumar Modukuri , Chaitanya Kulkarni , Logan Gunthorpe , Bjorn Helgaas , Shivaji Kant , Pranjal Shrivastava , Henrique Carvalho , linux-kernel@vger.kernel.org, linux-rdma@vger.kernel.org, linux-pci@vger.kernel.org Subject: [PATCH v2 3/7] md/raid1: serialize non-write-behind writes on CollisionCheck rdevs Date: Sun, 19 Jul 2026 10:53:23 +0000 Message-ID: <20260719105327.864949-4-mykola@meshstor.io> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260719105327.864949-1-mykola@meshstor.io> References: <20260719105327.864949-1-mykola@meshstor.io> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" A write-behind write is acknowledged to the caller while its copy to the write-mostly member is still in flight. wait_for_serialization() exists to order later overlapping writes behind it, but the non-behind clone path only takes it under serialize_policy. A write to a write-behind array skips the behind path whenever the behind queue is full, a reader is waiting on behind completion, or the behind bio allocation fails (raid1_start_write_behind()) -- and then races the in-flight behind write on the write-mostly member without any ordering: if the older behind data lands last, the member keeps stale data for sectors whose newer write the caller has already seen acknowledged. Gate the non-behind serialization on CollisionCheck alone. The flag marks exactly the rdevs that own a serial tree: every rdev present when serialize_policy is enabled, and write-mostly members when write-behind arms serialization, which is the case the current test misses. Take the matching remove_serial() at completion under the same condition. Discards take the same non-behind path and are now ordered as well. For an rdev hot-added while serialize_policy is already enabled, mddev_create_serial_pool() never builds a serial tree, so the old MD_SERIALIZE_POLICY gate sent it into wait_for_serialization() with rdev->serial =3D=3D NULL -- a latent oops that the per-rdev gate avoids by construction. REQ_NOWAIT writes can wait here when they overlap an in-flight behind write, as they already do on the behind path and under serialize_policy. Fixes: d0d2d8ba0494 ("md/raid1: introduce wait_for_serialization") Cc: stable@vger.kernel.org Assisted-by: Claude:claude-fable-5 Signed-off-by: Mykola Marzhan --- drivers/md/raid1.c | 11 +++++++++-- 1 file changed, 9 insertions(+), 2 deletions(-) diff --git a/drivers/md/raid1.c b/drivers/md/raid1.c index afe2ca96ad8c..8172df882f26 100644 --- a/drivers/md/raid1.c +++ b/drivers/md/raid1.c @@ -564,7 +564,7 @@ static void raid1_end_write_request(struct bio *bio) call_bio_endio(r1_bio); } } - } else if (test_bit(MD_SERIALIZE_POLICY, &rdev->mddev->flags)) + } else if (test_bit(CollisionCheck, &rdev->flags)) remove_serial(rdev, lo, hi); if (r1_bio->bios[mirror] =3D=3D NULL) rdev_dec_pending(rdev, conf->mddev); @@ -1677,7 +1677,14 @@ static bool raid1_write_request(struct mddev *mddev,= struct bio *bio, mbio =3D bio_alloc_clone(rdev->bdev, bio, GFP_NOIO, &mddev->bio_set); =20 - if (test_bit(MD_SERIALIZE_POLICY, &mddev->flags)) + /* + * Order against in-flight write-behind I/O: a + * behind write is acked early, and an unordered + * overwrite could land first, leaving its stale + * data on the member last. CollisionCheck marks + * every rdev that owns a serial tree. + */ + if (test_bit(CollisionCheck, &rdev->flags)) wait_for_serialization(rdev, r1_bio); } =20 --=20 2.43.0 From nobody Sat Jul 25 03:47:43 2026 Received: from mail-wr2-f0.google.com (mail-wr2-f0.google.com [74.125.225.64]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6C6C93911B5 for ; Sun, 19 Jul 2026 10:53:43 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.225.64 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784458425; cv=none; b=PxcnbhTQZJdHcc59zHCeaMyQcndcrSAbqWOdQPwbLOlgIMWzq0RZr1G9pZwZxc2N3XVRpjQ/QeNLP8y92uhs2DKi6IBF16usmo12lgE0+dfYvIB0gvcRdmFuYwSIkvw15ROcpi5y+tclltsJ30aqR2SQNKX5a1FrVPP8VdyPSnc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784458425; c=relaxed/simple; bh=x41+O7RfOQE1nIXUFKbgkPW7MiFw2TutvahKq8Yq6SE=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=tnjtAmwOUcrO67N1zrQ+tnCBhOFxF/ZErsC1ooW95fs09KCG4ftEuzcZmdUpbifKp9kXWQUnkch/NavuhqbVSfHqvk635RXwZSwWlbZPuPr52Nad2Ga4XqKE01NdB6Y/eN3DvwzJKb2DWG+nEYQfyG9vsTw9Rywk1flI2PG7NII= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=meshstor.io; spf=pass smtp.mailfrom=meshstor.io; dkim=pass (2048-bit key) header.d=meshstor.io header.i=@meshstor.io header.b=bUSiwg81; arc=none smtp.client-ip=74.125.225.64 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=meshstor.io Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=meshstor.io Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=meshstor.io header.i=@meshstor.io header.b="bUSiwg81" Received: by mail-wr2-f0.google.com with SMTP id ffacd0b85a97d-47f6c54e841so379532f8f.0 for ; Sun, 19 Jul 2026 03:53:43 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=meshstor.io; s=google; t=1784458422; x=1785063222; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=khPWV7pl2ttOdMmwNwArwe3lfR4e9+Vyg5lRTCmC9jU=; b=bUSiwg81KAOfM22LDx9gw+t50bGUmYodReq+fQquhtG/JsihiqGkf/3EopbJmmw7cJ q/lqm3v6njEwm3qFRJ8NjUlQ9XMuzATBbXRkO6YvPaHFlE0bOzoHa+/GR2h3q2xfddUr P/2HsgsnCiGmy0cgLTi1clx2QHR+DB4pxXz7RQ9rVmiDcV4N4dxGYXTF2CHTJF5D0CO0 tYZl4gZPuyXa1wtkj+pryj/f+9fev7qAubOycWs8WaIngEbzUL+Fxb1nexVnaBrRTYyZ mAGJj1JwFLzIFeALKINsW08EogotFPypaBgplE8TBfb5WdRxpoke/wHSsrKKG0Jkbbog E/6Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784458422; x=1785063222; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=khPWV7pl2ttOdMmwNwArwe3lfR4e9+Vyg5lRTCmC9jU=; b=Kq+zduXM6pZtAs8gcGXlptG9ffzLOn8IgBNCPH3X/pVR05NPxsu3nGBDmRjo3b0vA6 5iWQ+sbA4V2vk5zP5RYoUK14UN8osI97YA/L2bHG+oIwEp2rwHZ46Wkl8thAA8S81qHV 6yqyFHokqunvjQnposgKeL/chYjbQMZQ9FfFlBYE2KPDfVIZYuyAz0bzh0+WUZvAuaG9 G7t2g+0wgoIKL8UGy9B7YrcEklUzXEIQgvnYF9GY8BlW2EmGkUdtfVJjj7Pc75moLySf fAk7EVvXnHMqNl6t7dh/QR5MgFl/Q0XIsMsIW2BjIUgWIBGKYR6OuyxKdvuAFE9aPckt ebVQ== X-Forwarded-Encrypted: i=1; AHgh+RqiuIGJGO9eeLIFGcjl/wZWcaXFRl1We/z6E+IRwj1eSvNqcevzyV3+/PASvui7L6sfC9ODXb+Ue5+dxSI=@vger.kernel.org X-Gm-Message-State: AOJu0YwIMoerod9+6eQy1QK97xdr1HpsmETjhB/YbnmD8y3j0QJ6h4ZP GjBBMLEuruNfVxRgQM69JFB9W4RDPqGDJOtM7O17A3xEJEtKjzoVj97AhxNrQnnxsg== X-Gm-Gg: AfdE7ckvi3IG7KWdhSYjLAXD45911FFIAPilTgaqLNvOHHw1oSomzThqAsw1cxSWJXK Snlg1+MtiKwOQxX9Sr+LQEWC1esrJKdJbviUMgxisy13JT0UlYxzQSb/JOwd26CAkEptV30PEMV wAxSzyQK4aLq5qgyzAUUxNd2g/YLcxS5eHsI5PUmNn6r52by0EhKSUKVizpdv/C71QDjU4nT+MM zHgm+pIwxJ13P/9wCYj3BzQtRmkVU2kJARVz1v2yYscAGhOrpHIbsxzUKQKFliVPUNYGonFQBy2 dnMSfU4oYM04VqoZotLiJftfMMil/Vmt72EQrcxYee2CA/vC89l9wmErXCbs2CJC9O4ixZ95xdF dJg3TkiyfJg1v/so79gEyRGSFvBtBhvkuXahT0WGEFxhmDKqgf33RGGrnm0Y6Qw== X-Received: by 2002:a05:600c:6dc9:b0:490:688b:f9f8 with SMTP id 5b1f17b1804b1-4954a5151e5mr80660155e9.27.1784458421726; Sun, 19 Jul 2026 03:53:41 -0700 (PDT) Received: from mf-00-01.. ([194.220.239.180]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-4954a2a24b6sm197679575e9.3.2026.07.19.03.53.40 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sun, 19 Jul 2026 03:53:41 -0700 (PDT) From: Mykola Marzhan To: Jens Axboe , Song Liu , Yu Kuai , Keith Busch , Christoph Hellwig , Sagi Grimberg , linux-block@vger.kernel.org, linux-raid@vger.kernel.org, linux-nvme@lists.infradead.org Cc: Li Nan , Xiao Ni , Guoqing Jiang , Leon Romanovsky , Jason Gunthorpe , Kiran Kumar Modukuri , Chaitanya Kulkarni , Logan Gunthorpe , Bjorn Helgaas , Shivaji Kant , Pranjal Shrivastava , Henrique Carvalho , linux-kernel@vger.kernel.org, linux-rdma@vger.kernel.org, linux-pci@vger.kernel.org Subject: [PATCH v2 4/7] md/raid1: don't use write-behind for P2PDMA bios Date: Sun, 19 Jul 2026 10:53:24 +0000 Message-ID: <20260719105327.864949-5-mykola@meshstor.io> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260719105327.864949-1-mykola@meshstor.io> References: <20260719105327.864949-1-mykola@meshstor.io> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" alloc_behind_master_bio() copies the bio's data with bio_copy_data(), a CPU copy. P2PDMA pages are peer device (BAR) memory; generic code must not assume CPU load/store access to them is safe or fast on every architecture, and bouncing peer memory through the CPU defeats the point of a peer-to-peer transfer. Skip write-behind for P2PDMA bios: they are written directly to all members, including write-mostly ones. Ordering against write-behind I/O in flight to overlapping sectors is preserved: the non-behind clone path serializes on CollisionCheck rdevs (see the preceding fix), which covers these bios like any other write that bypasses write-behind. Fixes: 02666132403a ("md: propagate BLK_FEAT_PCI_P2PDMA from member devices= to RAID device") Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Mykola Marzhan Reviewed-by: Logan Gunthorpe --- drivers/md/raid1.c | 8 ++++++-- 1 file changed, 6 insertions(+), 2 deletions(-) diff --git a/drivers/md/raid1.c b/drivers/md/raid1.c index 8172df882f26..914fb86452c0 100644 --- a/drivers/md/raid1.c +++ b/drivers/md/raid1.c @@ -1523,6 +1523,7 @@ static bool raid1_write_request(struct mddev *mddev, = struct bio *bio, bool write_behind =3D false; bool nowait =3D bio->bi_opf & REQ_NOWAIT; bool is_discard =3D op_is_discard(bio->bi_opf); + bool is_p2pdma =3D md_bio_is_p2pdma(bio); sector_t sector =3D bio->bi_iter.bi_sector; =20 if (mddev_is_clustered(mddev) && @@ -1575,9 +1576,12 @@ static bool raid1_write_request(struct mddev *mddev,= struct bio *bio, /* * The write-behind io is only attempted on drives marked as * write-mostly, which means we could allocate write behind - * bio later. + * bio later. P2PDMA bios are excluded: write-behind copies + * the data with bio_copy_data(), a CPU copy that cannot be + * assumed safe or fast on P2PDMA (device BAR) pages. */ - if (!is_discard && rdev && test_bit(WriteMostly, &rdev->flags)) + if (!is_discard && !is_p2pdma && rdev && + test_bit(WriteMostly, &rdev->flags)) write_behind =3D true; =20 r1_bio->bios[i] =3D NULL; --=20 2.43.0 From nobody Sat Jul 25 03:47:43 2026 Received: from mail-wm2-f7.google.com (mail-wm2-f7.google.com [74.125.225.135]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C58C8396579 for ; Sun, 19 Jul 2026 10:53:46 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.225.135 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784458428; cv=none; b=ZH2CbyBhfQPT21/MmCv99Rw2W3ntadbWJ2xiTpzq4uiF+vit2x/E8E+P4xv7lm9AknxMSEmkHw/bmzd8Hz2sdf7kmH0xR9cQvLsyblEN2z6N4JAZrGoNUSBuJsnzf+3guyW85iA3ossgSJ5COWHpbG9JyfY/ZYs3diVC9mJeVus= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784458428; c=relaxed/simple; bh=ckOwh+/KVDM4Fp986EbTgYg+bPUMzv5VlFWjz/PbPJg=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=XZLV57viusEUMkWHP9bOFTUyNykGfWrwVxHxSwDzh6M9+GuHlZjbUk9b8btB9IC7dHqn8JoX/Lq/trHjhTKDOlvrnl+NOZbKHGahUMIcKiNK9fdon//+3P/wvcZrWHuqegzwcTGJDcv1JqgG6HOwHa4BQoHC5AhbDMc7+aCjY5A= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=meshstor.io; spf=pass smtp.mailfrom=meshstor.io; dkim=pass (2048-bit key) header.d=meshstor.io header.i=@meshstor.io header.b=JAFucgbw; arc=none smtp.client-ip=74.125.225.135 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=meshstor.io Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=meshstor.io Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=meshstor.io header.i=@meshstor.io header.b="JAFucgbw" Received: by mail-wm2-f7.google.com with SMTP id 5b1f17b1804b1-49557073245so2298355e9.1 for ; Sun, 19 Jul 2026 03:53:46 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=meshstor.io; s=google; t=1784458425; x=1785063225; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=2eFgBNGHZVs+4VEn+rFjG8dtsShNuRCRL19zmjqYsAk=; b=JAFucgbwqnuXms9cXGHT8sLvtkrldP4kQoLjBJrfyz5rDi2+4d5LG3rxQcsOb56Nlb 44iYz5BRwf+XIY1G3w6CKTzBUP8cY1ivHuE8tcO2q/7xURPinrCs3IqfItPOmCv+jYYH CpattxbHEKF4zH1H5AYZAy2ZImrZbNxRyr1chBAVeu907UB150gQktRQsNlBo5SSfTzV /T2jMiUfEKKpppOUMn42Ox4DrZtiuAlSG3NlshP8/okOzvoJ/KQJBcCtdORJHBJjSMmT QedaBihhuQXXHk9pc3tU+g6LiDE37OrCjUyYoBU+lFcx7fkjakAetCNySuduRRuq4vCr Obuw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784458425; x=1785063225; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=2eFgBNGHZVs+4VEn+rFjG8dtsShNuRCRL19zmjqYsAk=; b=cE0ey4kxdA7mrL2hQ0TUw1xoNNwXwGfv6k4AAfJwvDj4HiLLBOPAyAMxsq/UeTT1QB gjm5KucvviMVhkoU+6At4+9Mqsf+SlpsuMIHovNlK7JnAhgeE0vW0E2AQFic/9CqLohg QqzTLA4z2bx1DoQRxyjhsYj9inXea1ZZgW8F2qwyKkG4yhjQb8ZPTg8z9+kgLD9uijWM 6pZc3tNkSh5IKqktN7wzY9WhQb+P0UWnHNjGmYQO29cykaH2D4dxKSgDVZYa4EzhA1rh PfWXR/LBSGVPqonNrxmcBAXr+aXpXmDV/EqCOuZn5ALN3TTisVoW2a4zci0Z6bNVI5Qi SK7Q== X-Forwarded-Encrypted: i=1; AHgh+RqAkP4MpczzMEc1PoLWtCaPZPEtFVlh3sY4NiKRpwvr5+4jUVwb3GpepOudcaGIfI6iuY/LYIbdgpjNOkI=@vger.kernel.org X-Gm-Message-State: AOJu0Yw7OyENJJRRdFV8789A/abZmfaMzuylDTyaZFkDt0mGEgqDDRAf v19Y7QSckUgWO/tqWdsB93E65lVDE9Gn95NRXdjcveWm8cE3PNeMm9UU/XV5wzHDRg== X-Gm-Gg: AfdE7ckBTQva8JoNep82rUkU+gU/KNsXwv95WV/b6okoWEi9ClQQJqgCaIKuzBtTBIy xAFRMfiGndngefhI4WTji+FBG1AYrfkddKQxOsEUMf6fVAuQLaSD7U3PaDbLwbIm+P42GHZ2b4Y lNjQsTjzHvgjJYbb4tCCKRKEb/T/D2JpmY3tAL23iNIFfKC6Q2t8d9QnFQ40n6VwO3r1rPmNoFH Sof9gFL8fsFgrFzpwTN5yIfpysF0cvUR23e5J9EwjGLzTsd0ZUuVA0kS+KY7AzSQuOoidGP+/th 5YSJb7fJF33RPf7SJnlGnUR7O5Ot71pQdP39oCX14f4Bss1z5FbvY7/PXtfiWLFB6ndX6bSBrfy 3SkfxkzVO5N1mI6mK8V7Pux/faLASeo7xUr437AFE95ZgCNkB5bz310MU5QNN0UOOGoSCrwpR X-Received: by 2002:a05:600c:4585:b0:492:437a:a653 with SMTP id 5b1f17b1804b1-4954a50b811mr100979195e9.26.1784458425036; Sun, 19 Jul 2026 03:53:45 -0700 (PDT) Received: from mf-00-01.. ([194.220.239.180]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-4954a2a24b6sm197679575e9.3.2026.07.19.03.53.41 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sun, 19 Jul 2026 03:53:44 -0700 (PDT) From: Mykola Marzhan To: Jens Axboe , Song Liu , Yu Kuai , Keith Busch , Christoph Hellwig , Sagi Grimberg , linux-block@vger.kernel.org, linux-raid@vger.kernel.org, linux-nvme@lists.infradead.org Cc: Li Nan , Xiao Ni , Guoqing Jiang , Leon Romanovsky , Jason Gunthorpe , Kiran Kumar Modukuri , Chaitanya Kulkarni , Logan Gunthorpe , Bjorn Helgaas , Shivaji Kant , Pranjal Shrivastava , Henrique Carvalho , linux-kernel@vger.kernel.org, linux-rdma@vger.kernel.org, linux-pci@vger.kernel.org Subject: [PATCH v2 5/7] md/raid1,raid10: keep REQ_NOMERGE on narrow_write_error() retry clones Date: Sun, 19 Jul 2026 10:53:25 +0000 Message-ID: <20260719105327.864949-6-mykola@meshstor.io> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260719105327.864949-1-mykola@meshstor.io> References: <20260719105327.864949-1-mykola@meshstor.io> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" narrow_write_error() re-issues a failed write in badblock-granularity chunks, cloning from the master bio and resetting bi_opf to a bare REQ_OP_WRITE. For a P2PDMA bio that reset drops REQ_NOMERGE, which is the only request-level protection against the member queue merging P2PDMA segments across pgmaps or with host memory (see the preceding md_submit_bio() fix): the retry path would quietly reopen the hole the submission path closes. Restore the flag on P2PDMA retry clones. Fixes: 02666132403a ("md: propagate BLK_FEAT_PCI_P2PDMA from member devices= to RAID device") Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Mykola Marzhan Reviewed-by: Logan Gunthorpe --- drivers/md/raid1.c | 3 +++ drivers/md/raid10.c | 3 +++ 2 files changed, 6 insertions(+) diff --git a/drivers/md/raid1.c b/drivers/md/raid1.c index 914fb86452c0..f562b6bd438b 100644 --- a/drivers/md/raid1.c +++ b/drivers/md/raid1.c @@ -2573,6 +2573,9 @@ static void narrow_write_error(struct r1bio *r1_bio, = int i) } =20 wbio->bi_opf =3D REQ_OP_WRITE; + /* Keep P2PDMA retry bios unmergeable, like the original */ + if (md_bio_is_p2pdma(wbio)) + wbio->bi_opf |=3D REQ_NOMERGE; wbio->bi_iter.bi_sector =3D r1_bio->sector; wbio->bi_iter.bi_size =3D r1_bio->sectors << 9; =20 diff --git a/drivers/md/raid10.c b/drivers/md/raid10.c index 0a3cfdd3f5df..f7ef903a3d4e 100644 --- a/drivers/md/raid10.c +++ b/drivers/md/raid10.c @@ -2831,6 +2831,9 @@ static void narrow_write_error(struct r10bio *r10_bio= , int i) wbio->bi_iter.bi_sector =3D wsector + choose_data_offset(r10_bio, rdev); wbio->bi_opf =3D REQ_OP_WRITE; + /* Keep P2PDMA retry bios unmergeable, like the original */ + if (md_bio_is_p2pdma(wbio)) + wbio->bi_opf |=3D REQ_NOMERGE; =20 if (submit_bio_wait(wbio) && !rdev_set_badblocks(rdev, wsector, sectors, 0)) { --=20 2.43.0 From nobody Sat Jul 25 03:47:43 2026 Received: from mail-wm1-f46.google.com (mail-wm1-f46.google.com [209.85.128.46]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 371AC390231 for ; Sun, 19 Jul 2026 10:53:49 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.46 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784458431; cv=none; b=He0lMTcwBWco6yx8BQhlfsJahd52xvGs/W8Jk90Gus7zbG8QJvfShw8l1CkBMpPvLnYec/6XneUhEHR+XqRbxBWYlg5bXZkmhBBE5Ud5VOtoGasvLXVsDU087YUcCyDV54efdDpNW/R+pPWgHxRJr10fC5+07+YIh4jCR8FpIzI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784458431; c=relaxed/simple; bh=o6YIqEshj6myZ6eisq9eZWd8t50ZTZG9XsYkTI9G+RQ=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=D4qkyl8nSDiJV56oLbHT4skqYOhAR/KXZn3PkQ1dZUbI3QbnUU+c9RRq7zxlBmCijfNCGFvMeqp9uf8UrKdWDt5xbuj6ueqKK8xVSzJjVn1iyeX+/83r8SKbCtZ6AHQDmJ5aLG3mXrLYK3AZHHORPJ+vXx24IL595cyJ0POzyVM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=meshstor.io; spf=pass smtp.mailfrom=meshstor.io; dkim=pass (2048-bit key) header.d=meshstor.io header.i=@meshstor.io header.b=IfcC6xLt; arc=none smtp.client-ip=209.85.128.46 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=meshstor.io Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=meshstor.io Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=meshstor.io header.i=@meshstor.io header.b="IfcC6xLt" Received: by mail-wm1-f46.google.com with SMTP id 5b1f17b1804b1-49556f97a9dso3162345e9.1 for ; Sun, 19 Jul 2026 03:53:48 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=meshstor.io; s=google; t=1784458427; x=1785063227; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=DOCIvHQcUURC426TdmWIVqFzZor35QO+zXIWNPhmEXE=; b=IfcC6xLtAwcE3RRJGTLRlOefYRRatTBnmYKWEyCj1Kw47wrLY7br5+s8+r6kIUsTBI nj+C0GiE928YbJHsi3aPjhUi4Y/2/+wXre1GpLQpVJbdaBFWx2/pzZX5GNr2vkMnLq0O EYBY6eyYgglN0InQSZA+TLgpLxbsnTc8zDA+lGvybcXsWidDr4qsH3qRR1pTG8a1pi1Z 7eABjBU1Vm0F7UHf3dhjeh4gh5f0UPV4aeGh60oZnGAFuDn/DXOqNn0Y71iyxo0GKiwD QwVoqkuNXpb6ujBk25XTqe46w/Oxm9wJkz69O9DHdb99ZSwLyjlC5ISoLUaur9oDRYm1 /tuw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784458427; x=1785063227; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=DOCIvHQcUURC426TdmWIVqFzZor35QO+zXIWNPhmEXE=; b=tUo7RLXEQL2nlS3a7LEZRLvTyFM+4HUYHJoPEzh5a0vfnaNuI/iOghowjscqfuoV8R ihoFcz8qAH5q82LXM9D0UM1iuxRJbAX6lP+5bqIxPTpkfKszp2m8/fyCP5p4UlRHfBEM DhLKm3f5GEWkqdxhrfvfoXP31E42mn0SubR2lXL5dvmX7d4YZjjxugJB7KPBNiwN9h/Z /U1sy+t+v7LCz/2YOc/OmwsEdV5LYycknpQqXP+R9tBY1z5X5QE4fmMS6wDncEw+OufC EE/gNy5g/czGzrRGVApbtwcLw2ewvFmvVY+/ALOUjbnjltQiyLfX2JMPvKyvNRYcTgWD loMQ== X-Forwarded-Encrypted: i=1; AHgh+Rqi5t/fz1YmicDA4M4iH6KW95d3v3F26jpPJ2J8og8phMUFHW0KuwVXH4nreKed1o0SxX2Z6e7fQOg2XyE=@vger.kernel.org X-Gm-Message-State: AOJu0YxXKq+DAv8th+a61FD5jZzfJZhkmqs/wgpcTnn9miEy9oTKXB9c 3rkY+f1WL4rZTGry8XWukE3tuLMIUkQvKptKZS3UywXgqPTb8e1a6dHB1/o+dBkO8Q== X-Gm-Gg: AfdE7cntKAx3WQk2p5RPLJEUs76j79Pd56DgWJePX8M1dQoBdSCsLVo58yW96rgXn6F ZBOLhfkl1GzoZ356Coepx3vBKJ1agMGUnc2cBB5Ts0Q0JU4VN5iutgBU87hpZH6+XypXMb7jOYV +P5gzi/Pan08zu/GFCgju0GPjcyCbeOO1QOzlKDQ/76w1FQ9LMv0aalhiogyp/KKHLVVHRs0dki /RJ0ujnibT/QaUG0jByuvIGHyEgXxTO/92yHmMTVHuW41iacXFRlLfKS/sqxQD2iCcHRGEkNwPY 8uSBMMul0GQkHOvyxvqPiCol+6gqVRSQqCVUcXTMJBX/cPP8dNNdE9bvrTzQX7Gw6QABnONCgkE IQPRG3y6JnHfwC387lNHp/xBoza1izx4n+xY710oqaWhlfzlRy081WzJc8+1LEoHWvBqaey96 X-Received: by 2002:a05:600c:3143:b0:493:faf3:3ea5 with SMTP id 5b1f17b1804b1-4954a3d0925mr109816675e9.4.1784458427000; Sun, 19 Jul 2026 03:53:47 -0700 (PDT) Received: from mf-00-01.. ([194.220.239.180]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-4954a2a24b6sm197679575e9.3.2026.07.19.03.53.45 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sun, 19 Jul 2026 03:53:46 -0700 (PDT) From: Mykola Marzhan To: Jens Axboe , Song Liu , Yu Kuai , Keith Busch , Christoph Hellwig , Sagi Grimberg , linux-block@vger.kernel.org, linux-raid@vger.kernel.org, linux-nvme@lists.infradead.org Cc: Li Nan , Xiao Ni , Guoqing Jiang , Leon Romanovsky , Jason Gunthorpe , Kiran Kumar Modukuri , Chaitanya Kulkarni , Logan Gunthorpe , Bjorn Helgaas , Shivaji Kant , Pranjal Shrivastava , Henrique Carvalho , linux-kernel@vger.kernel.org, linux-rdma@vger.kernel.org, linux-pci@vger.kernel.org Subject: [PATCH v2 6/7] md/raid1,raid10: skip futile retries on P2PDMA mapping failures Date: Sun, 19 Jul 2026 10:53:26 +0000 Message-ID: <20260719105327.864949-7-mykola@meshstor.io> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260719105327.864949-1-mykola@meshstor.io> References: <20260719105327.864949-1-mykola@meshstor.io> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Since commit 02666132403a ("md: propagate BLK_FEAT_PCI_P2PDMA from member devices to RAID device") raid1 and raid10 arrays accept P2PDMA (peer device memory) bios, and a member that cannot DMA-map the peer's pages fails its leg bio with BLK_STS_TARGET (from blk_dma_map_iter_start()). That mapping failure is a property of the peer-device/member pairing, not of the medium: re-submitting the same peer pages to the same member cannot succeed, and there is nothing on the disk to repair. Routing it through the stock error machinery misfires on every path: - narrow_write_error() re-issues the failed write in badblock-sized chunks, serializing hundreds of guaranteed-to-fail synchronous bios through raid1d/raid10d per failed write (256 for a 1 MiB write on 4K logical blocks, 2048 on 512e). - the write-error handler sets WantReplacement, so a hot spare is pulled in, rebuilt onto, and the healthy member is then evicted by spare_active() -- and if the spare is also unreachable from the peer, the cycle consumes the next spare. - fix_read_error() probes members with kernel pages, starting with the failing member itself. That probe succeeds (the member is healthy for host memory), so the routine rewrites nothing, records nothing, and logs nothing -- but each invocation costs a full freeze_array() quiesce and a tick of the read-error budget. The budget exists to evict members whose medium keeps producing corrected errors; its hourly decay is sized for sporadic medium errors, while mapping failures are deterministic and arrive at I/O rate. On an asymmetric PCIe topology a mixed host/P2P read workload charges 20 errors to the unreachable leg within a second and kicks a perfectly healthy member. - On FailFast members both paths short-circuit into md_error(), so a single unroutable peer-memory I/O evicts a healthy mirror outright. Track P2PDMA masters with a new r1bio/r10bio state bit, set at submission where bio_has_data() is still meaningful; a leg's mapping failure is then "BLK_STS_TARGET and the master carries P2PDMA pages", cheap to test on every completion path. Handle it explicitly: - Writes: probe the whole range with one retry and record one bad range if it also fails. BLK_STS_TARGET is also produced for device conditions that narrow_write_error()'s retry loop recovers from (e.g. NVME_SC_NS_NOT_READY and NVME_SC_CMD_INTERRUPTED via nvme_error_status()), and md cannot tell the two apart from bi_status alone, so it must not skip the retry outright. A single whole-range probe observes the outcome instead of predicting it. WriteErrorSeen is still set (it gates the write-path badblocks consult), but WantReplacement is not: replacing a member cannot fix a topology property, and spending a spare on it destroys redundancy management for real failures. - Reads: mark the leg IO_BLOCKED so read_balance() picks another mirror, and skip the freeze/fix cycle and the budget charge. The same TARGET ambiguity exists here, but a redirected read leaves nothing to fence, and fix_read_error()'s host-page probe fails the same way under e.g. NVME_SC_NS_NOT_READY -- charging the budget would evict a healthy member for a firmware activation window. A genuinely failing member is still evicted via host reads, writes and BLK_STS_MEDIUM errors. If no mirror can serve the read the master bio fails with EIO as before (md completes masters by Uptodate state, as for any failed mirror I/O), now without kicking healthy members on the way. - FailFast: a mapping failure is rejected at map time and never reaches the wire, so it is no evidence of device unreliability; don't let it trigger the FailFast md_error() -- the same reasoning commit f7b24c7b41f2 ("md/raid1,raid10: don't fail devices for invalid IO errors") applied to BLK_STS_INVAL. raid10 replacement legs keep their stock policy: badblocks are never recorded on a replacement, so failing it is the only outcome that cannot leave a silent hole in a rebuilding replacement. This belongs in the same release as the BLK_STS_TARGET restoration: with that fix alone, an unroutable leg costs the retry storms and the healthy-member evictions above. Measured on QEMU q35 rigs (one CMB provider, two NVMe members), 8KiB peer-memory I/O, reads x90 under concurrent host reads: scenario stock error machinery with this patch asym reads 90/90 ok, healthy far 90/90 ok, no eviction, leg kicked after ~20 no budget charge unreach reads all EIO, one healthy all EIO, no members member kicked kicked p2p TARGET 16 chunk retries from 1 whole-range probe, write raid1d, then badblocks same badblocks, master ok A transient TARGET recovers through the single probe with no badblocks recorded, and a mapping failure no longer trips FailFast eviction while a genuine I/O error still does -- both verified on the same rig. Fixes: 02666132403a ("md: propagate BLK_FEAT_PCI_P2PDMA from member devices= to RAID device") Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Mykola Marzhan --- drivers/md/raid1.c | 56 +++++++++++++++++++++++++++++++----- drivers/md/raid1.h | 2 ++ drivers/md/raid10.c | 69 ++++++++++++++++++++++++++++++++++++++------- drivers/md/raid10.h | 2 ++ 4 files changed, 112 insertions(+), 17 deletions(-) diff --git a/drivers/md/raid1.c b/drivers/md/raid1.c index f562b6bd438b..61f635463475 100644 --- a/drivers/md/raid1.c +++ b/drivers/md/raid1.c @@ -483,12 +483,22 @@ static void raid1_end_write_request(struct bio *bio) * 'one mirror IO has finished' event handler: */ if (bio->bi_status && !ignore_error) { + /* + * A P2PDMA mapping failure reflects the peer/member + * pairing, not member health: don't pull in a spare + * or trip FailFast for it. + */ + bool p2pdma_unmappable =3D bio->bi_status =3D=3D BLK_STS_TARGET && + test_bit(R1BIO_P2PDMA, &r1_bio->state); + set_bit(WriteErrorSeen, &rdev->flags); - if (!test_and_set_bit(WantReplacement, &rdev->flags)) + if (!p2pdma_unmappable && + !test_and_set_bit(WantReplacement, &rdev->flags)) set_bit(MD_RECOVERY_NEEDED, & conf->mddev->recovery); =20 - if (test_bit(FailFast, &rdev->flags) && + if (!p2pdma_unmappable && + test_bit(FailFast, &rdev->flags) && (bio->bi_opf & MD_FAILFAST) && /* We never try FailFast to WriteMostly devices */ !test_bit(WriteMostly, &rdev->flags)) { @@ -1378,6 +1388,8 @@ static void raid1_read_request(struct mddev *mddev, s= truct bio *bio, else init_r1bio(r1_bio, mddev, bio); r1_bio->sectors =3D max_read_sectors; + if (md_bio_is_p2pdma(bio)) + set_bit(R1BIO_P2PDMA, &r1_bio->state); =20 /* * make_request() can abort the operation when read-ahead is being @@ -1557,6 +1569,8 @@ static bool raid1_write_request(struct mddev *mddev, = struct bio *bio, =20 r1_bio =3D alloc_r1bio(mddev, bio); r1_bio->sectors =3D max_sectors; + if (md_bio_is_p2pdma(bio)) + set_bit(R1BIO_P2PDMA, &r1_bio->state); =20 /* first select target devices under rcu_lock and * inc refcount on their rdev. Record them by setting @@ -2525,7 +2539,7 @@ static void fix_read_error(struct r1conf *conf, struc= t r1bio *r1_bio) } } =20 -static void narrow_write_error(struct r1bio *r1_bio, int i) +static void narrow_write_error(struct r1bio *r1_bio, int i, bool coarse) { struct mddev *mddev =3D r1_bio->mddev; struct r1conf *conf =3D mddev->private; @@ -2539,6 +2553,11 @@ static void narrow_write_error(struct r1bio *r1_bio,= int i) * It is conceivable that the bio doesn't exactly align with * blocks. We must handle this somehow. * + * With 'coarse', retry the whole range as one bio and record + * one bad range if it fails: for P2PDMA mapping failures, + * which fail every block identically, while the single retry + * still lets a cleared transient error recover. + * * We currently own a reference on the rdev. */ =20 @@ -2553,9 +2572,12 @@ static void narrow_write_error(struct r1bio *r1_bio,= int i) block_sectors =3D roundup(1 << rdev->badblocks.shift, lbs); =20 sector =3D r1_bio->sector; - sectors =3D ((sector + block_sectors) - & ~(sector_t)(block_sectors - 1)) - - sector; + if (coarse) + sectors =3D sect_to_write; + else + sectors =3D ((sector + block_sectors) + & ~(sector_t)(block_sectors - 1)) + - sector; =20 while (sect_to_write) { struct bio *wbio; @@ -2636,8 +2658,18 @@ static void handle_write_finished(struct r1conf *con= f, struct r1bio *r1_bio) * narrow down and record precise write * errors. */ + bool coarse =3D r1_bio->bios[m]->bi_status =3D=3D + BLK_STS_TARGET && + test_bit(R1BIO_P2PDMA, &r1_bio->state); + fail =3D true; - narrow_write_error(r1_bio, m); + /* + * A P2PDMA mapping failure fails the whole range + * identically: probe it once (coarse) instead of + * narrowing block by block. A transient TARGET + * recovers via the probe with nothing recorded. + */ + narrow_write_error(r1_bio, m, coarse); rdev_dec_pending(conf->mirrors[m].rdev, conf->mddev); } @@ -2664,6 +2696,9 @@ static void handle_read_error(struct r1conf *conf, st= ruct r1bio *r1_bio) { struct md_rdev *rdev =3D conf->mirrors[r1_bio->read_disk].rdev; struct bio *bio =3D r1_bio->bios[r1_bio->read_disk]; + /* evaluate before the bio_put() below */ + bool p2pdma_error =3D bio->bi_status =3D=3D BLK_STS_TARGET && + test_bit(R1BIO_P2PDMA, &r1_bio->state); struct mddev *mddev =3D conf->mddev; sector_t sector; =20 @@ -2683,6 +2718,13 @@ static void handle_read_error(struct r1conf *conf, s= truct r1bio *r1_bio) */ if (mddev->ro) { r1_bio->bios[r1_bio->read_disk] =3D IO_BLOCKED; + } else if (p2pdma_error) { + /* + * The peer cannot reach this member; nothing on the + * medium to fix. Skip the read-error budget and + * FailFast, just keep this leg out of the retry. + */ + r1_bio->bios[r1_bio->read_disk] =3D IO_BLOCKED; } else if (test_bit(FailFast, &rdev->flags)) { md_error(mddev, rdev); } else { diff --git a/drivers/md/raid1.h b/drivers/md/raid1.h index c98d43a7ae99..61b788a99d14 100644 --- a/drivers/md/raid1.h +++ b/drivers/md/raid1.h @@ -184,6 +184,8 @@ enum r1bio_state { R1BIO_MadeGood, R1BIO_WriteError, R1BIO_FailFast, +/* the master bio carries PCI P2PDMA (peer device memory) pages */ + R1BIO_P2PDMA, }; =20 static inline int sector_to_idx(sector_t sector) diff --git a/drivers/md/raid10.c b/drivers/md/raid10.c index f7ef903a3d4e..144d07b6029c 100644 --- a/drivers/md/raid10.c +++ b/drivers/md/raid10.c @@ -482,13 +482,24 @@ static void raid10_end_write_request(struct bio *bio) */ md_error(rdev->mddev, rdev); else { + /* + * A P2PDMA mapping failure reflects the + * peer/member pairing, not member health: don't + * pull in a spare or trip FailFast for it. + */ + bool p2pdma_unmappable =3D + bio->bi_status =3D=3D BLK_STS_TARGET && + test_bit(R10BIO_P2PDMA, &r10_bio->state); + set_bit(WriteErrorSeen, &rdev->flags); - if (!test_and_set_bit(WantReplacement, &rdev->flags)) + if (!p2pdma_unmappable && + !test_and_set_bit(WantReplacement, &rdev->flags)) set_bit(MD_RECOVERY_NEEDED, &rdev->mddev->recovery); =20 dec_rdev =3D 0; - if (test_bit(FailFast, &rdev->flags) && + if (!p2pdma_unmappable && + test_bit(FailFast, &rdev->flags) && (bio->bi_opf & MD_FAILFAST)) { md_error(rdev->mddev, rdev); } @@ -1170,6 +1181,9 @@ static void raid10_read_request(struct mddev *mddev, = struct bio *bio, */ gfp_t gfp =3D err_path ? (GFP_NOIO | __GFP_HIGH) : GFP_NOIO; =20 + if (md_bio_is_p2pdma(bio)) + set_bit(R10BIO_P2PDMA, &r10_bio->state); + if (slot >=3D 0 && r10_bio->devs[slot].rdev) { /* * This is an error retry, but we cannot @@ -1357,6 +1371,9 @@ static bool raid10_write_request(struct mddev *mddev,= struct bio *bio, sector_t sectors; int max_sectors; =20 + if (md_bio_is_p2pdma(bio)) + set_bit(R10BIO_P2PDMA, &r10_bio->state); + if ((mddev_is_clustered(mddev) && mddev->cluster_ops->area_resyncing(mddev, WRITE, bio->bi_iter.bi_sector, @@ -2786,7 +2803,7 @@ static void fix_read_error(struct r10conf *conf, stru= ct mddev *mddev, struct r10 } } =20 -static void narrow_write_error(struct r10bio *r10_bio, int i) +static void narrow_write_error(struct r10bio *r10_bio, int i, bool coarse) { struct bio *bio =3D r10_bio->master_bio; struct mddev *mddev =3D r10_bio->mddev; @@ -2800,6 +2817,11 @@ static void narrow_write_error(struct r10bio *r10_bi= o, int i) * It is conceivable that the bio doesn't exactly align with * blocks. We must handle this. * + * With 'coarse', retry the whole range as one bio and record + * one bad range if it fails: for P2PDMA mapping failures, + * which fail every block identically, while the single retry + * still lets a cleared transient error recover. + * * We currently own a reference to the rdev. */ =20 @@ -2814,9 +2836,12 @@ static void narrow_write_error(struct r10bio *r10_bi= o, int i) block_sectors =3D roundup(1 << rdev->badblocks.shift, lbs); =20 sector =3D r10_bio->sector; - sectors =3D ((r10_bio->sector + block_sectors) - & ~(sector_t)(block_sectors - 1)) - - sector; + if (coarse) + sectors =3D sect_to_write; + else + sectors =3D ((r10_bio->sector + block_sectors) + & ~(sector_t)(block_sectors - 1)) + - sector; =20 while (sect_to_write) { struct bio *wbio; @@ -2856,6 +2881,7 @@ static void handle_read_error(struct mddev *mddev, st= ruct r10bio *r10_bio) { int slot =3D r10_bio->read_slot; struct bio *bio; + bool p2pdma_error; struct r10conf *conf =3D mddev->private; struct md_rdev *rdev =3D r10_bio->devs[slot].rdev; =20 @@ -2868,17 +2894,28 @@ static void handle_read_error(struct mddev *mddev, = struct r10bio *r10_bio) * frozen. */ bio =3D r10_bio->devs[slot].bio; + /* evaluate before the bio_put() below */ + p2pdma_error =3D bio->bi_status =3D=3D BLK_STS_TARGET && + test_bit(R10BIO_P2PDMA, &r10_bio->state); bio_put(bio); r10_bio->devs[slot].bio =3D NULL; =20 if (mddev->ro) r10_bio->devs[slot].bio =3D IO_BLOCKED; - else if (!test_bit(FailFast, &rdev->flags)) { + else if (p2pdma_error) { + /* + * The peer cannot reach this member; nothing on the + * medium to fix. Skip the read-error budget and + * FailFast, just keep this leg out of the retry. + */ + r10_bio->devs[slot].bio =3D IO_BLOCKED; + } else if (test_bit(FailFast, &rdev->flags)) { + md_error(mddev, rdev); + } else { freeze_array(conf, 1); fix_read_error(conf, mddev, r10_bio); unfreeze_array(conf); - } else - md_error(mddev, rdev); + } =20 rdev_dec_pending(rdev, mddev); r10_bio->state =3D 0; @@ -2947,8 +2984,20 @@ static void handle_write_completed(struct r10conf *c= onf, struct r10bio *r10_bio) r10_bio->sectors, 0); rdev_dec_pending(rdev, conf->mddev); } else if (bio !=3D NULL && bio->bi_status) { + bool coarse =3D bio->bi_status =3D=3D + BLK_STS_TARGET && + test_bit(R10BIO_P2PDMA, + &r10_bio->state); + fail =3D true; - narrow_write_error(r10_bio, m); + /* + * A P2PDMA mapping failure fails the whole + * range identically: probe it once (coarse) + * instead of narrowing block by block. A + * transient TARGET recovers via the probe + * with nothing recorded. + */ + narrow_write_error(r10_bio, m, coarse); rdev_dec_pending(rdev, conf->mddev); } bio =3D r10_bio->devs[m].repl_bio; diff --git a/drivers/md/raid10.h b/drivers/md/raid10.h index ec79d87fb92f..a2e1554f77db 100644 --- a/drivers/md/raid10.h +++ b/drivers/md/raid10.h @@ -174,6 +174,8 @@ enum r10bio_state { R10BIO_Previous, /* failfast devices did receive failfast requests. */ R10BIO_FailFast, +/* the master bio carries PCI P2PDMA (peer device memory) pages */ + R10BIO_P2PDMA, R10BIO_Discard, }; #endif --=20 2.43.0 From nobody Sat Jul 25 03:47:43 2026 Received: from mail-wm1-f46.google.com (mail-wm1-f46.google.com [209.85.128.46]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BBAB339524E for ; Sun, 19 Jul 2026 10:53:50 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.46 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784458434; cv=none; b=XugOF7q15Mq+s99HNKrNZc0CgDYDiE8KWzxv1G4OfVRk3Wkm+vE3V3e3LwE7C7CKzxkoP2C5v8CT26J5J6h7beMe0J2qlCKqiqaYUSgsL33tS02vYWi6w+qfmWqcboAJN1biSQWWx2AiQbDaHZlfZ/GUAUHHe0NgABG5ZddtUL0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784458434; c=relaxed/simple; bh=Xid0awo2BPFgwYW3JA18b5cCaSsnbZ06zZ7HCLlf580=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=hx4zxkeOeMPtoGbhD6didGSLAdL1SFM3t0bTy1bSAwERcqpF13PabSS7t8t7Zv0fK/1xlkzySgBaWYBjtfONlQahS68z+MIcxR1IOpFjSXCS9pOZ1HUTAo5weBoy9AeCDZMUb73TlTALY25NEJ0Dfh/qyoPyzdeWi0sNV0S9Qdc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=meshstor.io; spf=pass smtp.mailfrom=meshstor.io; dkim=pass (2048-bit key) header.d=meshstor.io header.i=@meshstor.io header.b=J3s+1T1M; arc=none smtp.client-ip=209.85.128.46 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=meshstor.io Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=meshstor.io Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=meshstor.io header.i=@meshstor.io header.b="J3s+1T1M" Received: by mail-wm1-f46.google.com with SMTP id 5b1f17b1804b1-493d92b7db3so44326425e9.2 for ; Sun, 19 Jul 2026 03:53:50 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=meshstor.io; s=google; t=1784458429; x=1785063229; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=eF2tyKd52a4KRQgMXYm9m1FypwFqk/B2D9AlJ6vMM1U=; b=J3s+1T1My83t9a6sOGhu7PkD7F35JibCFhW0VmCBo2Ze29yjkDtpfX0+mleJw7Cpg2 l4IYr2m4Ar+yiBOqhWU1Kb8EiHYLUeVIn6lPnrS0gAtRnSKHddBuofM0ZJHKyujUv1rZ R5ygBA/muqr2tvz1ycLCaAjJOBOvD3ITy9WwVbjHu4I2MVYG4Vck2VRx1ZJ+RHWQKdXz eHCH0D7fF3BDoojIReQ/cpmb/61Wtp6CMYHcsHsCWDumwJQ+Ctwlfo38koAVxqRvHDOV XYvcLArdH/DuSb1+Naerha2LnDXiLPiZMFvMjhrdNiVN1hgze3ZQpu6RYqOzwdrWJF1N /NnA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784458429; x=1785063229; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=eF2tyKd52a4KRQgMXYm9m1FypwFqk/B2D9AlJ6vMM1U=; b=nPUgZ3DDMlk+WYkoZyyVUceuFE9cEsgAosbs17yQeWUb4fVR9JAOlILtMrqaSGewOy iGFlBhSeIsmrS6doSHusvnwtCIgavYLx7PgChHiexgWeke3pmgWdh74U6MJxTioVjpZo Q1AUCO0y55WdCSYjhFRK+Y7ikgt9Laff+PaOr1gWgQXMo2rOsyQki6+OvlOOQkUR96WT 6/r0v4RVlVGQqStDS3rk4Tlgdn/f8ZymHfbQpbNbvnKWdDdq8o2XCxUp5NUf3CZpjyOW O1OEyZA9fs3D4q0S2LqPGFx8xnnAiHP543nqN/opoYK5IS8KsrNdwtSCrAJzABVxT7oD RMKw== X-Forwarded-Encrypted: i=1; AHgh+RooZPowd39WrHlpbN6FJ2isf6HPku8P68+0HIJ6dKGDJmvJ75hUOxN3PGlTYchR6x8eG9TO+LQFnG2VrgQ=@vger.kernel.org X-Gm-Message-State: AOJu0YyLnHPsRV1sSikLwEZ3aSOwqvOpq3qJPYU9SVkDZiVb+h0i7F3T kQ4Tj/4R+S2cosFcylnJTF9CGC06mM2sk9dbGdqUO4YCKmprKangVWVZvoVtt/GN7A== X-Gm-Gg: AfdE7cmHlyVAs/QYz8cVl2b/t0MlkTbltzZevFC5wb5XlUMDDc/ITrG4jKZlejOSHCQ eSu6gn9n6A0AVKnaimGXZXSDi7akOQ2MdolCpv4I5PLjG4QjQCQsO3XpUItffaOOwCo6DiZDcHr BO8PaJkh2WqriPfDESFAsi5ZEzrd43JuaW1M5/XZIWDT1Ik40Fsgzza3eSGLsSZXx8yLQzE8gYc 0A1t41u5y+RlnXuSnchWTD3bSFUcmwQyJH57tDnZfd8MNSe+wqujPmgrM9SiwjdwgwTloLZxxaG Mq8SGK8kyxEXl760siQJd/K7wqiHDlYxLFHljbGh7S9Twr27BW3cUebJXPGqgAeN4qKhpLfNIIQ J1HVFYS75U9F4rbIlX/pqtGMw2UIEPXnIRgVKQLy4J8ZcsMaFvnkUPzzDpk7L4Q== X-Received: by 2002:a05:600c:528e:b0:495:4749:16a7 with SMTP id 5b1f17b1804b1-4954a3eff48mr107049845e9.14.1784458428769; Sun, 19 Jul 2026 03:53:48 -0700 (PDT) Received: from mf-00-01.. ([194.220.239.180]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-4954a2a24b6sm197679575e9.3.2026.07.19.03.53.47 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sun, 19 Jul 2026 03:53:48 -0700 (PDT) From: Mykola Marzhan To: Jens Axboe , Song Liu , Yu Kuai , Keith Busch , Christoph Hellwig , Sagi Grimberg , linux-block@vger.kernel.org, linux-raid@vger.kernel.org, linux-nvme@lists.infradead.org Cc: Li Nan , Xiao Ni , Guoqing Jiang , Leon Romanovsky , Jason Gunthorpe , Kiran Kumar Modukuri , Chaitanya Kulkarni , Logan Gunthorpe , Bjorn Helgaas , Shivaji Kant , Pranjal Shrivastava , Henrique Carvalho , linux-kernel@vger.kernel.org, linux-rdma@vger.kernel.org, linux-pci@vger.kernel.org Subject: [PATCH v2 7/7] nvme-rdma: return BLK_STS_TARGET for unsupported P2P transfers Date: Sun, 19 Jul 2026 10:53:27 +0000 Message-ID: <20260719105327.864949-8-mykola@meshstor.io> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260719105327.864949-1-mykola@meshstor.io> References: <20260719105327.864949-1-mykola@meshstor.io> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Since commit 23528aa3320a ("nvme: enable PCI P2PDMA support for RDMA transport") nvme-rdma accepts P2PDMA bios, but a mapping failure for peer memory the HCA cannot reach is misreported as a path error: ib_dma_map_sg() returns 0, discarding the -EREMOTEIO that dma_map_sgtable() documents for exactly this case, and the driver converts it to -EIO -> nvme_host_path_error(). Under the default multipath configuration (the multipath head node advertises BLK_FEAT_PCI_P2PDMA since commit fb0eeeed91f3 ("nvme-multipath: enable PCI P2PDMA for multipath devices")) the path-error status makes nvme_failover_req() requeue the bios with a fresh retry budget each cycle, and since the mapping failure is a deterministic property of the peer/device pairing the I/O simply never completes: a hot requeue livelock, and a stacked md mirror hangs instead of failing over to its other leg. With nvme_core.multipath=3DN the request burns nvme_max_retries requeues -- nothing ever reaching the wire -- and completes as BLK_STS_TRANSPORT, which blk_path_error() classifies as retryable and md/raid1,raid10 treat as a genuine device error: retry storms on writes, read-error-budget eviction of a healthy member on reads (see the preceding md patches, whose mapping-failure handling keys on BLK_STS_TARGET). Map both the data and the metadata scatterlists with ib_dma_map_sgtable_attrs() so the DMA layer's error code is preserved, and translate -EREMOTEIO to BLK_STS_TARGET: the classification nvme-pci established in commit 91fb2b6052f7 ("nvme-pci: convert to using dma_map_sgtable()") and the preceding blk-mq-dma patch restores. The metadata path needs the same treatment because integrity buffers are not always host memory: bio_integrity_map_user() allows P2P pages whenever the queue advertises P2PDMA, and nvme-pci already maps P2P metadata. Returning BLK_STS_TARGET from queue_rq is a deliberate trade-off: it forecloses failover for a configuration where a sibling path through a different HCA could reach the peer. This matches nvme-pci, which has returned BLK_STS_TARGET for unsupported P2P mappings since that same commit and equally forecloses cross-controller failover; distinguishing this-path-unreachable from all-paths-unreachable at completion time would cost a spurious requeue cycle per I/O, and a stacking consumer (md) routes around the leg one layer up. Start the request only after mapping succeeds, as nvme-pci does: a mapping failure now errors out of queue_rq on a not-yet-started request, so nvme_mpath_start_request() accounting is never taken and cannot leak on the direct blk-mq completion (this also closes the same latent leak for mapping-failure BLK_STS_IOERR returns; the post-send error tail is untouched). A mapping -ENOMEM now takes the existing BLK_STS_RESOURCE branch instead of masquerading as a path error, and a DMA-layer -EINVAL completes as BLK_STS_IOERR instead of a path error; generic -EIO keeps today's host-path-error behavior, and virt-DMA devices are unaffected. Ratelimit the map-failure message -- with an md mirror steering peer-memory I/O around an unreachable leg it fires per redirected I/O, not per rare event. Fixes: 23528aa3320a ("nvme: enable PCI P2PDMA support for RDMA transport") Cc: stable@vger.kernel.org # v7.1 Assisted-by: Claude:claude-fable-5 Signed-off-by: Mykola Marzhan --- drivers/nvme/host/rdma.c | 42 +++++++++++++++++++++++++--------------- 1 file changed, 26 insertions(+), 16 deletions(-) diff --git a/drivers/nvme/host/rdma.c b/drivers/nvme/host/rdma.c index 6909e3542794..9017d927edc4 100644 --- a/drivers/nvme/host/rdma.c +++ b/drivers/nvme/host/rdma.c @@ -1469,6 +1469,7 @@ static int nvme_rdma_dma_map_req(struct ib_device *ib= dev, struct request *rq, int *count, int *pi_count) { struct nvme_rdma_request *req =3D blk_mq_rq_to_pdu(rq); + struct sg_table sgt; int ret; =20 req->data_sgl.sg_table.sgl =3D (struct scatterlist *)(req + 1); @@ -1480,12 +1481,14 @@ static int nvme_rdma_dma_map_req(struct ib_device *= ibdev, struct request *rq, =20 req->data_sgl.nents =3D blk_rq_map_sg(rq, req->data_sgl.sg_table.sgl); =20 - *count =3D ib_dma_map_sg(ibdev, req->data_sgl.sg_table.sgl, - req->data_sgl.nents, rq_dma_dir(rq)); - if (unlikely(*count <=3D 0)) { - ret =3D -EIO; + sgt =3D (struct sg_table) { + .sgl =3D req->data_sgl.sg_table.sgl, + .orig_nents =3D req->data_sgl.nents, + }; + ret =3D ib_dma_map_sgtable_attrs(ibdev, &sgt, rq_dma_dir(rq), 0); + if (unlikely(ret)) goto out_free_table; - } + *count =3D sgt.nents; =20 if (blk_integrity_rq(rq)) { req->metadata_sgl->sg_table.sgl =3D @@ -1501,14 +1504,14 @@ static int nvme_rdma_dma_map_req(struct ib_device *= ibdev, struct request *rq, =20 req->metadata_sgl->nents =3D blk_rq_map_integrity_sg(rq, req->metadata_sgl->sg_table.sgl); - *pi_count =3D ib_dma_map_sg(ibdev, - req->metadata_sgl->sg_table.sgl, - req->metadata_sgl->nents, - rq_dma_dir(rq)); - if (unlikely(*pi_count <=3D 0)) { - ret =3D -EIO; + sgt =3D (struct sg_table) { + .sgl =3D req->metadata_sgl->sg_table.sgl, + .orig_nents =3D req->metadata_sgl->nents, + }; + ret =3D ib_dma_map_sgtable_attrs(ibdev, &sgt, rq_dma_dir(rq), 0); + if (unlikely(ret)) goto out_free_pi_table; - } + *pi_count =3D sgt.nents; } =20 return 0; @@ -2026,8 +2029,6 @@ static blk_status_t nvme_rdma_queue_rq(struct blk_mq_= hw_ctx *hctx, if (ret) goto unmap_qe; =20 - nvme_start_request(rq); - if (IS_ENABLED(CONFIG_BLK_DEV_INTEGRITY) && queue->pi_support && (c->common.opcode =3D=3D nvme_cmd_write || @@ -2039,11 +2040,13 @@ static blk_status_t nvme_rdma_queue_rq(struct blk_m= q_hw_ctx *hctx, =20 err =3D nvme_rdma_map_data(queue, rq, c); if (unlikely(err < 0)) { - dev_err(queue->ctrl->ctrl.device, - "Failed to map data (%d)\n", err); + dev_err_ratelimited(queue->ctrl->ctrl.device, + "Failed to map data (%d)\n", err); goto err; } =20 + nvme_start_request(rq); + sqe->cqe.done =3D nvme_rdma_send_done; =20 ib_dma_sync_single_for_device(dev, sqe->dma, @@ -2063,6 +2066,13 @@ static blk_status_t nvme_rdma_queue_rq(struct blk_mq= _hw_ctx *hctx, ret =3D nvme_host_path_error(rq); else if (err =3D=3D -ENOMEM || err =3D=3D -EAGAIN) ret =3D BLK_STS_RESOURCE; + /* + * The DMA layer refused to map peer memory to this device: a + * property of the pairing, not a path failure. Match nvme-pci + * and do not retry (see blk_path_error()). + */ + else if (err =3D=3D -EREMOTEIO) + ret =3D BLK_STS_TARGET; else ret =3D BLK_STS_IOERR; nvme_cleanup_cmd(rq); --=20 2.43.0