From nobody Fri Jul 24 22:18:41 2026 Received: from mail-yw1-f174.google.com (mail-yw1-f174.google.com [209.85.128.174]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 08CA1459ADA for ; Wed, 22 Jul 2026 23:39:38 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.174 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763581; cv=none; b=uGxSybfExb9PFezwJ4kQ2qxpXxO6GieIQw3BhtY54SozqNUy9ZwEv6XvOxEIVeQXZq8f+6IPv/ne6YkNZ9KrkE6lWJe3XG4lZg0v+4oHq4TzWoUL1I1zrgu/++GVTPsn5hEU423x0K1mneLxrE75eiZtTVCz7RRnT7tAPZzh2y4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763581; c=relaxed/simple; bh=482j/rlGL/MyUIfK5vi1HCJAff/SkFOJEYzr2v72u54=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=W+BBuw0tH/hy1p+mGydJBPmDrRBhXuvMcv8+y1ZqWssIdeCryM1XcPdlmKXSeiymHG6w9e5aRjsEzqTaCamnyz2wTb2CZH36nCeJ0T3agpaimATQNJVMcMuPID65USELq/8zcXQ0gkzEjN6JPMv/sLyu8KOhErSgIY88YCE1p3M= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=x6u.co; spf=pass smtp.mailfrom=x6u.co; dkim=pass (2048-bit key) header.d=x6u.co header.i=@x6u.co header.b=NHuFYO5t; arc=none smtp.client-ip=209.85.128.174 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=x6u.co Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=x6u.co Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=x6u.co header.i=@x6u.co header.b="NHuFYO5t" Received: by mail-yw1-f174.google.com with SMTP id 00721157ae682-81e86df8987so619947b3.3 for ; Wed, 22 Jul 2026 16:39:38 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=x6u.co; s=google; t=1784763578; x=1785368378; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=czUeoqytFwl2f4wSjedp4ua+inBVMuooqHyWYibjU+0=; b=NHuFYO5t9QVIFDAK2j2M2SWFNgOG833xmV89KfrwdmjcPiRvUmGoahuI7+0CAZFZFp NELGx44Nt0bswc2It390av/3NuXepaatTbyI+p0mzWWvX0IZrva7BLnYYeE2x4J0MsxA wp2fBOffL3ZO1dRTUnbAni73wYdHwNS5uu5pyE9eOR7YkyEtxXRDqHYy9rSi7sRxUP2X NfGbmVR6HHzT+2w5ueGEaW8hg1Z82IJZDhzT+FtiCF6eDsNcWmceV+AaeD33+L8o71w9 N7X4iWnD19RhHS241GmR/5WmebRaLvJtqj60rRYUFLjZ2/8JBAlgjscOn4GipEjt4xh+ sGmw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784763578; x=1785368378; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=czUeoqytFwl2f4wSjedp4ua+inBVMuooqHyWYibjU+0=; b=F+8SuUGre2IZGscnVgBijPQLVvWe481VooJMGS4Q6OcnJiI6OFoMerEmZJ5HswOiLf 6Zt5b1oF2Yhmgyrr4iyO0CmuI1GjRiCyVmwjlAJF4D7lJBFZZ+9RJ3ypYpNwPanTWvXT bv33OPSOpgESlQIr+/zt78/s+lJJUeytg4HzemFh0X9IKFVFNtKGtQcnso/JsAdpglnO EonA8iCUPD1QqNlaVQTH8JcIod31hcOQ/2MPBMMZ1JL9vzhXJyfOobpkzGomLOshyhRR bYmr1Qzfgo6BiMyTE36hqswGNMh9hudYn1z1UD4OzlV29KJnIJQsHLtLoHlW4v7bTlkZ j5Vw== X-Forwarded-Encrypted: i=1; AHgh+Rrv/rURcQLVNKVqFzH5QDg3b+cKGhCTAv4DIDzgb0OYa7Wv8VhzECMQiXleLBMK1QeQPaMf4iCnEG6LrSo=@vger.kernel.org X-Gm-Message-State: AOJu0Ywr+lSqGFy1Yc61/e+fMOv9JumlE9L6wKbsHK5VomIvKYquLJdQ M/PLa2KIOb5Ivl+acFf6mOAtirIyuXf1QmbzD9swlavtbrUHJFez0RfkGm714lW2b4AT X-Gm-Gg: AR+sD13NhOn/MwebafD4Fq0PirENtFBPPelq1Gp9xeWHKkAMIv4SIpbbD/DPN8DWBXP CQLVhge8RpkkQKh2yoSuEk+iCZ2xp5EnKdhH4Hn2PpVYAsjv85miqMSSvDhHc7qHnP91TIpq+tS qmbQUMgU3m+fCtQnfTi4T5qIPYZCcCl0Tfh+icHlsIK5/HaBhIuv6SVJWQW204ByVaGB1bB2vd8 b2f+yGH2GqQFqsj0pbcsHyu/yjZmQ5sXiOzRNSDAkfRdu53djW0YU4TGkow9qsVdixgyU0RTJMm YY+z2nnljG1TDqb3JdVUQ+4LmBUucTa2Tt8NZnEeIA6MuRy3XUE331AVQMrR7IUrzXfrkFKPYEw peDTDCOiqFJv1UP7dmoD6yMmTj30fJzs4mBqyUnM2ISFrYunPDGJFhXcdC4YmD6bbI3bPL3ybce G3bkg4rC/iRJtFw7hbTSNeTrEKnPLN6GV67MsJvJ5Ws7PvQTxbErAgWhCsTYgAcP29Noo= X-Received: by 2002:a05:690c:9a87:b0:7ff:19d6:74ee with SMTP id 00721157ae682-81f4c31fe46mr2273147b3.33.1784763577509; Wed, 22 Jul 2026 16:39:37 -0700 (PDT) Received: from worldpeace10.c.googlers.com.com (234.207.85.34.bc.googleusercontent.com. [34.85.207.234]) by smtp.gmail.com with ESMTPSA id 956f58d0204a3-66890794b1asm2152505d50.20.2026.07.22.16.39.36 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 22 Jul 2026 16:39:37 -0700 (PDT) From: dhu@x6u.co To: Sumit Semwal , =?UTF-8?q?Christian=20K=C3=B6nig?= Cc: Jason Gunthorpe , David Laight , Nicolin Chen , Leon Romanovsky , Kevin Tian , Ankit Agrawal , Alex Williamson , Pranjal Shrivastava , linux-media@vger.kernel.org, dri-devel@lists.freedesktop.org, linaro-mm-sig@lists.linaro.org, linux-kernel@vger.kernel.org, iommu@lists.linux.dev, stable@vger.kernel.org, jmoroni@google.com, kpberry@google.com, chriscli@google.com, viursachi@google.com, xuehaohu@google.com, sashiko-bot Subject: [PATCH v3] dma-buf: Split sgl by largest page-aligned chunk Date: Wed, 22 Jul 2026 23:39:32 +0000 Message-ID: <20260722233932.3997681-1-dhu@x6u.co> X-Mailer: git-send-email 2.55.0.229.g6434b31f56-goog In-Reply-To: <20260623015459.1153884-1-xuehaohu@google.com> References: <20260623015459.1153884-1-xuehaohu@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: David Hu Currently, `fill_sg_entry()` splits the scatterlist using `UINT_MAX`. This creates a non-page-aligned DMA length (`0xFFFFFFFF`) for the first entry, resulting in non-page-aligned DMA addresses for all subsequent entries. While the underlying IOMMU mapping may be contiguous, hardware DMA engines often require explicit address alignment (e.g., page, cacheline, or storage sector boundaries). Passing unaligned addresses and lengths can cause explicit failures in DMA descriptor creation or silent data corruption if lower unaligned bits are truncated. In addition, a non-page-aligned sgl length will trigger an edge case in `ib_umem_find_best_pgsz()`. In case of a discontinuity in later buffers, we will have a `va` with lowest bit set to 1. That will lead to `ib_umem_find_best_pgsz()` always return 0, and break the promise to find best page size for the mapping on the NIC side. Fix this by splitting the scatterlist by the largest possible page aligned chunk within `UINT_MAX` (`ALIGN_DOWN(UINT_MAX, PAGE_SIZE)`). This ensures all scatterlist DMA addresses and lengths remain page aligned, while minimizing the total number of sgl entries. Page-aligned entries allow the system to cleanly chunk payloads into PCIe MaxPayloadSize (MPS) (e.g., 128 bytes, 256 bytes, 512 bytes). As a result, this may help reduce TLP fragmentation in P2P transfers and alleviate potential congestion within a logical PCIe switch partition, especially when Relaxed Ordering is not possible due to hardware constraints. Reported-by: sashiko-bot Closes: https://lore.kernel.org/all/20260609165431.778061F00893@smtp.kernel= .org/ Fixes: 3aa31a8bb11e ("dma-buf: provide phys_vec to scatter-gather mapping r= outine") Cc: stable@vger.kernel.org Signed-off-by: David Hu Reviewed-by: Leon Romanovsky --- Changes in v3: - Removed the type cast for `min` (David Laight) - Reverted max ent size to be `ALIGN_DOWN(UINT_MAX, PAGE_SIZE)` and updated commit message to reflect that (Jason Gunthorpe) - Updated commit message to reflect that this also fixes an edge case in `ib_umem_find_best_pgsz()` Changes in v2: - Updated commit title and message to reflect the switch to 2G chunks - Switch to using 2G as the max sg entry size as it naturally aligns with most hardware boundaries, while allowing compiler optimizations with bit shifts (David Laight) - Optimized away division calculation for `nent`, and multiplication calculation for sgl address, by dropping the `for` loop in favor of a `while (length)` loop (David Laight) - Dropped `min_t` in favor of `min()` to maintain a strict type checking safety net (David Laight) drivers/dma-buf/dma-buf-mapping.c | 19 +++++++++++-------- 1 file changed, 11 insertions(+), 8 deletions(-) diff --git a/drivers/dma-buf/dma-buf-mapping.c b/drivers/dma-buf/dma-buf-ma= pping.c index 794acff2546a..50ded9daf5fb 100644 --- a/drivers/dma-buf/dma-buf-mapping.c +++ b/drivers/dma-buf/dma-buf-mapping.c @@ -5,16 +5,17 @@ */ #include #include +#include + +#define MAX_SG_ENT_SZ ALIGN_DOWN(UINT_MAX, PAGE_SIZE) =20 static struct scatterlist *fill_sg_entry(struct scatterlist *sgl, size_t l= ength, dma_addr_t addr) { - unsigned int len, nents; - int i; + size_t len; =20 - nents =3D DIV_ROUND_UP(length, UINT_MAX); - for (i =3D 0; i < nents; i++) { - len =3D min_t(size_t, length, UINT_MAX); + while (length) { + len =3D min(length, MAX_SG_ENT_SZ); length -=3D len; /* * DMABUF abuses scatterlist to create a scatterlist @@ -24,8 +25,10 @@ static struct scatterlist *fill_sg_entry(struct scatterl= ist *sgl, size_t length, * does not require the CPU list for mapping or unmapping. */ sg_set_page(sgl, NULL, 0, 0); - sg_dma_address(sgl) =3D addr + (dma_addr_t)i * UINT_MAX; + sg_dma_address(sgl) =3D addr; sg_dma_len(sgl) =3D len; + addr +=3D len; + /* Unconditionally advance. On last segment, this becomes NULL */ sgl =3D sg_next(sgl); } =20 @@ -41,14 +44,14 @@ static unsigned int calc_sg_nents(struct dma_iova_state= *state, =20 if (!state || !dma_use_iova(state)) { for (i =3D 0; i < nr_ranges; i++) - nents +=3D DIV_ROUND_UP(phys_vec[i].len, UINT_MAX); + nents +=3D DIV_ROUND_UP(phys_vec[i].len, MAX_SG_ENT_SZ); } else { /* * In IOVA case, there is only one SG entry which spans * for whole IOVA address space, but we need to make sure * that it fits sg->length, maybe we need more. */ - nents =3D DIV_ROUND_UP(size, UINT_MAX); + nents =3D DIV_ROUND_UP(size, MAX_SG_ENT_SZ); } =20 return nents; --=20 2.55.0.229.g6434b31f56-goog