From nobody Fri Jul 24 22:18:41 2026 Received: from mail-yx1-f47.google.com (mail-yx1-f47.google.com [74.125.224.47]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0A1D742123D for ; Wed, 22 Jul 2026 23:38:28 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.224.47 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763511; cv=none; b=DjZydu69ricn+3rq7o8BSE1E9srHxJHKsTUWnU387YPhv/3bKR1r6JinKI/0lZQpMsoQEz+PfFRy1caPi520CkZlizzueOiQO7z2pWzTMvOq6gEXcca+Res0mg3zXhUBKYy6XOsP/4YoQgo5XLoY/cxzbPU7vKaUBYK3gBC+ij8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763511; c=relaxed/simple; bh=482j/rlGL/MyUIfK5vi1HCJAff/SkFOJEYzr2v72u54=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=CVaPYdPZC679iTyZTN5EtLtTdvWpuZeEKAW4Y4jzegoG6MqG5rrFWZtbI2rwd2IVy0yzE2W33RPTExJztbo3stHBfjGGD/ltywVApDVXK3ssCRCYrLOc0daB9IcC54GTk0Lpz+twYriEndzEEKwb6FkGEghRdnJQUrIdbeIAjUk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=x6u.co; spf=pass smtp.mailfrom=x6u.co; dkim=pass (2048-bit key) header.d=x6u.co header.i=@x6u.co header.b=iA+/kGZo; arc=none smtp.client-ip=74.125.224.47 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=x6u.co Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=x6u.co Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=x6u.co header.i=@x6u.co header.b="iA+/kGZo" Received: by mail-yx1-f47.google.com with SMTP id 956f58d0204a3-664ce3000e6so15167d50.0 for ; Wed, 22 Jul 2026 16:38:28 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=x6u.co; s=google; t=1784763508; x=1785368308; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=czUeoqytFwl2f4wSjedp4ua+inBVMuooqHyWYibjU+0=; b=iA+/kGZobaT+lSDLZ45F6N0UzEEl1f46UkH+BD25EBWjM2bPmi86e2wxn3LogtyAxV FsOeHVfoRSucvTww5WRhSAjQlSpZzL73YNbklZ7woMak9x6Xe7mUlbhfutEUSjfXofwq LgN/GwezkBhzMxfoIQztgpFRpdIuTiLPp1Y6/PbJrHIpM4mj8uEvUdsCjPaHZeOS/VO8 dmfqEj7Q2OQZ3ZG6OYNEKaXdFXsmauhMMxxxN0W+Vl4l6Sff58mbBPN2kwvj5/mfQuvk kPmZi/YWoE0IlVhhtYdnf0fyh3x750v8YortRxwSOIbjgyMDE+pWAc4KOiC5Sp9xXKWT 2ElA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784763508; x=1785368308; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=czUeoqytFwl2f4wSjedp4ua+inBVMuooqHyWYibjU+0=; b=dKYj4nU1U+gKAH8Ged23O3em5JVAPnNZ4qPMLqTPVC3ve1N6HYuXMCZn2OCjAlbDhH LPWaNVDf5qDf6dF7jpLs048I56twu1+4hMWYzGkWvrSDfQOl9sl29nXQOgRUVXtgVomb JD1Tm039w30m04VKVYZ4+/lTo1nKfsyxxw7M7qgFz3A+aK38lU0JjFwJzBjbknp50xcY L5Zc0kL9IdRTf2z1P/6yVZKI8Sx3WTCUBJfJQlB2J3TC1byyZr4iPlmZ6k7rubyX1zWd ZAdW1af8DsQUsyF8w7zCYLbYVivf79BKktSJe99OH0oLPvYykYjp0VzsLXvtH2OXjcCK 1+zg== X-Forwarded-Encrypted: i=1; AHgh+Rrs9NUIkvUF2CvmPgs1maoGW8NQppCKFC7LlezlBDRjvSfpQ8qWTOgXhgJpYD+0QhPYDr4Gj5TetholgkM=@vger.kernel.org X-Gm-Message-State: AOJu0Yw0vE2gl6QYe3yE0FmEnJKxXKMADtb4rr7ibKigE1u1sZnuERax N0vzAJf/8XxJ8cNmLJ92mJoRgZADnNSvI+8RQD8uFFhtgDuDnL4+HhqGNZo7HJQDlR3n X-Gm-Gg: AR+sD13irI1RKST9RJvyl4X6XGDXWdiBJ7w/pU7MrBDzdS8IQm5SC7K8GNCoLHctSJQ 8AELnZXUf7WrJ/Oy4W/na1bjplgrrE43FJhrH3jrF8LUb2gvWNfoZUXZfQYf65qPLO2PyJWlH0B 0/Txy4lVBjgehgNgULtTQKUpGWcoRaAhk2faZE0h+27EKWBA3qv9iYL2ZbhujH+dS7urcpkvsRR nucxZMWasG0YL3zGJrjFVeEs199/oGgCPBsXN7bV0fRdnZWPaJGSKrqhp4DZ50oPcKP7Cj5F/ZC 06aF1no3EZ7vmE/YzxS2Y003DqjiL1kbC0K9C6X/S/WDFNLP8F/l4DqfB05qKIPkkKXgeUg39kW Ebkkpbgp2xcOlqOGDg+uAllGjUO1ZJfucq7HoL+ZCj19huxhgExP+/XQY8mBuPRg4WKJpiZdu2t 4MbXL4RIVtqiYeGtDr5pO1s9UWZ03K/R/UinW78Pmu14MBCIyCY2r1xs9H X-Received: by 2002:a53:ce84:0:b0:667:af05:2e8d with SMTP id 956f58d0204a3-668a4fbd261mr142804d50.79.1784763507854; Wed, 22 Jul 2026 16:38:27 -0700 (PDT) Received: from worldpeace10.c.googlers.com.com (234.207.85.34.bc.googleusercontent.com. [34.85.207.234]) by smtp.gmail.com with ESMTPSA id 956f58d0204a3-66890521188sm2193219d50.1.2026.07.22.16.38.26 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 22 Jul 2026 16:38:27 -0700 (PDT) From: dhu@x6u.co To: Sumit Semwal , =?UTF-8?q?Christian=20K=C3=B6nig?= Cc: Jason Gunthorpe , David Laight , Nicolin Chen , Leon Romanovsky , Kevin Tian , Ankit Agrawal , Alex Williamson , Pranjal Shrivastava , linux-media@vger.kernel.org, dri-devel@lists.freedesktop.org, linaro-mm-sig@lists.linaro.org, linux-kernel@vger.kernel.org, iommu@lists.linux.dev, stable@vger.kernel.org, jmoroni@google.com, kpberry@google.com, chriscli@google.com, viursachi@google.com, xuehaohu@google.com, sashiko-bot Subject: [PATCH v3] dma-buf: Split sgl by largest page-aligned chunk Date: Wed, 22 Jul 2026 23:38:06 +0000 Message-ID: <20260722233806.3922093-1-dhu@x6u.co> X-Mailer: git-send-email 2.55.0.229.g6434b31f56-goog In-Reply-To: <20260623015459.1153884-1-xuehaohu@google.com> References: <20260623015459.1153884-1-xuehaohu@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: David Hu Currently, `fill_sg_entry()` splits the scatterlist using `UINT_MAX`. This creates a non-page-aligned DMA length (`0xFFFFFFFF`) for the first entry, resulting in non-page-aligned DMA addresses for all subsequent entries. While the underlying IOMMU mapping may be contiguous, hardware DMA engines often require explicit address alignment (e.g., page, cacheline, or storage sector boundaries). Passing unaligned addresses and lengths can cause explicit failures in DMA descriptor creation or silent data corruption if lower unaligned bits are truncated. In addition, a non-page-aligned sgl length will trigger an edge case in `ib_umem_find_best_pgsz()`. In case of a discontinuity in later buffers, we will have a `va` with lowest bit set to 1. That will lead to `ib_umem_find_best_pgsz()` always return 0, and break the promise to find best page size for the mapping on the NIC side. Fix this by splitting the scatterlist by the largest possible page aligned chunk within `UINT_MAX` (`ALIGN_DOWN(UINT_MAX, PAGE_SIZE)`). This ensures all scatterlist DMA addresses and lengths remain page aligned, while minimizing the total number of sgl entries. Page-aligned entries allow the system to cleanly chunk payloads into PCIe MaxPayloadSize (MPS) (e.g., 128 bytes, 256 bytes, 512 bytes). As a result, this may help reduce TLP fragmentation in P2P transfers and alleviate potential congestion within a logical PCIe switch partition, especially when Relaxed Ordering is not possible due to hardware constraints. Reported-by: sashiko-bot Closes: https://lore.kernel.org/all/20260609165431.778061F00893@smtp.kernel= .org/ Fixes: 3aa31a8bb11e ("dma-buf: provide phys_vec to scatter-gather mapping r= outine") Cc: stable@vger.kernel.org Signed-off-by: David Hu --- Changes in v3: - Removed the type cast for `min` (David Laight) - Reverted max ent size to be `ALIGN_DOWN(UINT_MAX, PAGE_SIZE)` and updated commit message to reflect that (Jason Gunthorpe) - Updated commit message to reflect that this also fixes an edge case in `ib_umem_find_best_pgsz()` Changes in v2: - Updated commit title and message to reflect the switch to 2G chunks - Switch to using 2G as the max sg entry size as it naturally aligns with most hardware boundaries, while allowing compiler optimizations with bit shifts (David Laight) - Optimized away division calculation for `nent`, and multiplication calculation for sgl address, by dropping the `for` loop in favor of a `while (length)` loop (David Laight) - Dropped `min_t` in favor of `min()` to maintain a strict type checking safety net (David Laight) drivers/dma-buf/dma-buf-mapping.c | 19 +++++++++++-------- 1 file changed, 11 insertions(+), 8 deletions(-) diff --git a/drivers/dma-buf/dma-buf-mapping.c b/drivers/dma-buf/dma-buf-ma= pping.c index 794acff2546a..50ded9daf5fb 100644 --- a/drivers/dma-buf/dma-buf-mapping.c +++ b/drivers/dma-buf/dma-buf-mapping.c @@ -5,16 +5,17 @@ */ #include #include +#include + +#define MAX_SG_ENT_SZ ALIGN_DOWN(UINT_MAX, PAGE_SIZE) =20 static struct scatterlist *fill_sg_entry(struct scatterlist *sgl, size_t l= ength, dma_addr_t addr) { - unsigned int len, nents; - int i; + size_t len; =20 - nents =3D DIV_ROUND_UP(length, UINT_MAX); - for (i =3D 0; i < nents; i++) { - len =3D min_t(size_t, length, UINT_MAX); + while (length) { + len =3D min(length, MAX_SG_ENT_SZ); length -=3D len; /* * DMABUF abuses scatterlist to create a scatterlist @@ -24,8 +25,10 @@ static struct scatterlist *fill_sg_entry(struct scatterl= ist *sgl, size_t length, * does not require the CPU list for mapping or unmapping. */ sg_set_page(sgl, NULL, 0, 0); - sg_dma_address(sgl) =3D addr + (dma_addr_t)i * UINT_MAX; + sg_dma_address(sgl) =3D addr; sg_dma_len(sgl) =3D len; + addr +=3D len; + /* Unconditionally advance. On last segment, this becomes NULL */ sgl =3D sg_next(sgl); } =20 @@ -41,14 +44,14 @@ static unsigned int calc_sg_nents(struct dma_iova_state= *state, =20 if (!state || !dma_use_iova(state)) { for (i =3D 0; i < nr_ranges; i++) - nents +=3D DIV_ROUND_UP(phys_vec[i].len, UINT_MAX); + nents +=3D DIV_ROUND_UP(phys_vec[i].len, MAX_SG_ENT_SZ); } else { /* * In IOVA case, there is only one SG entry which spans * for whole IOVA address space, but we need to make sure * that it fits sg->length, maybe we need more. */ - nents =3D DIV_ROUND_UP(size, UINT_MAX); + nents =3D DIV_ROUND_UP(size, MAX_SG_ENT_SZ); } =20 return nents; --=20 2.55.0.229.g6434b31f56-goog